Research › Management Behavior
Artul.ai Research LibraryStudy No. 31Management BehaviorUpdated 2026-08-28

The Question Left Hanging: A Profile of High-Evasion Earnings Calls

By Artul.ai Research Group · n = 3,661 earnings calls · First published 2026-08-28
Abstract

This study examines earnings calls scoring 6 or higher on a 0-9 evasion meter, identifying 3,661 such calls among 165,182 transcripts spanning 1990-2026 (2.2%, 95% CI 2.1%-2.3%). High-evasion calls average 6.23 on the evasion scale versus 2.7 for the corpus, while scoring lower on candor (5.76 vs 6.86), specificity (6.26 vs 7.56), and confidence (6.48 vs 7.21), and higher on stress (4.11 vs 2.43). Guidance behavior diverges sharply: 7.9% of high-evasion calls withdrew guidance versus 2.7% baseline, while 21.1% raised it versus 8.7% overall. Among 371 high-evasion calls with forward returns, the median return was -13.0% versus -7.2% for the 22,449-call baseline, and 32.9% beat versus 39.5%.

Key findings
  • High-evasion calls number 3,661 of 165,182 transcripts (2.2%, 95% CI 2.1%-2.3%) across 1990-2026.
  • Guidance withdrawals appear on 7.9% of high-evasion calls versus 2.7% of all calls, roughly three times the base rate.
  • The median forward return for the 371 high-evasion calls with returns data was -13.0% versus -7.2% for the 22,449-call baseline.
  • The annual evasion-flag rate declined steadily from 3.93% of calls in 2015 to 1.38% in 2025 (through mid-year, 6,012 calls).

1Introduction

For anyone who parses earnings calls for a living, the hardest transcripts to read are not the ones where management stumbles, but the ones where nothing lands: questions get acknowledged, then quietly set down unanswered. A systematic evasion meter offers a way to find these calls at scale rather than by ear. If calls that dodge score measurably differently on related dimensions like candor, stress, and specificity, and if their guidance behavior looks unlike the rest of the corpus, that is a descriptive fingerprint worth having on record. It is also worth knowing what such calls look like in terms of language patterns and, descriptively, in forward returns. This study profiles the 3,661 calls scoring 6 or higher on the 0-9 evasion meter within a corpus of 165,182 transcripts from 1990-2026.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 evasion meter (n = 3,661; 2.2% of the reference set, 95% Wilson interval 2.1%–2.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The behavioral profile is coherent: high-evasion calls run 1.1 points lower on candor, 1.3 lower on specificity, and 1.68 higher on stress than the corpus, while promotion is nearly flat (0.13). Guidance mixes diverge most on withdrawals, 7.9% versus 2.7% baseline, and on raises, 21.1% versus 8.7%; lowered guidance is similar in both groups (11.9% vs 11.6%). Among over-represented language patterns, 'Scale-Dependent Advantage Claims' carries the largest lift at 2.63, followed by 'The Question Left Hanging' at 2.07; under-represented patterns include 'Calls That Resolve Doubts' (0.26) and 'Skeptic Reassured' (0.27). The annual flag rate fell from 3.93% in 2015 to 1.38% in 2025. Descriptively, median forward returns were -13.0% versus -7.2% baseline, with 32.9% beating versus 39.5%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor5.766.86-1.10
Evasion6.232.70+3.53
Specificity6.267.56-1.30
Stress4.112.43+1.68
Promotion5.185.05+0.13
Confidence6.487.21-0.73
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised8.7%21.1%
Maintained38.7%48.8%
Lowered11.9%11.6%
Withdrawn7.9%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims2.63×29.1%11.1%
The Question Left Hanging2.07×99.1%48.0%
Calls That Read Rehearsed1.35×52.5%38.7%
Results Worse Than Direction1.32×67.3%51.1%
The Finished-Story Tell1.27×5.6%4.4%
Calls That Resolve Doubts0.26×20.6%79.5%
Skeptic Reassured0.27×17.8%66.4%
Guidance Worth Underwriting0.42×29.7%71.5%
Confidence Proportionate to Evidence0.64×55.8%87.6%
Pricing Recovering0.67×14.3%21.5%
20153.93%
20163.07%
20172.90%
20182.67%
20192.55%
20202.22%
20211.81%
20221.71%
20231.60%
20241.54%
20251.38%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-13.0%-7.2%
Interquartile range-32.5% to +11.5%
Share beating SPY32.9% (95% CI 28%–38%)39.5%
Observations37122,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
HYMTFQ2 20252025-07-24F
UNPQ2 20252025-07-24D
WFGQ2 20252025-07-24F
NESRFQ4 20252025-07-24F
TSLAQ2 20252025-07-23F
THCQ2 20252025-07-23C+
VICRQ2 20252025-07-22F
SARTFQ2 20252025-07-22D

4Discussion

A careful reader should treat this as a descriptive portrait, not a warning label. High-evasion calls co-occur with guidance withdrawals, lower candor and specificity, and softer forward returns in this sample, but none of these tables establishes that evasion causes anything or predicts what comes next. The returns comparison especially invites over-reading: the flagged calls are a small, non-random slice of the corpus, and the companies that dodge questions may simply be the ones already in trouble. The declining flag rate over 2015-2025 could reflect changing disclosure norms, changing model behavior, or both. The honest conclusion is narrower: the evasion meter identifies a coherent, linguistically distinct, and comparatively rare slice of calls, and its readings travel together with other measures in ways that are internally consistent.

5Limitations

The candor, evasion, stress, and specificity scores are AI-generated fields and carry label noise, so the deltas reported here partly reflect measurement error. The returns sample covers 371 flagged calls against a baseline of 22,449 calls skewed toward liquid names, so the comparison is not representative of the full corpus. Our own forward tests falsified directional prediction: nothing here should be read as an edge. Finally, LLMs partially remember famous stocks' histories, which can contaminate any backtest, including the language-pattern lifts and the evasion labels themselves. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “The Question Left Hanging: A Profile of High-Evasion Earnings Calls.” Artul.ai Earnings-Call Research Library, Study No. 31. https://artul.ai/research/highly-evasive-management

Related studies

Specificity Is Next to Godliness: 161,829 Earnings CThe Truth Won't Set You Free: High Candor Is the NorNine Out of Ten Calls Are 'Confident' Now: Inside thStraight Talk, Stress Less: Low-Evasion Earnings CalThe Long Way to Say Less: Inside High-Complexity EarPromoted to the Front Page: High-Score Calls Across
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.