Dodging the Question, Formally: The Anatomy of an Unanswered Earnings Call
We examine 3,629 earnings calls (2.20% of a 165,182-call corpus spanning 1990-2026) where the model answered YES to "The Question Left Hanging" and evasion scored 6/9 or higher. These calls are behaviorally distinctive: evasion runs 6.23 vs. 2.70 baseline, while candor (5.75 vs. 6.86), specificity (6.25 vs. 7.56), and confidence (6.47 vs. 7.21) all fall short. Management guidance tells a matching story: 7.99% of these calls withdrew guidance versus 2.66% of the baseline. Subsequent one-quarter returns skew worse, with a median of -0.13 against -0.07 baseline, and only 31.9% beating expectations versus 39.5%. The share of such calls has also fallen steadily, from 3.93 in 2015 to 1.36 in 2025.
- Evasion scores average 6.23 on these calls versus 2.70 baseline, while candor (5.75 vs. 6.86) and specificity (6.25 vs. 7.56) sit well below typical.
- 7.99% of flagged calls withdrew guidance compared with 2.66% of all calls, and only 8.49% raised it versus 21.05%.
- Subsequent returns have a median of -0.13 versus -0.07 for the 22,449-call baseline, with 31.9% beating expectations (CI 27.3%-36.9%) against 39.5% overall.
- The flagged share of calls declined from 3.93 in 2015 to 1.36 in 2025, with the strongest top-tag lift being "Scale-Dependent Advantage Claims" at 2.65x.
1Introduction
An earnings call transcript can be fluent, confident, and still never answer the question that was asked. Analysts learn to listen for that gap, but reading thousands of calls by hand is impractical, and the signals are easy to talk yourself out of: was the silence evasion, or just discipline? This study isolates the extreme end of that behavior: calls where our model explicitly registered a question left hanging and where measured evasion reached 6 or higher on a 9-point scale. It profiles how those calls differ in tone, guidance behavior, analyst reception, and subsequent outcomes, and asks whether the pattern is becoming more or less common over time.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "The Question Left Hanging" AND evasion scored 6/9 or higher (n = 3,629; 2.2% of the reference set, 95% Wilson interval 2.1%–2.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile is coherent: evasion averages 6.23 against a 2.70 baseline, while candor (5.75 vs. 6.86), specificity (6.25 vs. 7.56), and confidence (6.47 vs. 7.21) all run below typical, and stress is elevated at 4.13 vs. 2.43. Guidance behavior mirrors the tone: 7.99% of these calls withdrew guidance versus 2.66% of the corpus, and 11.93% lowered it versus 11.56%. Tag-level lifts are modest, led by "Scale-Dependent Advantage Claims" at 2.65x, while "Calls That Resolve Doubts" and "Skeptic Reassured" appear far less often than expected (0.25x and 0.26x). In the returns sample, the median subsequent return is -0.13 versus -0.07 baseline, and 31.9% beat expectations against 39.5%. The annual share has fallen from 3.93 in 2015 to 1.36 in 2025.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 5.75 | 6.86 | -1.11 |
| Evasion | 6.23 | 2.70 | +3.53 |
| Specificity | 6.25 | 7.56 | -1.31 |
| Stress | 4.13 | 2.43 | +1.70 |
| Promotion | 5.18 | 5.05 | +0.13 |
| Confidence | 6.47 | 7.21 | -0.74 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 8.5% | 21.1% |
| Maintained | 38.6% | 48.8% |
| Lowered | 11.9% | 11.6% |
| Withdrawn | 8.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 2.65× | 29.4% | 11.1% |
| Calls That Read Rehearsed | 1.36× | 52.6% | 38.7% |
| Results Worse Than Direction | 1.32× | 67.6% | 51.1% |
| The Finished-Story Tell | 1.28× | 5.6% | 4.4% |
| Underused Fixed Costs | 1.27× | 52.8% | 41.6% |
| Calls That Resolve Doubts | 0.25× | 19.9% | 79.5% |
| Skeptic Reassured | 0.26× | 17.0% | 66.4% |
| Guidance Worth Underwriting | 0.41× | 29.1% | 71.5% |
| Confidence Proportionate to Evidence | 0.63× | 55.4% | 87.6% |
| Pricing Recovering | 0.66× | 14.1% | 21.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -13.3% | -7.2% |
| Interquartile range | -32.5% to +11.7% | — |
| Share beating SPY | 31.9% (95% CI 27%–37%) | 39.5% |
| Observations | 360 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| HYMTF | Q2 2025 | 2025-07-24 | F |
| UNP | Q2 2025 | 2025-07-24 | D |
| WFG | Q2 2025 | 2025-07-24 | F |
| NESRF | Q4 2025 | 2025-07-24 | F |
| TSLA | Q2 2025 | 2025-07-23 | F |
| THC | Q2 2025 | 2025-07-23 | C+ |
| VICR | Q2 2025 | 2025-07-22 | F |
| SARTF | Q2 2025 | 2025-07-22 | D |
4Discussion
A careful reader should conclude that calls flagged for evasion and an unanswered question are measurably different in tone, guidance behavior, and analyst reception, and that reception is on the whole less favorable. What follows is association, not cause: nothing here shows that evasion itself produces weaker outcomes, and the overlapping returns CIs caution against over-reading the gap. The declining annual share is a descriptive trend, not a forecast. Treat this as a descriptive anatomy of a communication style, useful for close reading of individual calls, not as a mechanical screen or a signal to trade on.
5Limitations
The fields underlying these scores are produced by AI readers and are noisy; individual calls can be mis-scored even when aggregate patterns hold. The returns comparison covers only 360 flagged calls against a 22,449-call baseline skewed toward liquid names, so the sample is not representative of the full corpus. Our own forward tests on similar constructs falsified directional prediction, and no result here should be read as an edge. Finally, LLMs partially remember the published history of famous stocks, which can contaminate any backtest of model-scored transcripts. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.