Bad News Comes in Sixes: What High-Uncertainty Calls Sound Like
We examine 18,704 earnings calls scoring 6 or higher on a 0-9 uncertainty meter, drawn from a corpus of 165,182 calls spanning 1990 to 2026. These high-uncertainty calls make up 11.32% of the corpus (95% CI 11.17% to 11.48%). Management teams on them show sharply different behavior: confidence averages 5.96 versus 7.21 on all other calls, promotion framing drops to 4.15 versus 5.05, and stress rises to 3.66 versus 2.43. Guidance behavior diverges strongly: 32.59% lower guidance versus 11.56% in the comparison group, and 17.24% withdrawn versus 2.66%. Among 1,656 calls with measured next-day returns, the median return is -0.0967% versus -0.0716% for the 22,449-call base sample, with 36.96% beating versus 39.47%.
- High-uncertainty calls (score 6+) account for 11.32% of the 165,182-call corpus, with a 95% confidence interval of 11.17% to 11.48%.
- Confidence averages 5.96 on high-uncertainty calls versus 7.21 elsewhere, and stress averages 3.66 versus 2.43.
- Guidance is lowered on 32.59% of high-uncertainty calls and withdrawn on 17.24%, compared with 11.56% and 2.66% respectively in the base population.
- The trend peaks at 30.59% of calls in 2020, with 2025 at 14.47% on 6,012 calls.
1Introduction
Earnings calls are a performance, and uncertainty shows through the script. When management teams are unsure of their own numbers, the tell is not what they say but how they say it: less promotion, less confidence, more stress, more evasion. For anyone who reads calls closely, distinguishing routine hedging from genuine uncertainty could sharpen how they interpret a quarter. Using a 0-9 uncertainty meter, this study isolates 18,704 calls out of 165,182 that scored 6 or higher, spanning 1990 through 2026. It examines how these calls differ in language profile, guidance behavior, recurring narrative patterns, and measured next-day returns.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 uncertainty meter (n = 18,704; 11.3% of the reference set, 95% Wilson interval 11.2%–11.5%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The language profile is distinctive: confidence runs 5.96 versus 7.21, promotion 4.15 versus 5.05, and stress 3.66 versus 2.43, while evasion rises to 3.31 from 2.70. Guidance tells the same story: lowered on 32.59% of high-uncertainty calls versus 11.56% otherwise, and withdrawn on 17.24% versus 2.66%; raised guidance falls to 4.57% from 21.05%. Overrepresented narrative patterns include 'The Question Left Hanging' at 1.6x lift and 'Results Worse Than Direction' at 1.45x, while 'Guidance Worth Underwriting' appears at 0.49x. The trend spikes to 30.59% in 2020 and sits at 14.47% for 2025 on 6,012 calls. Median measured returns are -0.0967% versus -0.0716% in the base sample.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 7.10 | 6.86 | +0.24 |
| Evasion | 3.31 | 2.70 | +0.61 |
| Specificity | 7.31 | 7.56 | -0.25 |
| Stress | 3.66 | 2.43 | +1.23 |
| Promotion | 4.15 | 5.05 | -0.90 |
| Confidence | 5.96 | 7.21 | -1.26 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 4.6% | 21.1% |
| Maintained | 26.8% | 48.8% |
| Lowered | 32.6% | 11.6% |
| Withdrawn | 17.2% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| The Question Left Hanging | 1.60× | 76.9% | 48.0% |
| Results Worse Than Direction | 1.45× | 74.4% | 51.1% |
| Underused Fixed Costs | 1.45× | 60.5% | 41.6% |
| Scale-Dependent Advantage Claims | 1.37× | 15.1% | 11.1% |
| When the CFO Dominates | 1.32× | 18.6% | 14.1% |
| Guidance Worth Underwriting | 0.49× | 34.9% | 71.5% |
| Skeptic Reassured | 0.51× | 34.0% | 66.4% |
| Volume About to Step Up | 0.54× | 15.3% | 28.5% |
| Deferred Revenue Growing | 0.61× | 5.4% | 8.9% |
| Calls That Resolve Doubts | 0.64× | 50.5% | 79.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -9.7% | -7.2% |
| Interquartile range | -32.1% to +11.4% | — |
| Share beating SPY | 37.0% (95% CI 35%–39%) | 39.5% |
| Observations | 1,656 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| CNC | Q2 2025 | 2025-07-25 | F |
| MTH | Q2 2025 | 2025-07-25 | C |
| WF | Q2 2025 | 2025-07-25 | C |
| WZZAF | Q1 2026 | 2025-07-25 | F |
| MLLGF | Q2 2025 | 2025-07-25 | C+ |
| RNECF | Q2 2025 | 2025-07-25 | D |
| INTC | Q2 2025 | 2025-07-24 | D |
| SAM | Q2 2025 | 2025-07-24 | D |
4Discussion
High-uncertainty calls look, sound, and behave differently from the rest of the corpus: quieter promotion, more stress, more withdrawn and lowered guidance, and narrative patterns about unanswered questions. These are associations in a large descriptive dataset, not a rule for reading any single call. The return comparison is descriptive too, and a small median gap does not establish any trading edge. Careful readers should treat these figures as a map of what high-uncertainty calls tend to contain, not as a signal that predicts outcomes.
5Limitations
The uncertainty score and all language fields are AI-read and inherently noisy, so individual call labels should be treated as approximate. The returns sample covers 22,449 calls and is skewed toward liquid names, so measured returns may not generalize. Our own forward tests falsified directional prediction, and LLMs partially remember famous stocks' histories, which can contaminate any backtest of these labels. The study is descriptive throughout, and no causal or predictive claim is supported. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.