The Question Was Answered. Just Not in This Room: Hanging Questions on Earnings Calls
We asked a language model a single question about 165,182 earnings calls from 1990 to 2026: was a question left hanging? It answered YES for 79,206 calls, or 47.95% of the corpus (95% CI 47.71% to 48.19%). Calls with a hanging question show lower candor (6.58 vs 6.86), lower specificity (7.19 vs 7.56), and higher stress (2.99 vs 2.43) than the rest. Guidance behavior skews negative: 15.4% lowered guidance versus 11.6% in the base. The share fell from 52.51% in 2015 to 44.47% in 2021, sitting near 46% since. Median subsequent return was -10.46% versus -7.16% for all calls.
- The model flagged a hanging question on 79,206 of 165,182 calls, a 47.95% share with a 95% CI of 47.71% to 48.19%.
- Flagged calls score lower on candor (6.58 vs 6.86), specificity (7.19 vs 7.56), and confidence (6.88 vs 7.21), and higher on stress (2.99 vs 2.43).
- Guidance was lowered on 15.40% of flagged calls versus 11.56% of base calls, and withdrawn on 4.16% versus 2.66%.
- The 'Skeptic Reassured' tag appears at 0.54x its base rate on flagged calls, the largest under-representation among top tags.
- Median subsequent return after flagged calls was -10.46% versus -7.16% for the 22,449-call returns sample overall.
1Introduction
Anyone who listens to earnings calls knows the feeling: an analyst presses, the executive pivots, and the clock runs out with the question dangling. Whether that dangling is noise or signal is an empirical question. Using a language model to score 165,182 calls from 1990 through 2026, we identify the 79,206 calls where a question was left hanging, roughly 48% of the corpus. We compare their language profiles, guidance actions, topical signature, and prevalence over time, and we report how subsequent returns distribute for a subset. The study examines what tends to co-occur with the unanswered question, not what causes what.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "The Question Left Hanging", tracked by year (n = 79,206; 48.0% of the reference set, 95% Wilson interval 47.7%–48.2%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Flagged calls read differently: candor runs 6.58 versus 6.86, specificity 7.19 versus 7.56, and stress 2.99 versus 2.43. Their guidance mix is heavier on the defensive side: 15.40% lowered and 4.16% withdrew guidance, against 11.56% and 2.66% in the base. The topical fingerprint is striking: 'Scale-Dependent Advantage Claims' appears at 2.02x its base rate, while 'Skeptic Reassured' appears at just 0.54x. The annual share peaked at 52.51% in 2015, bottomed at 44.47% in 2021, and has hovered between 45.57% and 47.62% since 2022. Median forward return for flagged calls was -10.46% versus -7.16% overall, though the distributions overlap heavily (Q1 -31.93%, Q3 +11.73%).
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.58 | 6.86 | -0.28 |
| Evasion | 3.27 | 2.70 | +0.58 |
| Specificity | 7.19 | 7.56 | -0.37 |
| Stress | 2.99 | 2.43 | +0.56 |
| Promotion | 5.21 | 5.05 | +0.16 |
| Confidence | 6.88 | 7.21 | -0.33 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 12.9% | 21.1% |
| Maintained | 42.8% | 48.8% |
| Lowered | 15.4% | 11.6% |
| Withdrawn | 4.2% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 2.02× | 22.3% | 11.1% |
| Results Worse Than Direction | 1.29× | 65.8% | 51.1% |
| Underused Fixed Costs | 1.25× | 52.1% | 41.6% |
| Skeptic Reassured | 0.54× | 35.9% | 66.4% |
| Guidance Worth Underwriting | 0.65× | 46.5% | 71.5% |
| Calls That Resolve Doubts | 0.72× | 57.5% | 79.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -10.5% | -7.2% |
| Interquartile range | -31.9% to +11.7% | — |
| Share beating SPY | 36.7% (95% CI 36%–38%) | 39.5% |
| Observations | 8,049 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| DOC | Q2 2025 | 2025-07-25 | C |
| HCA | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| ULH | Q2 2025 | 2025-07-25 | C+ |
| DBOEY | Q2 2025 | 2025-07-25 | B+ |
| TBBK | Q2 2025 | 2025-07-25 | C |
| LBTSF | Q2 2025 | 2025-07-25 | C+ |
| TRATF | Q2 2025 | 2025-07-25 | F |
4Discussion
A careful reader should treat these as co-occurrences. Calls where a question hangs tend to accompany lower-scoring candor and more guidance cuts, and their subsequent returns skew somewhat worse, but the interquartile range shows enormous overlap between groups. Nothing here establishes that the hanging question caused the guidance cut, the language profile, or the returns; the flagged calls may simply be calls about companies already under pressure. The right conclusion is descriptive: when a question goes unanswered, the surrounding call tends to look more defensive.
5Limitations
The flags come from AI-read fields, which are noisy and reflect the model's judgments rather than ground truth. The returns sample covers 22,449 calls and is skewed toward liquid names, so return statistics may not generalize. Our own forward tests falsified directional prediction from these signals. Finally, language models partially remember famous stocks' histories, which can contaminate any backtest of the same corpus; the comparisons here should be read as descriptive associations only. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.