The Question Left Hanging, Fewer and Farther Between: A Study of Calls That Resolve Doubts
This study examines earnings calls where the model answered YES to the battery item "Calls That Resolve Doubts" across a corpus spanning 1990 to 2026. Of 165,182 calls, 131,335 (79.51%) met the definition. These calls show a notably different behavioral profile: specificity of 7.75 versus a 7.56 baseline, confidence of 7.42 versus 7.21, and lower stress (2.12 vs 2.43) and evasion (2.49 vs 2.70). They are also overrepresented among "The Question Left Hanging" (0.72 observed vs 0.35 expected). Guidance was raised on 24.90% of these calls versus 21.05% in the base. In the returns sample of 19,998 calls, median excess return was -0.064% versus -0.072% baseline.
- 79.51% of the 165,182-call corpus (131,335 calls) were flagged as resolving doubts, with a 95% CI of 79.31% to 79.70%.
- These calls score higher on specificity (7.75 vs 7.56), confidence (7.42 vs 7.21), and candor (6.99 vs 6.86), and lower on stress (2.12 vs 2.43) and evasion (2.49 vs 2.70).
- "The Question Left Hanging" is overrepresented at 0.72 observed versus 0.35 expected (lift 2.11x), the largest lift of any flagged item.
- Guidance was raised on 24.90% of doubt-resolving calls versus 21.05% of the base, and median excess return was -0.064% versus -0.072%.
1Introduction
Anyone who listens to earnings calls knows the feeling of a call that settles the room: questions answered directly, uncertainty visibly reduced. Whether a call resolves doubts is one of the most basic qualities an analyst judges, yet it is rarely measured at scale. Because our corpus spans 1990 to 2026 and covers 165,182 calls, we can ask how common these calls are, what behavioral signatures accompany them, and how their guidance and returns compare to the rest. This study examines calls where the model answered YES to the battery item "Calls That Resolve Doubts" and profiles them against all other calls.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Calls That Resolve Doubts" (n = 131,335; 79.5% of the reference set, 95% Wilson interval 79.3%–79.7%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Doubt-resolving calls differ most sharply on stress and evasion, both lower than baseline (2.12 vs 2.43 and 2.49 vs 2.70), with specificity and confidence higher (7.75 vs 7.56; 7.42 vs 7.21). The overrepresentation table is striking: "The Question Left Hanging" appears at 0.72 versus 0.35 expected, suggesting calls can resolve most doubts while still leaving one open. Guidance actions lean positive (24.90% raised vs 21.05%; 8.93% lowered vs 11.56%). Returns show essentially no difference: median excess return of -0.064% versus -0.072%, with 40.31% beating versus 39.47% in the base. The trend series peaks at 83.34 in 2020 and falls to 74.25 in 2025.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.99 | 6.86 | +0.13 |
| Evasion | 2.49 | 2.70 | -0.21 |
| Specificity | 7.75 | 7.56 | +0.19 |
| Stress | 2.12 | 2.43 | -0.31 |
| Promotion | 5.01 | 5.05 | -0.04 |
| Confidence | 7.42 | 7.21 | +0.21 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 24.9% | 21.1% |
| Maintained | 52.3% | 48.8% |
| Lowered | 8.9% | 11.6% |
| Withdrawn | 2.2% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 0.33× | 3.6% | 11.1% |
| The Question Left Hanging | 0.72× | 34.7% | 48.0% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -6.4% | -7.2% |
| Interquartile range | -24.4% to +12.2% | — |
| Share beating SPY | 40.3% (95% CI 40%–41%) | 39.5% |
| Observations | 19,998 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| AON | Q2 2025 | 2025-07-25 | C |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| OMF | Q2 2025 | 2025-07-25 | A |
4Discussion
A careful reader should conclude that calls resolving doubts are common (79.51% of the corpus) and carry a coherent behavioral fingerprint: more specific, more confident, less stressed, less evasive. They also leave more hanging questions on average, which is a curious but descriptive pattern. What one should not conclude is that resolving doubts causes better outcomes, that the return differences of -0.064% versus -0.072% represent any exploitable signal, or that the 2020-2021 peaks (83.34 and 83.05) tell us anything about future call behavior. All findings here are descriptive.
5Limitations
The battery items and profile scores are AI-read fields and are inherently noisy; the model may misjudge candor, stress, or evasion on any given call. The returns sample covers only 22,449 calls overall (19,998 here) and is skewed toward liquid names, so results may not generalize. Our own forward tests falsified directional prediction, so nothing here should be read as an edge. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of model judgments against realized outcomes. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.