The Guidance Was Fine All Along: What Model-Endorsed Calls Look Like
We examine 165,182 earnings-call transcripts from 1990 through 2026 and isolate the 118,164 calls (71.5%, 95% CI 71.3% to 71.8%) where the model answered YES to the battery item 'Guidance Worth Underwriting'. These calls skew more candid (7.0 vs 6.86), more specific (7.77 vs 7.56), and more confident (7.42 vs 7.21), while showing less stress (2.1 vs 2.43) and less evasion (2.5 vs 2.7). Guidance actions differ too: 26.8% of these calls raised guidance versus 21.1% in the base set. Post-call returns medians were -0.063 versus -0.072, and 40.5% beat versus 39.5% in the base sample.
- 71.5% of 165,182 calls (CI 71.3%-71.8%) received a YES on 'Guidance Worth Underwriting' (n=118,164).
- YES calls show higher specificity (7.77 vs 7.56) and confidence (7.42 vs 7.21), and lower stress (2.1 vs 2.43) and evasion (2.5 vs 2.7).
- 26.8% of YES calls raised guidance versus 21.1% of base calls, while only 0.31% withdrew guidance versus 2.66% in the base set.
- In the matched returns sample (n=19,751), the YES-group median post-call return was -0.063 versus -0.072 for the base sample, with a beat rate of 40.5% versus 39.5%.
1Introduction
Anyone who follows earnings calls knows the ritual: management walks through the quarter, then offers a forward view the Street must decide whether to trust. The battery item 'Guidance Worth Underwriting' asks a simple question of each call - does the forward guidance hold up as something an analyst could actually lean on? Because the answer is available for 165,182 calls spanning 1990 through 2026, it offers a way to characterize what a credible-guidance call sounds like in language, what actions follow it, and how those calls behave afterward. This study describes the 118,164 YES calls and compares them with the full corpus on language profile, guidance actions, discourse markers, annual trends, and post-call returns.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Guidance Worth Underwriting" (n = 118,164; 71.5% of the reference set, 95% Wilson interval 71.3%–71.8%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The YES group reads differently before any outcome is known: candor runs 7.0 versus 6.86, specificity 7.77 versus 7.56, and confidence 7.42 versus 7.21, while stress (2.1 vs 2.43) and evasion (2.5 vs 2.7) run lower. Guidance actions line up with the label - 26.8% raised versus 21.1% in the base set, and withdrawals are rare (0.31% vs 2.66%). Two discourse markers lift: 'The Question Left Hanging' appears in 0.65% of YES calls versus 0.31% of base calls, and 'Scale-Dependent Advantage Claims' in 0.20% versus 0.11%. The annual share is mostly stable in the low-to-mid 70s but dips sharply to 58.58% in 2020, recovering to 73.23% in 2021. In the matched returns sample, medians were -0.063 versus -0.072 and beat rates 40.5% versus 39.5% - small, descriptive gaps.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 7.00 | 6.86 | +0.14 |
| Evasion | 2.50 | 2.70 | -0.19 |
| Specificity | 7.77 | 7.56 | +0.21 |
| Stress | 2.10 | 2.43 | -0.33 |
| Promotion | 4.95 | 5.05 | -0.11 |
| Confidence | 7.42 | 7.21 | +0.21 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 26.8% | 21.1% |
| Maintained | 56.0% | 48.8% |
| Lowered | 8.9% | 11.6% |
| Withdrawn | 0.3% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 0.20× | 2.2% | 11.1% |
| The Question Left Hanging | 0.65× | 31.2% | 48.0% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -6.3% | -7.2% |
| Interquartile range | -24.2% to +12.2% | — |
| Share beating SPY | 40.5% (95% CI 40%–41%) | 39.5% |
| Observations | 19,751 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
4Discussion
A careful reader should conclude that calls the model flags as having underwritable guidance tend to sound more candid, specific, and confident, and are more often followed by raised guidance. These are descriptions of co-occurrence within one labeling system, not evidence that the label causes anything or that the label helps predict returns. The returns gaps (median -0.063 vs -0.072; beat rate 40.5% vs 39.5%) are modest and could reflect many confounds, including which stocks have liquid trading data at all. The 2020 dip to 58.58% likely reflects the pandemic environment rather than a change in disclosure quality. Nothing here should be read as a trading signal.
5Limitations
The language scores are produced by AI readers and are inherently noisy; a YES label may partly reflect fluent writing rather than genuinely underwritable guidance. The returns sample covers only 19,751 of the YES calls out of 22,449 base calls and is skewed toward liquid names, so return comparisons may not generalize. Our own forward tests falsified directional prediction from these labels, so no claim of predictive power is made. Finally, LLMs partially remember famous stocks' histories from training data, which can contaminate any backtest by leaking future information into the labels themselves. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.