Straight Talk, Stress Less: Low-Evasion Earnings Calls, 1990–2026
We examined 86,188 earnings calls—52.2% of a 165,182-call corpus spanning 1990 to 2026—that scored 2 or lower on a 0–9 evasion meter. These candid calls ran higher on candor (7.12 vs 6.86), specificity (7.85 vs 7.56), and confidence (7.36 vs 7.21), while showing less evasion (1.79 vs 2.70) and stress (2.02 vs 2.43) than the base. Guidance profiles leaned more positive: 24.1% raised guidance versus 21.1% in the base, with only 1.7% withdrawn versus 2.7%. In the matched sample of 10,966 calls with returns, the median return was -0.061 versus -0.072 for the base, and 40.6% beat versus 39.5%.
- Low-evasion calls made up 52.2% of the 165,182-call corpus (86,188 calls, 95% CI 51.9%–52.4%).
- These calls scored higher on candor (7.12 vs 6.86) and specificity (7.85 vs 7.56) and lower on stress (2.02 vs 2.43) than the base.
- Guidance was raised after 24.1% of low-evasion calls versus 21.1% of base calls, and withdrawn after just 1.7% versus 2.7%.
- Among 10,966 calls with post-call returns, the median return was -0.061 versus -0.072 for the 22,449-call base, with 40.6% beating versus 39.5%.
1Introduction
Evasion on earnings calls is easy to sense and hard to quantify: dodged questions, vague scale claims, answers that end nowhere near where they started. Because call transcripts are now read by analysts, journalists, and machines alike, systematic differences in how directly management speaks are worth measuring at scale. Low-evasion calls are the natural contrast group—management answering what was asked. This study examines 86,188 calls scoring 2 or lower on a 0–9 evasion meter, drawn from a corpus of 165,182 calls covering 1990 through 2026, and compares their language profiles, guidance actions, and post-call returns against the rest of the corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 2 or lower on the 0–9 evasion meter (n = 86,188; 52.2% of the reference set, 95% Wilson interval 51.9%–52.4%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Low-evasion calls differ from the base on every language dimension measured: candor 7.12 vs 6.86, specificity 7.85 vs 7.56, confidence 7.36 vs 7.21, with evasion itself at 1.79 vs 2.70 and stress at 2.02 vs 2.43. The only overrepresented evasion topics are 'The Question Left Hanging' (55% over the base's 26.5% share) and 'Scale-Dependent Advantage Claims' (6.7% vs 11.1%). Guidance skews favorable: 24.1% raised versus 21.1% in the base, and 1.7% withdrawn versus 2.7%. Returns are modestly less negative—median -0.061 vs -0.072—and 40.6% beat versus 39.5%. The annual share rose from 48.0% in 2015 to 57.95% in 2025.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 7.12 | 6.86 | +0.26 |
| Evasion | 1.79 | 2.70 | -0.91 |
| Specificity | 7.85 | 7.56 | +0.29 |
| Stress | 2.02 | 2.43 | -0.41 |
| Promotion | 4.89 | 5.05 | -0.16 |
| Confidence | 7.36 | 7.21 | +0.15 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 24.1% | 21.1% |
| Maintained | 50.1% | 48.8% |
| Lowered | 10.4% | 11.6% |
| Withdrawn | 1.7% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| The Question Left Hanging | 0.55× | 26.5% | 48.0% |
| Scale-Dependent Advantage Claims | 0.60× | 6.7% | 11.1% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -6.1% | -7.2% |
| Interquartile range | -23.3% to +11.6% | — |
| Share beating SPY | 40.6% (95% CI 40%–42%) | 39.5% |
| Observations | 10,966 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| GBCI | Q2 2025 | 2025-07-25 | A |
| MOG.A | Q3 2025 | 2025-07-25 | B+ |
| FLG | Q2 2025 | 2025-07-25 | B |
| FRST | Q2 2025 | 2025-07-25 | A |
4Discussion
A careful reader should conclude that calls where management avoids dodging are, in this dataset, associated with slightly stronger language profiles, somewhat more favorable guidance actions, and marginally less negative median post-call returns. These are descriptive associations, not proof that candor causes better outcomes—firms that answer directly may simply differ in other ways from firms that do not. The gaps are also small in practical terms, and the returns comparison covers only a subset of calls. Nothing here should be read as a signal for predicting returns or picking trades.
5Limitations
Evasion, candor, and related scores are AI-read fields and inherently noisy; misclassification is expected at scale. The returns sample covers 22,449 base calls and is skewed toward liquid names, limiting generalizability. Our own forward tests falsified directional prediction from these features, so no predictive claim is made. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest built on model-scored transcripts, including ours. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.