No One Saw This Coming: What 142 'Unpredictable' Quarters Sound Like
This study examines earnings calls that answered YES to the hypothesis "This quarter could not have been described last quarter" — calls whose content genuinely exceeded what was knowable one quarter earlier. Of 318 calls scored between 2015 and 2024, 142 (44.65%) qualified. Compared with other calls, these quarters show slightly lower candor (6.78 vs 6.85) but higher specificity (7.72 vs 7.62), lower stress (1.98 vs 2.40), and higher confidence (7.64 vs 7.29). Guidance behavior diverges sharply: 28.87% raised guidance versus 21.07% elsewhere, and only 7.75% lowered it versus 15.72%. Forward 90-day returns for a 47-call subset had a median of -7.18% versus -10.54% for the base. The study is descriptive and offers no predictive claim.
- 142 of 318 calls (44.65%, 95% CI 39.29%-50.15%) answered YES to the hypothesis that the quarter could not have been described last quarter.
- These calls show lower stress (1.98 vs 2.40) and higher confidence (7.64 vs 7.29) than other calls, with specificity up (7.72 vs 7.62) and candor down slightly (6.78 vs 6.85).
- Guidance was raised on 28.87% of these calls versus 21.07% of others, and lowered on only 7.75% versus 15.72%; none withdrew guidance (0.0%).
- Among 47 calls with forward returns, the median 90-day return was -7.18% versus -10.54% for the base, with 42.55% beating the market versus 41.32% (CI 29.51%-56.72%).
1Introduction
Most earnings-call language is pre-scripted: themes, metrics, and even surprises can be previewed in the prior quarter's outlook. A small set of calls breaks that pattern — management describes a quarter that, by construction, could not have been telegraphed. For anyone who parses calls for a living, these moments are the interesting edge cases: do speakers become more guarded when they are off-script, or more expansive? Do they raise or cut guidance more often? Using AI-scored behavioral fields on calls from 2015 through 2024, this study characterizes the 142 calls (of 318) that answered YES to the hypothesis "This quarter could not have been described last quarter," and compares their language, guidance actions, and subsequent returns with the rest.
2Data & methodology
The corpus comprises 318 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "This quarter could not have been described last quarter" (n = 142; 44.7% of the reference set, 95% Wilson interval 39.3%–50.1%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile of these genuinely surprising quarters is distinctive: stress is lower (1.98 vs 2.40) and confidence higher (7.64 vs 7.29) than on other calls, while specificity edges up (7.72 vs 7.62) and candor slips slightly (6.78 vs 6.85). Promotion rises the most (5.61 vs 5.19, +0.42). Guidance actions diverge notably: 28.87% raised guidance versus 21.07% elsewhere, and 7.75% lowered it versus 15.72%, with no withdrawals. Over-indexed phrases include "Early Products Growing Fast" (lift 1.36) and "A Tiny Fraction of the Market" (1.32). Returns offer no clean edge: the 47-call subset's median 90-day return of -7.18% sits above the base's -10.54%, but the beat rate, 42.55%, is statistically indistinguishable from the base's 41.32%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.78 | 6.85 | -0.06 |
| Evasion | 2.64 | 2.73 | -0.09 |
| Specificity | 7.72 | 7.62 | +0.10 |
| Stress | 1.98 | 2.40 | -0.42 |
| Promotion | 5.61 | 5.19 | +0.42 |
| Confidence | 7.64 | 7.29 | +0.35 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 28.9% | 21.1% |
| Maintained | 52.1% | 50.3% |
| Lowered | 7.7% | 15.7% |
| Withdrawn | 0.0% | 1.3% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Early Products Growing Fast | 1.36× | 54.2% | 39.9% |
| A Tiny Fraction of the Market | 1.32× | 43.7% | 33.0% |
| Founder-Led Companies | 1.30× | 30.3% | 23.3% |
| Volume About to Step Up | 1.29× | 32.4% | 25.2% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -7.2% | -10.5% |
| Interquartile range | -34.4% to +12.5% | — |
| Share beating SPY | 42.6% (95% CI 30%–57%) | 41.3% |
| Observations | 47 | 121 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SFIX | Q3 2024 | 2024-06-04 | C+ |
| CRGO | Q1 2024 | 2024-05-20 | C+ |
| WRBY | Q1 2024 | 2024-05-09 | A |
| HCKT | Q1 2024 | 2024-05-08 | C |
| LINC | Q1 2024 | 2024-05-06 | B+ |
| AES | Q1 2024 | 2024-05-03 | C+ |
| PPC | Q1 2024 | 2024-05-03 | A |
| CTRA | Q1 2024 | 2024-05-03 | A |
4Discussion
The honest reading is descriptive. Calls answering YES to this hypothesis tend to coincide with confident, promotion-heavy language and more frequent guidance raises — an association, not a mechanism. The returns comparison is deliberately underwhelming: a higher median return against the base does not imply these calls outperform, because the 47-call returns sample is small, the confidence interval on the beat rate (29.51%-56.72%) spans the base rate, and medians can differ for many reasons. A careful reader should treat the phrase and guidance patterns as characterization of a corpus, not as signals. Nothing here supports predicting returns, selecting stocks, or assuming the labeled surprise is economically meaningful.
5Limitations
The behavioral fields (candor, evasion, specificity, stress, promotion, confidence) are AI-assigned and noisy; they should not be read as ground truth. The returns sample covers 47 calls out of 22,449 in the underlying corpus and is skewed toward liquid, widely covered names. Our own forward tests of this labeling approach falsified any directional prediction — the apparent return differences do not survive honest out-of-sample evaluation. Additionally, LLMs partially remember famous stocks' histories, which can contaminate labels and any backtest-looking statistic. All figures here are descriptive summaries of a scored corpus, not evidence of a tradable effect. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.
Companion page: every company matching this hypothesis is listed at the question’s own page.