The Order Book Is Fine, Thank You for Asking: Calls Where Backlog Reads as Shrinking
This study examines 7,699 earnings calls (4.7% of a 165,182-call corpus spanning 1990 to 2026) where backlog was read as shrinking. These calls sound measurably different: stress runs 3.07 versus a 2.43 baseline, while confidence (6.47 vs 7.21) and promotion (4.39 vs 5.05) fall short of typical calls. Guidance behavior diverges sharply: 28.7% of these calls lowered guidance versus 11.6% overall, and 6.8% withdrew it versus 2.7%. Language clusters skew toward 'Underused Fixed Costs' (1.44x) and away from 'Deferred Revenue Growing' (0.50x). Among 1,046 calls with matched returns, the median was -9.9% versus -7.2% for the 22,449-call baseline.
- Backlog-shrinking calls account for 4.7% of the corpus (7,699 of 165,182 calls), with a 95% confidence band of 4.6% to 4.8%.
- Guidance was lowered on 28.7% of these calls versus 11.6% of all calls, and withdrawn on 6.8% versus 2.7%.
- Stress scores average 3.07 versus a 2.43 baseline, while confidence averages 6.47 versus 7.21.
- The most overused phrase cluster is 'Underused Fixed Costs' at 1.44x baseline frequency; 'Deferred Revenue Growing' appears at 0.50x.
- Median next-period return on the 1,046-call matched sample was -9.9% versus -7.2% for the 22,449-call baseline, with 37.0% beating versus 39.5% overall.
1Introduction
Backlog is one of the few forward-looking numbers management teams volunteer unprompted, which makes it a favorite anchor for bullish narratives. When the backlog story turns negative, the way a call is conducted often changes with it: tone tightens, promotion drops, and guidance gets adjusted more often. For anyone who parses earnings calls for a living, distinguishing a routine order-book dip from a broader softening matters, and the linguistic fingerprint of a shrinking-backlog call is a useful reference point. This study examines 7,699 calls where backlog was read as shrinking, comparing their tone, guidance behavior, phrase usage, and next-period returns against the full 165,182-call corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where backlog was read as shrinking (n = 7,699; 4.7% of the reference set, 95% Wilson interval 4.6%–4.8%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The tonal profile is distinctive: stress is elevated (3.07 vs 2.43) and candor slightly higher (7.18 vs 6.86), while promotion (4.39 vs 5.05) and confidence (6.47 vs 7.21) sit below baseline, with specificity unchanged at 7.56. Guidance skews negative: 28.7% of these calls lowered guidance versus 11.6% corpus-wide, and 6.8% withdrew it versus 2.7%. Phrase usage points the same direction, with 'Underused Fixed Costs' (1.44x) and 'Results Worse Than Direction' (1.33x) overrepresented, and 'Deferred Revenue Growing' (0.50x) and 'Volume About to Step Up' (0.58x) underrepresented. The annual trend is volatile rather than steady, ranging from 1.9% of calls in 2021 to 7.26% in 2015.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 7.18 | 6.86 | +0.32 |
| Evasion | 2.74 | 2.70 | +0.04 |
| Specificity | 7.56 | 7.56 | +0.00 |
| Stress | 3.07 | 2.43 | +0.64 |
| Promotion | 4.39 | 5.05 | -0.66 |
| Confidence | 6.47 | 7.21 | -0.74 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 10.6% | 21.1% |
| Maintained | 37.9% | 48.8% |
| Lowered | 28.7% | 11.6% |
| Withdrawn | 6.8% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Underused Fixed Costs | 1.44× | 60.1% | 41.6% |
| Results Worse Than Direction | 1.33× | 68.0% | 51.1% |
| The Hidden Segment | 1.25× | 26.5% | 21.1% |
| Deferred Revenue Growing | 0.50× | 4.4% | 8.9% |
| Volume About to Step Up | 0.58× | 16.6% | 28.5% |
| Early Products Growing Fast | 0.67× | 25.9% | 38.5% |
| A Tiny Fraction of the Market | 0.68× | 20.4% | 30.0% |
| Pricing Recovering | 0.70× | 15.0% | 21.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -9.9% | -7.2% |
| Interquartile range | -30.5% to +10.9% | — |
| Share beating SPY | 37.0% (95% CI 34%–40%) | 39.5% |
| Observations | 1,046 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| HMDPF | Q2 2025 | 2025-07-25 | B |
| MTH | Q2 2025 | 2025-07-25 | C |
| WZZAF | Q1 2026 | 2025-07-25 | F |
| RGP | Q4 2025 | 2025-07-24 | F |
| DAIO | Q2 2025 | 2025-07-24 | D |
| FPH | Q2 2025 | 2025-07-24 | D |
| HZO | Q3 2025 | 2025-07-24 | D |
| JAKK | Q2 2025 | 2025-07-24 | D |
4Discussion
A careful reader should conclude that calls where backlog was read as shrinking travel together with a recognizable package: more stressed tone, more guidance cuts and withdrawals, and phrase patterns emphasizing spare capacity rather than growth. What should not be concluded is that the backlog reading caused any of this, or that it predicts returns. The median return gap (-9.9% vs -7.2%) is a descriptive difference in a non-random sample, and the 37.0% beat rate sits close to the 39.5% baseline. These are co-occurrences in historical data, not signals.
5Limitations
Backlog readings come from AI parsing of transcripts and are noisy; misclassifications are inevitable. The returns sample covers 1,046 of these calls against a 22,449-call baseline skewed toward liquid names, so price reactions may not generalize. Our own forward tests falsified directional prediction, and LLMs partially remember famous stocks' histories, contaminating any backtest. The share estimates carry a confidence band of roughly 4.6% to 4.8%, and the annual trend partly reflects shifting corpus composition across years. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.