Research › Business Verdicts
Artul.ai Research LibraryStudy No. 98Business VerdictsUpdated 2026-08-28

The Order Book Is Fine, Thank You for Asking: Calls Where Backlog Reads as Shrinking

By Artul.ai Research Group · n = 7,699 earnings calls · First published 2026-08-28
Abstract

This study examines 7,699 earnings calls (4.7% of a 165,182-call corpus spanning 1990 to 2026) where backlog was read as shrinking. These calls sound measurably different: stress runs 3.07 versus a 2.43 baseline, while confidence (6.47 vs 7.21) and promotion (4.39 vs 5.05) fall short of typical calls. Guidance behavior diverges sharply: 28.7% of these calls lowered guidance versus 11.6% overall, and 6.8% withdrew it versus 2.7%. Language clusters skew toward 'Underused Fixed Costs' (1.44x) and away from 'Deferred Revenue Growing' (0.50x). Among 1,046 calls with matched returns, the median was -9.9% versus -7.2% for the 22,449-call baseline.

Key findings
  • Backlog-shrinking calls account for 4.7% of the corpus (7,699 of 165,182 calls), with a 95% confidence band of 4.6% to 4.8%.
  • Guidance was lowered on 28.7% of these calls versus 11.6% of all calls, and withdrawn on 6.8% versus 2.7%.
  • Stress scores average 3.07 versus a 2.43 baseline, while confidence averages 6.47 versus 7.21.
  • The most overused phrase cluster is 'Underused Fixed Costs' at 1.44x baseline frequency; 'Deferred Revenue Growing' appears at 0.50x.
  • Median next-period return on the 1,046-call matched sample was -9.9% versus -7.2% for the 22,449-call baseline, with 37.0% beating versus 39.5% overall.

1Introduction

Backlog is one of the few forward-looking numbers management teams volunteer unprompted, which makes it a favorite anchor for bullish narratives. When the backlog story turns negative, the way a call is conducted often changes with it: tone tightens, promotion drops, and guidance gets adjusted more often. For anyone who parses earnings calls for a living, distinguishing a routine order-book dip from a broader softening matters, and the linguistic fingerprint of a shrinking-backlog call is a useful reference point. This study examines 7,699 calls where backlog was read as shrinking, comparing their tone, guidance behavior, phrase usage, and next-period returns against the full 165,182-call corpus.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where backlog was read as shrinking (n = 7,699; 4.7% of the reference set, 95% Wilson interval 4.6%–4.8%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The tonal profile is distinctive: stress is elevated (3.07 vs 2.43) and candor slightly higher (7.18 vs 6.86), while promotion (4.39 vs 5.05) and confidence (6.47 vs 7.21) sit below baseline, with specificity unchanged at 7.56. Guidance skews negative: 28.7% of these calls lowered guidance versus 11.6% corpus-wide, and 6.8% withdrew it versus 2.7%. Phrase usage points the same direction, with 'Underused Fixed Costs' (1.44x) and 'Results Worse Than Direction' (1.33x) overrepresented, and 'Deferred Revenue Growing' (0.50x) and 'Volume About to Step Up' (0.58x) underrepresented. The annual trend is volatile rather than steady, ranging from 1.9% of calls in 2021 to 7.26% in 2015.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor7.186.86+0.32
Evasion2.742.70+0.04
Specificity7.567.56+0.00
Stress3.072.43+0.64
Promotion4.395.05-0.66
Confidence6.477.21-0.74
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised10.6%21.1%
Maintained37.9%48.8%
Lowered28.7%11.6%
Withdrawn6.8%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Underused Fixed Costs1.44×60.1%41.6%
Results Worse Than Direction1.33×68.0%51.1%
The Hidden Segment1.25×26.5%21.1%
Deferred Revenue Growing0.50×4.4%8.9%
Volume About to Step Up0.58×16.6%28.5%
Early Products Growing Fast0.67×25.9%38.5%
A Tiny Fraction of the Market0.68×20.4%30.0%
Pricing Recovering0.70×15.0%21.5%
20157.26%
20165.81%
20173.59%
20182.81%
20194.10%
20206.91%
20211.90%
20224.61%
20236.95%
20244.93%
20254.03%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-9.9%-7.2%
Interquartile range-30.5% to +10.9%
Share beating SPY37.0% (95% CI 34%–40%)39.5%
Observations1,04622,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
HMDPFQ2 20252025-07-25B
MTHQ2 20252025-07-25C
WZZAFQ1 20262025-07-25F
RGPQ4 20252025-07-24F
DAIOQ2 20252025-07-24D
FPHQ2 20252025-07-24D
HZOQ3 20252025-07-24D
JAKKQ2 20252025-07-24D

4Discussion

A careful reader should conclude that calls where backlog was read as shrinking travel together with a recognizable package: more stressed tone, more guidance cuts and withdrawals, and phrase patterns emphasizing spare capacity rather than growth. What should not be concluded is that the backlog reading caused any of this, or that it predicts returns. The median return gap (-9.9% vs -7.2%) is a descriptive difference in a non-random sample, and the 37.0% beat rate sits close to the 39.5% baseline. These are co-occurrences in historical data, not signals.

5Limitations

Backlog readings come from AI parsing of transcripts and are noisy; misclassifications are inevitable. The returns sample covers 1,046 of these calls against a 22,449-call baseline skewed toward liquid names, so price reactions may not generalize. Our own forward tests falsified directional prediction, and LLMs partially remember famous stocks' histories, contaminating any backtest. The share estimates carry a confidence band of roughly 4.6% to 4.8%, and the annual trend partly reflects shifting corpus composition across years. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “The Order Book Is Fine, Thank You for Asking: Calls Where Backlog Reads as Shrinking.” Artul.ai Earnings-Call Research Library, Study No. 98. https://artul.ai/research/when-the-backlog-shrinks-earnings-calls

Related studies

Steady As She Goes: What Happens When Guidance HoldsMargins Are Expanding, and So Is the Confidence: ReaCapex Up, Stress Down: Reading 64,183 Calls Where SpGuidance Is a Vibe: Calls Where Demand Reads as AcceBacklog Is Growing, and So Is the Confidence: EarninMargins Are Fine, Thanks for Asking: What Contractio
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.