Research › Signal Combinations
Artul.ai Research LibraryStudy No. 11Signal CombinationsUpdated 2026-08-28

Loud Backlogs and Busy Phones: Calls Where the Model Saw a Volume Step-Up

By Artul.ai Research Group · n = 23,140 earnings calls · First published 2026-08-28
Abstract

This study examines 23,140 earnings calls—14.01% of the 165,182-call Artul.ai corpus spanning 1990 to 2026—where the model answered YES to "Volume About to Step Up" and the backlog was read as growing. Compared with the corpus baseline, these calls score higher on promotion (5.60 vs 5.05) and confidence (7.63 vs 7.21), with stress lower (2.21 vs 2.43). Guidance was raised on 32.36% of these calls versus 21.05% in the base. Language lifts are led by "Deferred Revenue Growing" at 2.17x. Among 2,598 calls with return data, the median next-window return was -7.76%, versus -7.16% for the 22,449-call base, and the beat rate was 39.53% versus 39.47%—no directional edge.

Key findings
  • The pattern covers 23,140 calls, 14.01% of the 165,182-call corpus (95% CI: 13.84% to 14.18%).
  • Confidence averages 7.63 vs 7.21 in the base, promotion 5.60 vs 5.05, and stress 2.21 vs 2.43.
  • Guidance was raised on 32.36% of these calls versus 21.05% of base calls, and lowered on 7.52% versus 11.56%.
  • Median next-window return among 2,598 eligible calls was -7.76% versus -7.16% for the base, with beat rates of 39.53% and 39.47% respectively.

1Introduction

Few signals on an earnings call sound as promising as a growing backlog paired with a claim that volume is about to step up. Management teams using this language talk a particular talk: more confident, more promotional, less stressed. Whether that tone accompanies better outcomes is a separate question from whether it appears. For anyone who parses calls for a living, the combination is common enough—about one call in seven in the Artul.ai library—that its behavioral fingerprint deserves description. This study examines the profile, guidance behavior, language lifts, and subsequent-return statistics of those calls.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "Volume About to Step Up" AND backlog was read as growing (n = 23,140; 14.0% of the reference set, 95% Wilson interval 13.8%–14.2%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Calls in this set run higher on promotion (5.60 vs 5.05) and confidence (7.63 vs 7.21) than the corpus baseline, with lower stress (2.21 vs 2.43) and lower evasion (2.58 vs 2.70). They lean optimistic on guidance: 32.36% raised versus 21.05% in the base, while only 7.52% lowered versus 11.56%. Language lifts are led by "Deferred Revenue Growing" at 2.17x and "Scale-Dependent Advantage Claims" at 1.42x; "When the CFO Dominates" appears underrepresented at 0.64x. The pattern peaks in 2021 at 21.72% of calls. Subsequent returns show a median of -7.76% versus -7.16% in the base—a descriptive gap, not an edge.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.836.86-0.04
Evasion2.582.70-0.11
Specificity7.667.56+0.10
Stress2.212.43-0.22
Promotion5.605.05+0.55
Confidence7.637.21+0.42
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised32.4%21.1%
Maintained46.0%48.8%
Lowered7.5%11.6%
Withdrawn1.7%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Deferred Revenue Growing2.17×19.2%8.9%
Early Products Growing Fast1.56×60.0%38.5%
Underused Fixed Costs1.44×60.1%41.6%
Scale-Dependent Advantage Claims1.42×15.7%11.1%
A Tiny Fraction of the Market1.36×40.8%30.0%
When the CFO Dominates0.64×9.1%14.1%
The Finished-Story Tell0.69×3.0%4.4%
20158.18%
201611.36%
201713.79%
201812.93%
201911.67%
202013.31%
202121.72%
202215.42%
202312.68%
202414.04%
202512.77%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-7.8%-7.2%
Interquartile range-28.0% to +13.0%
Share beating SPY39.5% (95% CI 38%–41%)39.5%
Observations2,59822,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
FLGQ2 20252025-07-25B
FRSTQ2 20252025-07-25A
FFBCQ2 20252025-07-25B+
DBOEYQ2 20252025-07-25B+
SSBQ2 20252025-07-25B+
STELQ2 20252025-07-25B
BAHQ1 20262025-07-25C+
VLOUFQ2 20252025-07-25C

4Discussion

A careful reader should conclude that this language pattern is common, tonally distinctive, and associated with more guidance raises. They should not conclude that it causes better outcomes or predicts them. The beat rate (39.53% vs 39.47%) and median returns (-7.76% vs -7.16%) are statistically indistinguishable in practical terms from the base. Every figure here is a description of co-occurrence in transcript data, not a trading signal, and the AI-read fields that define the pattern carry their own measurement error.

5Limitations

All fields—the YES/NO volume answer, backlog reading, and behavioral scores—are AI-generated and noisy; misclassification attenuates every comparison. The returns sample covers only 2,598 of 23,140 calls within a base of 22,449 calls skewed toward liquid names, so return statistics may not generalize. Our own forward tests falsified directional prediction on this pattern. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of transcript-derived signals, including the comparisons reported here. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Loud Backlogs and Busy Phones: Calls Where the Model Saw a Volume Step-Up.” Artul.ai Earnings-Call Research Library, Study No. 11. https://artul.ai/research/backlog-building-into-a-volume-step-up

Related studies

Small Pond, Big Fish: Calls Claiming Fast Growth in Trust the Raise: When Skeptics Get Reassured and GuiThe Founder Is Present: Candor and Confidence on EarFounder Knows Best: High-Promotion Founder-Led EarniThe Quiet Hand on the Microphone: CFO Dominance MeetDodging the Question, Formally: The Anatomy of an Un
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.