Ready, Set, Step Up: The Calls Where Volume Was Supposed to Arrive
This study examines earnings calls where Artul.ai's model answered YES to "Volume About to Step Up" — 47,036 of 165,182 calls, or 28.5% of the corpus. These calls show a promotional tilt: promotion scores run 5.58 vs 5.05 and confidence 7.51 vs 7.21 versus other calls, while stress is lower (2.33 vs 2.43). Guidance outcomes skew slightly positive (26.4% raised vs 21.1% baseline) but lowered guidance is also less common (8.7% vs 11.6%). Forward returns are weaker: a median of -10.4% vs -7.2% baseline, with 37.6% beating expectations vs 39.5%. The pattern describes tone and outcomes, not a tradable signal.
- Model-YES calls make up 28.5% of the corpus (47,036 of 165,182), with a confidence interval of 28.3%-28.7%.
- Tone skews promotional: promotion 5.58 vs 5.05 and confidence 7.51 vs 7.21, while stress is lower at 2.33 vs 2.43.
- Guidance was raised on 26.4% of YES calls vs 21.1% of baseline calls, and lowered on 8.7% vs 11.6%.
- Median forward return on YES calls (n=5,213) was -10.4% vs -7.2% for baseline calls, and 37.6% beat expectations vs 39.5% baseline.
1Introduction
Anyone who reads earnings calls knows a certain cadence: the moment a management team starts telegraphing that trading volume, order flow, or customer activity is about to accelerate. It is a confident, forward-leaning claim, and confident claims deserve scrutiny. If the model's YES to "Volume About to Step Up" marks a distinctive management posture, that posture should show up in tone scores, in guidance behavior, and — possibly — in what happens afterward. This study measures exactly that: 47,036 flagged calls across a 1990-2026 corpus, their language profile, their guidance outcomes, and their subsequent returns.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "Volume About to Step Up", tracked by year (n = 47,036; 28.5% of the reference set, 95% Wilson interval 28.3%–28.7%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The flagged calls sound more upbeat than the rest of the corpus: promotion scores of 5.58 vs 5.05 and confidence of 7.51 vs 7.21, with stress lower at 2.33 vs 2.43 and candor slightly lower at 6.78 vs 6.86. Guidance leans positive — 26.4% raised vs 21.1% baseline — though 8.7% still lowered guidance. Phrase lifts cluster around growth framing, with "Scale-Dependent Advantage Claims" over-represented 1.61x and "Early Products Growing Fast" 1.45x, while "When the CFO Dominates" appears 0.65x. Notably, the returns picture does not flatter the optimism: median forward return of -10.4% vs -7.2% baseline, and beats at 37.6% vs 39.5%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.78 | 6.86 | -0.08 |
| Evasion | 2.67 | 2.70 | -0.03 |
| Specificity | 7.58 | 7.56 | +0.02 |
| Stress | 2.33 | 2.43 | -0.10 |
| Promotion | 5.58 | 5.05 | +0.53 |
| Confidence | 7.51 | 7.21 | +0.30 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 26.4% | 21.1% |
| Maintained | 47.4% | 48.8% |
| Lowered | 8.7% | 11.6% |
| Withdrawn | 2.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 1.61× | 17.9% | 11.1% |
| Underused Fixed Costs | 1.46× | 60.6% | 41.6% |
| Early Products Growing Fast | 1.45× | 55.7% | 38.5% |
| A Tiny Fraction of the Market | 1.42× | 42.5% | 30.0% |
| Deferred Revenue Growing | 1.29× | 11.4% | 8.9% |
| When the CFO Dominates | 0.65× | 9.2% | 14.1% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -10.4% | -7.2% |
| Interquartile range | -32.0% to +12.2% | — |
| Share beating SPY | 37.6% (95% CI 36%–39%) | 39.5% |
| Observations | 5,213 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| FLG | Q2 2025 | 2025-07-25 | B |
| FRST | Q2 2025 | 2025-07-25 | A |
| HMDPF | Q2 2025 | 2025-07-25 | B |
| FFBC | Q2 2025 | 2025-07-25 | B+ |
| DBOEY | Q2 2025 | 2025-07-25 | B+ |
| TBBK | Q2 2025 | 2025-07-25 | C |
| SSB | Q2 2025 | 2025-07-25 | B+ |
| STEL | Q2 2025 | 2025-07-25 | B |
4Discussion
A careful reader should conclude that flagged calls carry a distinct communication style — more promotional, more confident, less stressed — and that this style coexists with slightly better guidance actions but slightly worse subsequent return medians. These are descriptive associations within one corpus. What should not be concluded is that the flag causes anything, predicts returns, or offers an edge; the beat-rate gap (37.6% vs 39.5%) is modest and the samples are not randomized. The year trend, peaking at 37.7% in 2021 before settling near 25-29%, is context, not signal.
5Limitations
The language fields are AI-read and inherently noisy, so small deltas like -0.08 in candor should not be over-read. The returns sample covers 5,213 flagged calls against a 22,449-call baseline skewed toward liquid names, so return comparisons may reflect listing characteristics rather than call content. Our own forward tests falsified directional prediction from these signals. Additionally, LLMs partially remember the histories of famous stocks, which can contaminate any backtest built on model-flagged language. Treat every figure here as descriptive of this dataset only. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.