Smooth Talkers, Full Backlogs: Calls With No Critic Ammo and Growing Deferred Revenue
This study profiles earnings calls where the model found no critic ammunition (answered NO) and saw deferred revenue growing (answered YES) — 2,343 calls out of 165,182, or 1.42%. Compared with the base population, these calls show lower evasion (2.13 vs. 2.70), lower stress (1.29 vs. 2.43), higher specificity (7.89 vs. 7.56), and higher confidence (8.03 vs. 7.21). Guidance behavior diverges sharply: 55.23% raised guidance versus 21.06% in the base, while only 0.77% lowered it versus 11.56%. Among the 468 calls with measurable next-day outcomes, the median return was 10.06% versus -7.16% for the base, and 62.82% beat versus 39.47%. Lifts are strongest for "Early Products Growing Fast" (1.85x).
- Guidance was raised on 55.23% of these calls versus 21.06% in the base population, and lowered on just 0.77% versus 11.56%.
- Confidence scores average 8.03 versus 7.21 in the base, with stress at 1.29 versus 2.43 and evasion at 2.13 versus 2.70.
- The median next-day return in the 468-call returns sample was 10.06% versus -7.16% for the base, with 62.82% beating versus 39.47%.
- The strongest overrepresented language pattern is "Early Products Growing Fast" at 1.85x lift, while "Critic Ammunition" is absent by construction (0.0x).
1Introduction
Most earnings-call research hunts for smoke: evasion, hedging, the question left hanging. This study looks at the opposite corner of the corpus — calls where the model found no critic ammunition at all and observed deferred revenue growing. Deferred revenue is one of the few line items management cannot easily massage, so its growth alongside unusually confident, low-stress language makes these calls a distinctive population. They account for 2,343 of 165,182 calls spanning 1990 to 2026. The study examines how these calls differ in tone, guidance behavior, language patterns, and measured outcomes relative to the broader base.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered NO to "Critic Ammunition" AND the model answered YES to "Deferred Revenue Growing" (n = 2,343; 1.4% of the reference set, 95% Wilson interval 1.4%–1.5%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile is striking: confidence runs 8.03 versus 7.21, specificity 7.89 versus 7.56, while stress (1.29 vs. 2.43) and evasion (2.13 vs. 2.70) sit well below base. Guidance tilts heavily upward — 55.23% raised versus 21.06% base, and only 0.77% lowered versus 11.56%. Language lifts concentrate on growth storytelling: "Early Products Growing Fast" at 1.85x, "A Tiny Fraction of the Market" at 1.59x, and "Founder-Led Companies" at 1.57x. In the 468-call returns subset, the median return was 10.06% versus -7.16% base, and 62.82% beat versus 39.47% — though selection, not prediction, may explain it.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.92 | 6.86 | +0.05 |
| Evasion | 2.13 | 2.70 | -0.57 |
| Specificity | 7.89 | 7.56 | +0.33 |
| Stress | 1.29 | 2.43 | -1.14 |
| Promotion | 5.52 | 5.05 | +0.47 |
| Confidence | 8.03 | 7.21 | +0.82 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 55.2% | 21.1% |
| Maintained | 38.4% | 48.8% |
| Lowered | 0.8% | 11.6% |
| Withdrawn | 0.3% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Early Products Growing Fast | 1.85× | 71.3% | 38.5% |
| A Tiny Fraction of the Market | 1.59× | 47.8% | 30.0% |
| Founder-Led Companies | 1.57× | 31.4% | 20.0% |
| Skeptic Reassured | 1.51× | 100.0% | 66.4% |
| Guidance Worth Underwriting | 1.39× | 99.7% | 71.5% |
| Critic Ammunition | 0.00× | 0.0% | 87.8% |
| Scale-Dependent Advantage Claims | 0.03× | 0.3% | 11.1% |
| The Question Left Hanging | 0.17× | 8.4% | 48.0% |
| Results Worse Than Direction | 0.24× | 12.1% | 51.1% |
| The Hidden Segment | 0.41× | 8.7% | 21.1% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | +10.1% | -7.2% |
| Interquartile range | -8.7% to +28.3% | — |
| Share beating SPY | 62.8% (95% CI 58%–67%) | 39.5% |
| Observations | 468 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| DBOEY | Q2 2025 | 2025-07-25 | B+ |
| COUR | Q2 2025 | 2025-07-24 | A |
| PEGA | Q2 2025 | 2025-07-23 | A |
| TMSNY | Q2 2025 | 2025-07-23 | B |
| PITAF | Q2 2025 | 2025-07-22 | A |
| AGYS | Q1 2026 | 2025-07-21 | A |
| TTAN | Q1 2026 | 2025-06-06 | B |
| VEEV | Q1 2026 | 2025-05-28 | A |
4Discussion
A careful reader should conclude that these calls are, descriptively, a calmer and more confident population: less evasion, less stress, more specificity, and far more guidance raises than the base. That is an association in the data, not a causal mechanism, and it is not a trading signal. The returns gap reflects which companies end up in this bucket — often those already performing well — rather than a demonstrated edge. Our own forward tests falsified directional prediction, so the honest reading is a profile of how confident, backlog-backed calls look, not a recipe for beating the market.
5Limitations
All fields are AI-read and noisy; labels like "Critic Ammunition" or "Deferred Revenue Growing" reflect model judgment, not ground truth. The returns sample covers 468 of these calls against 22,449 base calls, skewed toward liquid names, so outcomes may not generalize. Critically, our own forward tests falsified directional prediction — the strong median-return gap should not be read as an exploitable signal. LLMs also partially remember famous stocks' histories, which can contaminate any backtest of this kind. Treat all figures as descriptive descriptions of this dataset, not forecasts or recommendations. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.