Big Claims, Thin Results: Scale-Dependent Advantage Talk on Earnings Calls
This study examines 18,297 earnings calls (11.1% of a 165,182-call corpus spanning 1990 to 2026) where a model flagged claims that a company's advantages depend on operating at scale. These calls differ sharply in tone: promotion runs 6.07 versus a 5.05 baseline, while specificity sits at 6.83 against 7.56 and candor at 6.21 against 6.86. Guidance behavior skews negative: 8.8% raised guidance versus 21.1% in the base group, and 13.0% lowered it versus 11.6%. Post-call price reactions were modestly worse, with a median return of -0.32% versus -0.07% baseline and only 26.0% beating expectations against 39.5% baseline. The theme's prevalence peaked at 13.5% of calls in 2022 before falling to 9.8% in 2025.
- Calls with scale-dependent advantage claims account for 11.1% of the corpus (18,297 of 165,182 calls).
- Promotion tone is elevated on these calls (6.07 versus 5.05 baseline) while specificity (6.83 versus 7.56) and candor (6.21 versus 6.86) are depressed.
- Only 26.0% of these calls beat expectations versus 39.5% of the 22,449-call base sample, with a median post-call return of -0.32% versus -0.07%.
- The share of such calls rose from 9.1% in 2020 to 13.5% in 2022, then declined to 9.8% in 2025.
1Introduction
Executives often argue that scale itself is the moat: more volume lowers unit costs, entrenches the leader, and starves rivals. That framing is seductive because it converts a promise into a physics lesson. But talk of scale-dependent advantage is also a convenient place to hide vagueness, and a careful reader of earnings calls should know what tends to accompany it. Using Artul.ai's library of 165,182 transcripts from 1990 through 2026, this study profiles the 18,297 calls where the model answered YES to the battery item 'Scale-Dependent Advantage Claims,' comparing their tone, guidance behavior, thematic profile, and post-call price reactions against the rest of the corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Scale-Dependent Advantage Claims" (n = 18,297; 11.1% of the reference set, 95% Wilson interval 10.9%–11.2%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The tonal profile is the most striking pattern: promotion is up 1.02 points (6.07 versus 5.05) while specificity is down 0.73 (6.83 versus 7.56) and candor down 0.65 (6.21 versus 6.86), and stress runs 0.93 points higher. Guidance activity skews cautious: 8.8% raised versus 21.1% baseline, while 13.0% lowered versus 11.6%. Thematically, these calls over-index on 'A Tiny Fraction of the Market' (2.48x), 'The Question Left Hanging' (2.02x), and 'Underused Fixed Costs' (1.97x), and strongly under-index on 'Skeptic Reassured' (0.09x) and 'Guidance Worth Underwriting' (0.2x). Post-call returns are weaker: a median of -0.32% versus -0.07%, with 26.0% beating versus 39.5%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.21 | 6.86 | -0.65 |
| Evasion | 3.26 | 2.70 | +0.57 |
| Specificity | 6.83 | 7.56 | -0.73 |
| Stress | 3.35 | 2.43 | +0.93 |
| Promotion | 6.07 | 5.05 | +1.02 |
| Confidence | 6.89 | 7.21 | -0.32 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 8.8% | 21.1% |
| Maintained | 35.5% | 48.8% |
| Lowered | 13.0% | 11.6% |
| Withdrawn | 3.2% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| A Tiny Fraction of the Market | 2.48× | 74.5% | 30.0% |
| The Question Left Hanging | 2.02× | 96.7% | 48.0% |
| Underused Fixed Costs | 1.97× | 82.0% | 41.6% |
| Founder-Led Companies | 1.76× | 35.2% | 20.0% |
| Results Worse Than Direction | 1.70× | 87.2% | 51.1% |
| Skeptic Reassured | 0.09× | 6.0% | 66.4% |
| Guidance Worth Underwriting | 0.20× | 14.2% | 71.5% |
| Calls That Resolve Doubts | 0.33× | 26.2% | 79.5% |
| Confidence Proportionate to Evidence | 0.38× | 33.5% | 87.6% |
| Pricing Recovering | 0.59× | 12.7% | 21.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -31.9% | -7.2% |
| Interquartile range | -58.5% to +1.3% | — |
| Share beating SPY | 26.0% (95% CI 23%–30%) | 39.5% |
| Observations | 565 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| VRTS | Q2 2025 | 2025-07-25 | C+ |
| ASPS | Q2 2025 | 2025-07-25 | D |
| INTC | Q2 2025 | 2025-07-24 | D |
| DAIO | Q2 2025 | 2025-07-24 | D |
| FPH | Q2 2025 | 2025-07-24 | D |
| RPT | Q2 2025 | 2025-07-24 | D |
| MBLY | Q2 2025 | 2025-07-24 | B |
| IRDM | Q2 2025 | 2025-07-24 | F |
4Discussion
A careful reader should conclude that calls emphasizing scale-dependent advantage are, in this dataset, associated with more promotional, less specific language, more guarded guidance, and weaker subsequent price reactions. These are co-occurrences, not causes: the scale talk does not make outcomes worse, and worse outcomes do not necessarily produce scale talk. The pattern is consistent with companies leaning on structural narratives during harder periods, but the data here cannot confirm intent. No conclusion in this study should be read as a signal for predicting any individual call or trade.
5Limitations
The battery items are AI-read and inherently noisy; the same transcript could plausibly be labeled differently on another pass. The returns sample covers 22,449 calls skewed toward liquid names, so the 565-call flagged subset may not generalize. Our own forward tests falsified directional prediction, and LLMs partially remember famous stocks' histories, contaminating any backtest. The trend figures mix regimes and sample sizes, including a partial 2025. Treat every number here as descriptive of this corpus only. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.