Ammunition, Not Smoke: Inside the Calls Where Critics Load Up
This study examines earnings calls where the model flagged the battery item "Critic Ammunition" with a YES, covering 145,051 of 165,182 calls (87.8%, 95% CI 87.7%–88.0%) across 1990–2026. Flagged calls skew slightly toward defensive language: stress averages 2.58 versus 2.43, and confidence 7.13 versus 7.21, while candor (6.83 vs 6.86) and specificity (7.51 vs 7.56) run marginally lower. Guidance mix is modestly weaker (18.6% raised, 13.0% lowered). The returns sample (22,449 base calls) shows a median post-call return of -0.0818 versus -0.0716 baseline, with 38.4% beating versus 39.5% baseline.
- 87.8% of 165,182 calls (145,051) were flagged YES for "Critic Ammunition", with a 95% CI of 87.7% to 88.0%.
- Flagged calls show elevated stress (2.58 vs 2.43) and reduced confidence (7.13 vs 7.21) relative to the base.
- Guidance was lowered on 13.0% of flagged calls versus 11.6% of baseline calls, and raised on 18.6% versus 21.1%.
- The 19,045 flagged calls with returns data had a median post-call return of -0.0818 versus -0.0716 for the 22,449-call baseline, with 38.4% beating versus 39.5%.
1Introduction
Earnings calls are performances under scrutiny, and analysts' questions often arrive pre-loaded with ammunition. This study asks how often the model detects that critical framing, and what such calls look like. The answer is: nearly always. 87.8% of the 165,182 calls in the corpus, spanning 1990 to 2026, were flagged YES for "Critic Ammunition", making this one of the most pervasive patterns in the library. Because flagged calls are the norm rather than the exception, the interesting question is how the small behavioral differences between flagged and unflagged calls line up with tone, guidance choices, and subsequent returns. This study examines those differences.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Critic Ammunition" (n = 145,051; 87.8% of the reference set, 95% Wilson interval 87.7%–88.0%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral gaps are small but consistent: flagged calls run higher on stress (2.58 vs 2.43) and evasion (2.79 vs 2.70), and lower on confidence (7.13 vs 7.21), candor (6.83 vs 6.86), and specificity (7.51 vs 7.56). Guidance mix tilts weaker, with 13.0% lowered versus 11.6% baseline and 18.6% raised versus 21.1%. Trend data shows the flag rate ranging from 84.29% (2021) to 90.98% (2015), hovering near 88–90% in recent years. In the returns sample, flagged calls show a median of -0.0818 versus -0.0716 baseline, and a 38.4% beat rate versus 39.5%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.83 | 6.86 | -0.03 |
| Evasion | 2.79 | 2.70 | +0.10 |
| Specificity | 7.51 | 7.56 | -0.04 |
| Stress | 2.58 | 2.43 | +0.15 |
| Promotion | 5.08 | 5.05 | +0.03 |
| Confidence | 7.13 | 7.21 | -0.09 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 18.6% | 21.1% |
| Maintained | 48.5% | 48.8% |
| Lowered | 13.0% | 11.6% |
| Withdrawn | 2.9% | 2.7% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -8.2% | -7.2% |
| Interquartile range | -27.5% to +11.2% | — |
| Share beating SPY | 38.4% (95% CI 38%–39%) | 39.5% |
| Observations | 19,045 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
4Discussion
A careful reader should conclude only that calls carrying the "Critic Ammunition" flag differ slightly in tone, guidance mix, and median returns from unflagged calls. These are descriptive associations computed from the same corpus; the flag identifies how questions are framed, not why management responds as it does, and none of these deltas establish any causal or predictive relationship. The flag's sheer prevalence (87.8%) also means flagged calls closely resemble the overall corpus. The differences in returns and beat rates are small and within ranges easily explained by sample composition.
5Limitations
The behavioral scores are AI-read fields and are inherently noisy, so small deltas should be treated cautiously. The returns sample covers 22,449 calls and is skewed toward liquid names, limiting generalizability. Our own forward tests falsified directional prediction, so nothing here should be read as a trading signal. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of this kind. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.