Give Them Nothing: The Calls Where Critics Feast
We examine 145,051 earnings calls from 1990 to 2026 where Artul.ai's model answered YES to "Critic Ammunition" — material a critic could seize on. These calls make up 87.81% of the 165,182-call corpus (95% CI: 87.65%–87.97%). Relative to other calls, they show slightly lower candor (6.83 vs 6.86), higher evasion (2.79 vs 2.70), and higher stress (2.58 vs 2.43). Post-call returns for a 19,045-call subset have a median of -0.08 and a 38.38% beat rate versus 39.47% for the 22,449-call baseline. Guidance is lowered in 12.96% of these calls versus 11.56% otherwise. This is descriptive, not predictive.
- 87.81% of the 165,182-call corpus (145,051 calls) was flagged as containing critic ammunition.
- Flagged calls show higher evasion (2.79 vs 2.70) and stress (2.58 vs 2.43) than the rest of the corpus.
- Median post-call return in the flagged returns sample is -0.08 vs -0.07 for the 22,449-call baseline.
- Guidance is lowered in 12.96% of flagged calls versus 11.56% of the remainder.
1Introduction
Every earnings call hands skeptics something: a hedged answer, a soft metric, an awkward pause. Whether that material is common — and what else it travels with — is a fair empirical question for anyone who reads transcripts closely. Using Artul.ai's "Critic Ammunition" field, we identify calls where the model answered YES and track them by year across a 1990–2026 corpus of 165,182 calls. The study describes how frequent the flag is, how tone scores and guidance behavior differ, and how returns distribute — without claiming any predictive value.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "Critic Ammunition", tracked by year (n = 145,051; 87.8% of the reference set, 95% Wilson interval 87.7%–88.0%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The flag is nearly universal: 87.81% of calls (n=145,051) qualify, with a 95% CI of 87.65% to 87.97%. Annual rates range from 84.29% (2021) to 90.98% (2015). Flagged calls skew slightly toward defensiveness: evasion 2.79 vs 2.70, stress 2.58 vs 2.43, candor 6.83 vs 6.86, confidence 7.13 vs 7.21. Guidance is lowered in 12.96% of flagged calls versus 11.56% otherwise, while raised-guidance rates sit at 18.59% vs 21.05%. In the 19,045-call returns subset, the median return is -0.08 vs -0.07 in the 22,449-call baseline, with beat rates of 38.38% vs 39.47%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.83 | 6.86 | -0.03 |
| Evasion | 2.79 | 2.70 | +0.10 |
| Specificity | 7.51 | 7.56 | -0.04 |
| Stress | 2.58 | 2.43 | +0.15 |
| Promotion | 5.08 | 5.05 | +0.03 |
| Confidence | 7.13 | 7.21 | -0.09 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 18.6% | 21.1% |
| Maintained | 48.5% | 48.8% |
| Lowered | 13.0% | 11.6% |
| Withdrawn | 2.9% | 2.7% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -8.2% | -7.2% |
| Interquartile range | -27.5% to +11.2% | — |
| Share beating SPY | 38.4% (95% CI 38%–39%) | 39.5% |
| Observations | 19,045 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
4Discussion
A careful reader should conclude only that critic-ammunition flags co-occur with modestly more evasive, more stressed language and slightly weaker guidance and return outcomes. These are small descriptive gaps, not evidence that the flag causes anything or forecasts returns. The flag's near-universality (87.81%) limits its discriminating power: it may simply describe the nature of earnings calls generally. No causal claim, timing signal, or trading edge is supported by this data.
5Limitations
AI-read fields like Critic Ammunition are noisy and unvalidated against ground truth. The returns sample covers only 22,449 calls and is skewed toward liquid names, so return comparisons may not generalize. Our own forward tests falsified directional prediction, so nothing here should be read as an edge. Additionally, LLMs partially remember famous stocks' histories, contaminating any backtest that mixes model outputs with known outcomes. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.