Read From the Same Script, Say Almost Nothing: Rehearsed Calls With Low Specificity
This study profiles the 378 earnings calls (0.23% of a 165,182-call corpus spanning 1990-2026) where a language model flagged the call as reading rehearsed AND scored specificity at 2/9 or lower. Compared with corpus baselines, these calls score 1.42 versus 7.56 on specificity, 3.61 versus 6.86 on candor, and 4.65 versus 7.21 on confidence. Themes that typically signal substance are nearly absent: 'Calls That Resolve Doubts', 'Guidance Worth Underwriting', and 'Skeptic Reassured' each appear in 0% of these calls versus roughly 60-80% of the broader corpus. Their prevalence was near zero through 2023 but reached 4.41% of 6,012 calls in 2025.
- The 378 flagged calls represent 0.23% of the 165,182-call corpus, with a 95% confidence interval of 0.21%-0.25%.
- Specificity averages 1.42 versus a corpus baseline of 7.56, the largest behavioral gap of any measured trait (delta -6.14).
- 'Skeptic Reassured' appears in 0% of these calls versus 66% of the corpus, and 'Guidance Worth Underwriting' in 0% versus 72%.
- The share was 0% or 1% every year from 2015 through 2023, then jumped to 4.41% of 6,012 calls in 2025.
1Introduction
Analysts who read hundreds of transcripts a season develop a feel for calls that sound polished but convey little. This study attempts to quantify that feeling: calls a model judged rehearsed AND nearly devoid of specifics. These calls combine two traits usually studied separately - scripted delivery and low informational content - and their near-absence of reassuring moments makes them a useful contrast group for what normal, substantive calls contain. The pattern is rare but concentrated in recent years. We examine how these 378 calls differ across candor, evasion, guidance behavior, and recurring transcript themes.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "Calls That Read Rehearsed" AND specificity scored 2/9 or lower (n = 378; 0.2% of the reference set, 95% Wilson interval 0.2%–0.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile is uniformly depressed: specificity 1.42 versus 7.56, candor 3.61 versus 6.86, confidence 4.65 versus 7.21, and evasion 0.30 versus 2.70. Guidance actions are rare - 0% raised guidance versus 21% corpus-wide. The theme table is starker: 'The Question Left Hanging' appears 2.09 times per call against a 1.00 baseline, while 'Calls That Resolve Doubts', 'Guidance Worth Underwriting', and 'Skeptic Reassured' each appear 0% of the time against baselines of 66%-80%. The trend shows near-total absence through 2023 followed by 4.41% in 2025.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 3.61 | 6.86 | -3.25 |
| Evasion | 0.30 | 2.70 | -2.39 |
| Specificity | 1.42 | 7.56 | -6.14 |
| Stress | 1.38 | 2.43 | -1.05 |
| Promotion | 2.51 | 5.05 | -2.54 |
| Confidence | 4.65 | 7.21 | -2.56 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 0.0% | 21.1% |
| Maintained | 0.8% | 48.8% |
| Lowered | 0.3% | 11.6% |
| Withdrawn | 0.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| The Question Left Hanging | 2.09× | 100.0% | 48.0% |
| Scale-Dependent Advantage Claims | 1.72× | 19.0% | 11.1% |
| Calls That Resolve Doubts | 0.00× | 0.0% | 79.5% |
| Guidance Worth Underwriting | 0.00× | 0.0% | 71.5% |
| Skeptic Reassured | 0.00× | 0.0% | 66.4% |
| Pricing Recovering | 0.02× | 0.5% | 21.5% |
| Deferred Revenue Growing | 0.03× | 0.3% | 8.9% |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| CRDOF | Q1 2025 | 2025-05-29 | F |
| REX | Q1 2025 | 2025-05-28 | F |
| TPTA | Q1 2025 | 2025-05-22 | F |
| WB | Q1 2025 | 2025-05-21 | F |
| WRD | Q1 2025 | 2025-05-21 | F |
| MDWD | Q1 2025 | 2025-05-21 | D |
| BIOX | Q3 2025 | 2025-05-21 | D |
| XPEV | Q1 2025 | 2025-05-21 | F |
4Discussion
A careful reader should conclude that these calls are describable: low on specifics, light on guidance, and heavy on unresolved questions. What one should not conclude is that the label explains anything about outcomes. The 2025 spike could reflect model behavior, transcript availability, or genuine shifts in call style - the data cannot separate these. Nothing here establishes that rehearsed-sounding calls precede good or bad results, and the rarity of the pattern means small sample noise shapes many of the year-level figures.
5Limitations
AI-read fields such as rehearsed-ness and specificity are noisy model judgments, not ground truth, and the rarity of the pattern amplifies labeling errors. Any returns comparison would rest on 22,449 calls skewed toward liquid names, limiting generalizability. Our own forward tests falsified directional prediction from these signals, so no trading implication is offered. Finally, LLMs partially remember famous stocks' histories, contaminating any backtest of transcript-based labels. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.