Speak Softly and Carry the Spreadsheet: CFO-Dominated Calls
We examined 23,336 earnings calls (14.1% of a 165,182-call corpus spanning 1990-2026) where the model answered YES to the battery item "When the CFO Dominates." These calls skew toward specifics: specificity runs 7.71 vs. 7.56 in the base, while promotion language drops to 4.46 vs. 5.05 and confidence to 6.95 vs. 7.21. Guidance is maintained on 53.0% of calls vs. 48.8% base. Language flags also differ: "The Finished-Story Tell" appears at 1.27x base, while "Scale-Dependent Advantage Claims" appear at 0.62x. In the 3,556-call subset with measured returns, the median move is -4.75% vs. -7.16% base, with 42.7% beating vs. 39.5% base.
- CFO-dominated calls show specificity of 7.71 vs. 7.56 base and promotion language of 4.46 vs. 5.05 base.
- Guidance is maintained on 53.0% of these calls vs. 48.8% in the base corpus.
- The rehearsed-language flag "The Finished-Story Tell" appears at 1.27x base, while "Scale-Dependent Advantage Claims" appear at only 0.62x.
- In the 3,556-call returns subset, the median post-call move is -4.75% vs. -7.16% base, and 42.7% beat vs. 39.5% base.
1Introduction
Anyone who listens to earnings calls knows the two voices: the CEO narrating the future and the CFO defending the numbers. When the CFO dominates the call, the texture of the conversation changes - less selling, more reconciling, more line items and fewer adjectives. Whether that shift is measurable, and how it coexists with guidance decisions and market reactions, is worth quantifying rather than assuming. This study examines 23,336 calls from a 165,182-call corpus spanning 1990-2026 where the model answered YES to "When the CFO Dominates," comparing their language profile, guidance mix, flag rates, and post-call return distribution against the rest of the corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "When the CFO Dominates" (n = 23,336; 14.1% of the reference set, 95% Wilson interval 14.0%–14.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile fits the stereotype: specificity is higher (7.71 vs. 7.56) while promotion language (4.46 vs. 5.05) and confidence (6.95 vs. 7.21) run lower. Guidance is maintained more often (53.0% vs. 48.8%) and lowered slightly more (13.1% vs. 11.6%). Flag rates tell a consistent story: "The Finished-Story Tell" and "Calls That Read Rehearsed" are overrepresented at 1.27x and 1.34x, while hype-adjacent items are underrepresented - "Scale-Dependent Advantage Claims" at 0.62x and "A Tiny Fraction of the Market" at 0.64x. The trend shows the share falling from 16.4% in 2015 to 12.7% in 2025. Returns skew milder: median -4.75% vs. -7.16%, with 42.7% beating vs. 39.5%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.94 | 6.86 | +0.08 |
| Evasion | 2.68 | 2.70 | -0.01 |
| Specificity | 7.71 | 7.56 | +0.15 |
| Stress | 2.58 | 2.43 | +0.15 |
| Promotion | 4.46 | 5.05 | -0.59 |
| Confidence | 6.95 | 7.21 | -0.26 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 17.7% | 21.1% |
| Maintained | 53.0% | 48.8% |
| Lowered | 13.1% | 11.6% |
| Withdrawn | 2.8% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Calls That Read Rehearsed | 1.34× | 51.8% | 38.7% |
| The Finished-Story Tell | 1.27× | 5.6% | 4.4% |
| Scale-Dependent Advantage Claims | 0.62× | 6.8% | 11.1% |
| A Tiny Fraction of the Market | 0.64× | 19.3% | 30.0% |
| Volume About to Step Up | 0.65× | 18.5% | 28.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -4.8% | -7.2% |
| Interquartile range | -23.6% to +14.0% | — |
| Share beating SPY | 42.7% (95% CI 41%–44%) | 39.5% |
| Observations | 3,556 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| GBCI | Q2 2025 | 2025-07-25 | A |
| UVE | Q2 2025 | 2025-07-25 | C+ |
| FLG | Q2 2025 | 2025-07-25 | B |
4Discussion
A careful reader should conclude that CFO-dominated calls carry a measurably different texture: more specificity, less promotional language, more maintained guidance, and a somewhat milder return distribution. None of this establishes why. The association could reflect company circumstances, sector mix, or which quarters lead a CFO to take the microphone. The differences are modest in size, and the flag-rate lifts involve small base rates. These are descriptive patterns in one model's annotations, not rules about any individual call, company, or quarter.
5Limitations
The battery fields are AI-read and noisy; a YES label reflects the model's judgment, not ground truth about who spoke. The returns sample covers 22,449 calls and skews toward liquid names, so it is not representative of the full corpus. Our own forward tests falsified directional prediction - these statistics describe the past, not the future. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest. Treat every figure here as descriptive. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.