Research › 20-Year Trends
Artul.ai Research LibraryStudy No. 6620-Year TrendsUpdated 2026-08-28

Speak Softly and Carry the Spreadsheet: CFO-Dominated Earnings Calls

By Artul.ai Research Group · n = 23,336 earnings calls · First published 2026-08-28
Abstract

We studied 23,336 earnings calls out of 165,182 (14.13%, 95% CI 13.96%–14.30%) where the model answered YES to "When the CFO Dominates." CFO-dominated calls read less promotional (promotion 4.46 vs 5.05) and less confident (6.95 vs 7.21), but slightly more specific (7.71 vs 7.56) and under more stress (2.58 vs 2.43). Guidance is maintained more often (53.03% vs 48.79%) and raised slightly less often (17.69% vs 21.05%). Rehearsed-language flags are over-represented (lift 1.34), while hype phrases like "A Tiny Fraction of the Market" are under-represented (lift 0.64). Post-call returns skew modestly better than baseline.

Key findings
  • CFO-dominated calls are 14.13% of the 165,182-call corpus (95% CI 13.96%–14.30%).
  • They score lower on promotion (4.46 vs 5.05) and confidence (6.95 vs 7.21) but higher on specificity (7.71 vs 7.56).
  • Guidance is maintained on 53.03% of CFO-dominated calls versus 48.79% of baseline, and raised on 17.69% versus 21.05%.
  • The 'Calls That Read Rehearsed' flag is over-represented with a lift of 1.34, while 'A Tiny Fraction of the Market' is under-represented at 0.64.

1Introduction

Earnings-call watchers often treat the CFO as the adult in the room: the one who answers the question actually asked, hedges honestly, and skips the victory lap. Whether that stereotype survives contact with data is worth checking. If CFO-dominated calls really do sound different, their language profile, guidance behavior, and market reception should diverge measurably from the average call. Using Artul.ai's library of 165,182 calls spanning 1990 to 2026, this study examines the 23,336 calls where the model answered YES to "When the CFO Dominates," comparing their language, guidance moves, phrase-level flags, and post-call returns against the rest of the corpus.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "When the CFO Dominates", tracked by year (n = 23,336; 14.1% of the reference set, 95% Wilson interval 14.0%–14.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The language profile fits the stereotype: CFO-dominated calls are less promotional (4.46 vs 5.05) and less confident (6.95 vs 7.21), but slightly more specific (7.71 vs 7.56) and measured under somewhat more stress (2.58 vs 2.43). Guidance skews conservative: maintained guidance appears on 53.03% of these calls versus 48.79% overall, while raises are rarer (17.69% vs 21.05%). Phrase flags reinforce the tone: rehearsed-sounding language is over-represented (lift 1.34), as is "The Finished-Story Tell" (1.27), while hype-adjacent phrases like "Scale-Dependent Advantage Claims" (0.62) are under-represented. The annual share drifts from 16.42% in 2015 to 12.74% in 2025. Among 3,556 calls with return data, the median post-call return is -4.75% versus -7.16% for the 22,449-call baseline.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.946.86+0.08
Evasion2.682.70-0.01
Specificity7.717.56+0.15
Stress2.582.43+0.15
Promotion4.465.05-0.59
Confidence6.957.21-0.26
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised17.7%21.1%
Maintained53.0%48.8%
Lowered13.1%11.6%
Withdrawn2.8%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Calls That Read Rehearsed1.34×51.8%38.7%
The Finished-Story Tell1.27×5.6%4.4%
Scale-Dependent Advantage Claims0.62×6.8%11.1%
A Tiny Fraction of the Market0.64×19.3%30.0%
Volume About to Step Up0.65×18.5%28.5%
201516.42%
201615.37%
201714.54%
201815.01%
201915.28%
202013.86%
202111.72%
202213.70%
202314.49%
202413.28%
202512.74%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-4.8%-7.2%
Interquartile range-23.6% to +14.0%
Share beating SPY42.7% (95% CI 41%–44%)39.5%
Observations3,55622,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
USCBQ2 20252025-07-25B+
HCAQ2 20252025-07-25C
NWGQ2 20252025-07-25B+
BFHQ2 20252025-07-25B
FFICQ2 20252025-07-25B+
GBCIQ2 20252025-07-25A
UVEQ2 20252025-07-25C+
FLGQ2 20252025-07-25B

4Discussion

A careful reader should conclude that CFO-dominated calls have a distinct statistical fingerprint: flatter affect, fewer superlatives, more maintained guidance, and modestly less negative median post-call returns. That is a description of co-occurring characteristics, not a mechanism. Nothing here shows that CFO dominance causes better outcomes, better candor, or better returns; companies whose CFOs dominate may simply differ in sector, size, or circumstances. The lifts and deltas are descriptive associations in one corpus, and the annual shares fluctuate without a stable directional story. Treat the pattern as context for reading calls, not as a signal to act on.

5Limitations

The YES/NO labels come from AI-read fields and inherit their noise; a model's judgment of who 'dominates' a transcript is imperfect. The returns comparison rests on 22,449 calls skewed toward liquid names, and the 3,556-call sample shares that skew. Our own forward tests falsified directional prediction from these features, so no trading implication survives. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest-style comparison of labels and outcomes. The share CI is tight only because the corpus is large, not because the labels are precise. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Speak Softly and Carry the Spreadsheet: CFO-Dominated Earnings Calls.” Artul.ai Earnings-Call Research Library, Study No. 66. https://artul.ai/research/is-the-cfo-taking-over-the-earnings-call

Related studies

Give Them Nothing: The Calls Where Critics FeastYes Means Yes: Inside the Earnings Calls Where the SThe Question Was Answered. Just Not in This Room: HaRehearsal Season: The Earnings Calls That Sound a LiReady, Set, Step Up: The Calls Where Volume Was SuppLike Father, Like Firm: Founder-Led Earnings Calls,
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.