Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 37Hypotheses TestedUpdated 2026-08-28

Slow and Steady Wins the Call: A Study of Compounding Evidence on Earnings Calls

By Artul.ai Research Group · n = 84 earnings calls · First published 2026-08-28
Abstract

This study examines earnings calls that answered YES to the research hypothesis "Compounding evidence" — management teams presenting steadily accumulating supporting facts. Of 470 calls in the 2015–2024 corpus, 84 calls (17.9%, 95% CI 14.7%–21.6%) met the definition. Compared with the broader corpus, these calls showed higher promotion (5.58 vs 5.11, +0.48), higher confidence (7.64 vs 7.33, +0.31), and lower stress (2.13 vs 2.39, -0.26). Language lifts were led by "Early Products Growing Fast" (1.68x) and "A Tiny Fraction of the Market" (1.59x). Guidance was raised on 26.2% of these calls versus 21.9% overall. Among 28 calls with follow-on returns data, the median forward return was -7.1% versus -10.0% for the base sample.

Key findings
  • 84 of 470 calls (17.9%) answered YES to the "Compounding evidence" hypothesis, with a 95% CI of 14.7% to 21.6%.
  • These calls scored higher on promotion (5.58 vs 5.11) and confidence (7.64 vs 7.33) while showing lower stress (2.13 vs 2.39).
  • The strongest language lift was "Early Products Growing Fast" at 1.68x the base rate, followed by "A Tiny Fraction of the Market" at 1.59x and "Founder-Led Companies" at 1.5x.
  • Median forward returns on the 28 calls with returns data were -7.1% versus -10.0% for the 184-call base sample, with 42.9% beating the base rate versus 39.7%.

1Introduction

Most earnings-call language decays fast: a claim is made, a quarter passes, and the narrative resets. The rarer pattern is the compounding call, where each statement appears to build on prior evidence — an early product gaining traction, a small share of a large market, volume ready to step up. For anyone who parses these transcripts, such calls are a distinct genre worth characterizing: they carry a recognizable tonal signature, a distinct guidance posture, and their own frequency over time. This study examines what the 84 compounding-evidence calls in a 470-call corpus spanning 2015–2024 look like across tone, guidance behavior, language themes, and subsequent returns.

2Data & methodology

The corpus comprises 470 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "Compounding evidence" (n = 84; 17.9% of the reference set, 95% Wilson interval 14.7%–21.6%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The tonal profile of compounding-evidence calls skews promotional: promotion runs +0.48 above the base (5.58 vs 5.11) and confidence +0.31 higher (7.64 vs 7.33), while stress sits 0.26 lower (2.13 vs 2.39); candor (-0.11) and evasion (-0.02) barely move. Guidance leans slightly positive — 26.2% raised versus 21.9% in the base — with withdrawals essentially unchanged (1.2% vs 1.1%). The theme lifts are the most distinctive feature: "Early Products Growing Fast" appears at 1.68x base frequency, "A Tiny Fraction of the Market" at 1.59x, "Founder-Led Companies" at 1.5x, and "Volume About to Step Up" at 1.36x. The annual trend peaked at 0.11 in 2022, fell to 0.09 in 2023 and 0.06 in 2024, with no qualifying calls in the partial 2025 data.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.806.91-0.11
Evasion2.702.72-0.02
Specificity7.737.64+0.09
Stress2.132.39-0.26
Promotion5.585.11+0.48
Confidence7.647.33+0.31
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised26.2%21.9%
Maintained51.2%52.6%
Lowered11.9%13.6%
Withdrawn1.2%1.1%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Early Products Growing Fast1.68×64.3%38.3%
A Tiny Fraction of the Market1.59×48.8%30.6%
Founder-Led Companies1.50×31.0%20.6%
Volume About to Step Up1.36×34.5%25.3%
20150.03%
20160.04%
20170.06%
20180.06%
20190.01%
20200.00%
20210.06%
20220.11%
20230.09%
20240.06%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-7.1%-10.0%
Interquartile range-26.9% to +14.1%
Share beating SPY42.9% (95% CI 27%–61%)39.7%
Observations28184
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
NOAHQ1 20242024-05-30D
CRGOQ1 20242024-05-20C+
FWONKQ1 20242024-05-08C+
LINCQ1 20242024-05-06B+
AESQ1 20242024-05-03C+
PPCQ1 20242024-05-03A
ROCKQ1 20242024-05-01B+
LTRXQ3 20242024-04-29C

4Discussion

A careful reader should treat this as a description of a recognizable call archetype, not a signal. Compounding-evidence calls do talk differently — more promotion, more confidence, less stress — and they lean on growth-narrative phrases at elevated rates. Their forward returns in this sample (median -7.1% vs -10.0% base) are less bad but still negative, and the returns subsample is small at 28 calls with overlapping confidence intervals on the beat rate (42.9%, CI 26.5%–60.9%). Nothing here establishes that the language caused any outcome, or that identifying this pattern in a new transcript would help an investor. The pattern is real as a description; its usefulness is unproven.

5Limitations

The tone, theme, and hypothesis labels are produced by AI reading call transcripts and are inherently noisy; misclassification at the call level is expected. The returns analysis covers only 28 of these calls within a 22,449-call base skewed toward liquid names, so the comparison is not representative of all listed companies. Our own forward tests falsified directional prediction from these features, and no claim of predictive power survives that result. Additionally, LLMs partially remember famous stocks' history, which contaminates any backtest of language against subsequent returns; measured lift may reflect memorized narratives rather than genuine inference from the text. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “Slow and Steady Wins the Call: A Study of Compounding Evidence on Earnings Calls.” Artul.ai Earnings-Call Research Library, Study No. 37. https://artul.ai/research/hypothesis-compounding-evidence

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecDon't Stick to the Script: Earnings Calls Where AnswHave Your Cake and Expand the Base Too: Repeat GrowtPressed Harder, Answered Straight: A Profile of 183
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.