Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 34Hypotheses TestedUpdated 2026-08-28

Don't Stick to the Script: Earnings Calls Where Answers Outrun the Page

By Artul.ai Research Group · n = 200 earnings calls · First published 2026-08-28
Abstract

Of 1,445 earnings calls scored between 2015 and 2025, 200 (13.8%) answered yes to a simple hypothesis: the answers went deeper than the script. What do those calls look like? Their language profile leans candid and specific — confidence 7.57 versus 7.35, specificity 7.76 versus 7.65, evasion 2.42 versus 2.66, stress 2.09 versus 2.33 on other calls. Guidance skews constructive: 33.5% raised guidance against a 24.4% base rate, while 7.5% lowered it versus 12.0%. The phrase 'Volume About to Step Up' appears on 39.5% of flagged calls, 1.55 times its base rate, while 'The Question Left Hanging' (27.5% versus 41.7%) recedes. The payoff test is blunt: across 68 flagged calls with return data, the median post-call return was -5.03% versus -6.92%, and a 42.6% beat rate sits within noise of the 41.4% base rate.

Key findings
  • Flagged calls score higher on confidence (7.57 versus 7.35) and specificity (7.76 versus 7.65), and lower on evasion (2.42 versus 2.66) and stress (2.09 versus 2.33).
  • Guidance leans positive: 33.5% of flagged calls raised guidance versus a 24.4% base rate, and only 7.5% lowered it versus 12.0%.
  • 'Volume About to Step Up' appears on 39.5% of flagged calls (1.55x its 25.5% base rate), while 'The Question Left Hanging' (0.66x) and 'Calls That Read Rehearsed' (0.72x) recede.
  • Across 68 flagged calls with return data, the median return was -5.03% versus -6.92% overall, and the 42.6% beat rate (CI 31.6%-54.5%) spans the 41.4% base rate.

1Introduction

Anyone who follows earnings calls has absorbed the same advice for years: ignore the prepared remarks, watch the Q&A. The script is choreographed; the answers are where management either knows its business or doesn't. Yet the analyst's core judgment — that some calls genuinely go deeper than their script — is rarely quantified. How often does it happen? What do those calls sound like? Does the label attach to anything measurable downstream, like guidance or post-call returns? This study examines the 200 of 1,445 scored earnings calls (13.8%, 2015-2025) that answered yes to 'Answers go deeper than the script,' profiling their language, guidance behavior, characteristic phrases, yearly frequency, and returns — with an explicit check on whether any of it predicts anything.

2Data & methodology

The corpus comprises 1,445 earnings-call transcripts published between 2015 and 2025, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "Answers go deeper than the script" (n = 200; 13.8% of the reference set, 95% Wilson interval 12.2%–15.7%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The behavioral gaps all point the same way. Flagged calls score higher on candor (7.0 versus 6.92) and confidence (7.57 versus 7.35), and lower on evasion (2.42 versus 2.66) and stress (2.09 versus 2.33); promotion is flat (5.08 versus 5.09), so the shift reads as substance rather than salesmanship. Guidance leans constructive: 33.5% raised guidance versus a 24.4% base rate, and 7.5% lowered it versus 12.0%. Phrase lifts agree — 'Volume About to Step Up' appears on 39.5% of flagged calls against 25.5% (1.55x), while 'The Question Left Hanging' (27.5% versus 41.7%) and 'Calls That Read Rehearsed' (27.1% versus 37.8%) recede. The yearly trend is choppy, bottoming at 0% in 2020 and 2025 and peaking at 24% in 2023. Returns flatten the story: median -5.03% versus -6.92%, and a 42.6% beat rate whose interval spans the 41.4% baseline.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor7.006.92+0.08
Evasion2.422.66-0.25
Specificity7.767.65+0.12
Stress2.092.33-0.24
Promotion5.085.09-0.02
Confidence7.577.35+0.22
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised33.5%24.4%
Maintained46.5%51.4%
Lowered7.5%12.0%
Withdrawn0.5%1.0%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Volume About to Step Up1.55×39.5%25.5%
The Question Left Hanging0.66×27.5%41.7%
Calls That Read Rehearsed0.72×27.1%37.8%
20150.09%
20160.18%
20170.23%
20180.11%
20190.01%
20200.00%
20210.13%
20220.15%
20230.24%
20240.10%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-5.0%-6.9%
Interquartile range-20.2% to +12.7%
Share beating SPY42.6% (95% CI 32%–54%)41.4%
Observations68556
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
TXRHQ2 20242024-07-25A
FTAIQ2 20242024-07-24A
DAOQ1 20242024-05-23B
VITLQ1 20242024-05-09A
TALOQ1 20242024-05-07B+
FSPQ1 20242024-05-01D
WECQ1 20242024-05-01A
ROCKQ1 20242024-05-01B+

4Discussion

A careful reader can conclude that these calls sound different — more candid, more specific, less evasive — and that their guidance mix leans positive. A careful reader cannot conclude that the label caused any of it, that deeper answers predict returns, or that the flag identifies better investments. The 42.6% beat rate sits inside a confidence interval that comfortably includes the 41.4% base rate, and the median return gap (-5.03% versus -6.92%) is modest in a volatile sample. Treat 'answers deeper than the script' as a description of communication style, not a verdict on the business — and not a reason to trade on it.

5Limitations

These fields are AI-read, and AI-read fields are noisy: a 'candor' score is a model's opinion, not an instrument reading. The returns panel covers 22,449 calls and skews toward liquid names, so the comparison overweights widely followed stocks. Our own forward tests falsified directional prediction — flagged calls did no better out of sample. And large language models partially remember famous stocks' histories, which contaminates any backtest that pits language against outcomes. Treat every delta here as a descriptive pattern, not a working signal. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “Don't Stick to the Script: Earnings Calls Where Answers Outrun the Page.” Artul.ai Earnings-Call Research Library, Study No. 34. https://artul.ai/research/hypothesis-answers-go-deeper-than-the-script

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecHave Your Cake and Expand the Base Too: Repeat GrowtPressed Harder, Answered Straight: A Profile of 183 Second Front, Funded and Fueled: Who Says Yes on the
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.