Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 58Hypotheses TestedUpdated 2026-08-28

Show, Don't Tell: Calls With Strong Facts and a Held-Back Story

By Artul.ai Research Group · n = 101 earnings calls · First published 2026-08-28
Abstract

This study examines 101 earnings calls (out of 460, a 0.22 share) that fit the pattern "Strong facts, held-back story" — calls where management presents solid specifics while underplaying narrative promotion. These calls show higher confidence (7.67 vs 7.33) and specificity (7.79 vs 7.65), with lower stress (2.2 vs 2.4) and slightly lower promotion (5.05 vs 5.12). Guidance behavior differs sharply: 0.40 raised guidance versus 0.22 in the base, and only 0.01 lowered versus 0.13. Overrepresented themes include "Pricing Recovering" (1.58x) and "Skeptic Reassured" (1.29x); underrepresented are "Results Worse Than Direction" (0.45x) and "The Hidden Segment" (0.7x). Post-call returns for 51 calls show a median of -0.07 versus -0.08 for 185 base calls, with 0.43 beats.

Key findings
  • Calls matching the pattern make up 0.22 of the corpus (101 of 460), with a confidence delta of +0.34 and a stress delta of -0.2 versus the base.
  • Guidance was raised on 0.40 of these calls versus 0.22 in the base, and lowered on only 0.01 versus 0.13.
  • The theme "Pricing Recovering" appears 1.58x more often than in the base corpus, while "Results Worse Than Direction" appears at 0.45x.
  • Among 51 calls with returns data, the median post-call return was -0.07 versus -0.08 for the 185-call base, with 0.43 beating versus 0.39.

1Introduction

Some earnings calls read like a drawer of receipts: plenty of verifiable specifics, little salesmanship. Calls that answer YES to "Strong facts, held-back story" let the numbers carry the argument — confidence runs higher, stress lower, and promotion actually below average. For anyone who parses earnings calls for a living, that combination is interesting precisely because it inverts the usual assumption that enthusiasm signals strength. If management is quietly specific while talking down the narrative, what does the rest of the call look like? This study profiles those 101 calls across 2015-2024: their behavioral markers, guidance actions, recurring themes, and post-call outcomes.

2Data & methodology

The corpus comprises 460 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "Strong facts, held-back story" (n = 101; 22.0% of the reference set, 95% Wilson interval 18.4%–26.0%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The behavioral profile is the headline: specificity of 7.79 versus 7.65 and confidence of 7.67 versus 7.33, paired with promotion of 5.05 versus 5.12 — more assurance, less selling. Guidance skews positive: 0.40 raised versus 0.22 in the base, 0.53 maintained, and essentially none lowered (0.01 vs 0.13). Theme lifts point the same direction — "Skeptic Reassured" at 1.29x and "Pricing Recovering" at 1.58x — while "The Hidden Segment" (0.7x) and "Results Worse Than Direction" (0.45x) fade. The pattern's prevalence has varied over time, from 0.0 in 2019-2020 to 0.11 in 2016 and 2022. Post-call returns are modestly better than base: median -0.07 versus -0.08, with 0.43 beats.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.966.92+0.05
Evasion2.832.72+0.11
Specificity7.797.65+0.14
Stress2.202.40-0.20
Promotion5.055.12-0.07
Confidence7.677.33+0.34
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised39.6%22.0%
Maintained53.5%53.9%
Lowered1.0%12.8%
Withdrawn0.0%1.1%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Pricing Recovering1.58×25.7%16.3%
Skeptic Reassured1.29×91.1%70.9%
Results Worse Than Direction0.45×20.8%46.3%
The Hidden Segment0.70×16.8%24.1%
20150.09%
20160.11%
20170.09%
20180.06%
20190.00%
20200.00%
20210.08%
20220.11%
20230.10%
20240.02%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-6.7%-7.6%
Interquartile range-26.7% to +12.5%
Share beating SPY43.1% (95% CI 31%–57%)38.9%
Observations51185
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
CRGOQ1 20242024-05-20C+
AZEKQ2 20242024-05-08B+
GLQ1 20242024-04-23F
PBRQ4 20232024-03-08D
NICEQ4 20232024-02-22B+
EGPQ4 20232024-02-08B
DXCMQ4 20232024-02-08B+
METQ4 20232024-02-01B+

4Discussion

A careful reader should treat this as a description of a call style, not a signal. These calls coincide with more guidance raises and reassured-skeptic themes, and their post-call returns were slightly better than base — but coincidence of traits is not a cause, and the returns gap (median -0.07 vs -0.08) is small enough to sit inside ordinary noise. Nothing here predicts how a future call or stock will behave. The fair conclusion is narrower: this style of call, in this dataset, tended to accompany constructive guidance actions and calmer delivery.

5Limitations

The behavioral scores and themes are AI-read fields and inherently noisy, so deltas like +0.34 confidence should be read as tendencies, not measurements. The returns sample covers only 51 of these calls (185 base) and is skewed toward liquid names. Our own forward tests of this pipeline falsified directional prediction, so no edge should be inferred. Finally, LLMs partially remember famous stocks' histories, which can contaminate any backtest-style comparison, including the returns figures reported here. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “Show, Don't Tell: Calls With Strong Facts and a Held-Back Story.” Artul.ai Earnings-Call Research Library, Study No. 58. https://artul.ai/research/hypothesis-strong-facts-held-back-story

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecDon't Stick to the Script: Earnings Calls Where AnswHave Your Cake and Expand the Base Too: Repeat GrowtPressed Harder, Answered Straight: A Profile of 183
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.