Research › Management Behavior
Artul.ai Research LibraryStudy No. 68Management BehaviorUpdated 2026-08-28

Bad News Comes in Sixes: What High-Uncertainty Calls Sound Like

By Artul.ai Research Group · n = 18,704 earnings calls · First published 2026-08-28
Abstract

We examine 18,704 earnings calls scoring 6 or higher on a 0-9 uncertainty meter, drawn from a corpus of 165,182 calls spanning 1990 to 2026. These high-uncertainty calls make up 11.32% of the corpus (95% CI 11.17% to 11.48%). Management teams on them show sharply different behavior: confidence averages 5.96 versus 7.21 on all other calls, promotion framing drops to 4.15 versus 5.05, and stress rises to 3.66 versus 2.43. Guidance behavior diverges strongly: 32.59% lower guidance versus 11.56% in the comparison group, and 17.24% withdrawn versus 2.66%. Among 1,656 calls with measured next-day returns, the median return is -0.0967% versus -0.0716% for the 22,449-call base sample, with 36.96% beating versus 39.47%.

Key findings
  • High-uncertainty calls (score 6+) account for 11.32% of the 165,182-call corpus, with a 95% confidence interval of 11.17% to 11.48%.
  • Confidence averages 5.96 on high-uncertainty calls versus 7.21 elsewhere, and stress averages 3.66 versus 2.43.
  • Guidance is lowered on 32.59% of high-uncertainty calls and withdrawn on 17.24%, compared with 11.56% and 2.66% respectively in the base population.
  • The trend peaks at 30.59% of calls in 2020, with 2025 at 14.47% on 6,012 calls.

1Introduction

Earnings calls are a performance, and uncertainty shows through the script. When management teams are unsure of their own numbers, the tell is not what they say but how they say it: less promotion, less confidence, more stress, more evasion. For anyone who reads calls closely, distinguishing routine hedging from genuine uncertainty could sharpen how they interpret a quarter. Using a 0-9 uncertainty meter, this study isolates 18,704 calls out of 165,182 that scored 6 or higher, spanning 1990 through 2026. It examines how these calls differ in language profile, guidance behavior, recurring narrative patterns, and measured next-day returns.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 uncertainty meter (n = 18,704; 11.3% of the reference set, 95% Wilson interval 11.2%–11.5%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The language profile is distinctive: confidence runs 5.96 versus 7.21, promotion 4.15 versus 5.05, and stress 3.66 versus 2.43, while evasion rises to 3.31 from 2.70. Guidance tells the same story: lowered on 32.59% of high-uncertainty calls versus 11.56% otherwise, and withdrawn on 17.24% versus 2.66%; raised guidance falls to 4.57% from 21.05%. Overrepresented narrative patterns include 'The Question Left Hanging' at 1.6x lift and 'Results Worse Than Direction' at 1.45x, while 'Guidance Worth Underwriting' appears at 0.49x. The trend spikes to 30.59% in 2020 and sits at 14.47% for 2025 on 6,012 calls. Median measured returns are -0.0967% versus -0.0716% in the base sample.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor7.106.86+0.24
Evasion3.312.70+0.61
Specificity7.317.56-0.25
Stress3.662.43+1.23
Promotion4.155.05-0.90
Confidence5.967.21-1.26
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised4.6%21.1%
Maintained26.8%48.8%
Lowered32.6%11.6%
Withdrawn17.2%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
The Question Left Hanging1.60×76.9%48.0%
Results Worse Than Direction1.45×74.4%51.1%
Underused Fixed Costs1.45×60.5%41.6%
Scale-Dependent Advantage Claims1.37×15.1%11.1%
When the CFO Dominates1.32×18.6%14.1%
Guidance Worth Underwriting0.49×34.9%71.5%
Skeptic Reassured0.51×34.0%66.4%
Volume About to Step Up0.54×15.3%28.5%
Deferred Revenue Growing0.61×5.4%8.9%
Calls That Resolve Doubts0.64×50.5%79.5%
201512.66%
201610.06%
20176.92%
20186.09%
20199.93%
202030.59%
20218.18%
202211.79%
20239.09%
20247.51%
202514.47%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-9.7%-7.2%
Interquartile range-32.1% to +11.4%
Share beating SPY37.0% (95% CI 35%–39%)39.5%
Observations1,65622,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
CNCQ2 20252025-07-25F
MTHQ2 20252025-07-25C
WFQ2 20252025-07-25C
WZZAFQ1 20262025-07-25F
MLLGFQ2 20252025-07-25C+
RNECFQ2 20252025-07-25D
INTCQ2 20252025-07-24D
SAMQ2 20252025-07-24D

4Discussion

High-uncertainty calls look, sound, and behave differently from the rest of the corpus: quieter promotion, more stress, more withdrawn and lowered guidance, and narrative patterns about unanswered questions. These are associations in a large descriptive dataset, not a rule for reading any single call. The return comparison is descriptive too, and a small median gap does not establish any trading edge. Careful readers should treat these figures as a map of what high-uncertainty calls tend to contain, not as a signal that predicts outcomes.

5Limitations

The uncertainty score and all language fields are AI-read and inherently noisy, so individual call labels should be treated as approximate. The returns sample covers 22,449 calls and is skewed toward liquid names, so measured returns may not generalize. Our own forward tests falsified directional prediction, and LLMs partially remember famous stocks' histories, which can contaminate any backtest of these labels. The study is descriptive throughout, and no causal or predictive claim is supported. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Bad News Comes in Sixes: What High-Uncertainty Calls Sound Like.” Artul.ai Earnings-Call Research Library, Study No. 68. https://artul.ai/research/management-admitting-uncertainty

Related studies

Specificity Is Next to Godliness: 161,829 Earnings CThe Truth Won't Set You Free: High Candor Is the NorNine Out of Ten Calls Are 'Confident' Now: Inside thStraight Talk, Stress Less: Low-Evasion Earnings CalThe Long Way to Say Less: Inside High-Complexity EarPromoted to the Front Page: High-Score Calls Across
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.