Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 47Hypotheses TestedUpdated 2026-08-28

The Storm Chasers: Earnings Calls That Lean Into Turbulence

By Artul.ai Research Group · n = 86 earnings calls · First published 2026-08-28
Abstract

This study examines earnings calls that answered YES to the research hypothesis 'Leaning into the storm' — management teams that confront rather than sidestep turbulent conditions. Of 382 calls in the 2015–2024 corpus, 86 (22.5%, 95% CI 18.6%–27.0%) fit the pattern. Compared to the corpus baseline, these calls show higher stress (2.78 vs 2.24) and lower confidence (6.85 vs 7.43), while candor (7.0 vs 6.91) and specificity (7.55 vs 7.68) barely move. They lower guidance far more often than the base rate (24.4% vs 10.2%). The 'Results Worse Than Direction' theme appears 1.62x more often than in typical calls. In the 34-call returns sample, the median one-year return was -18.98% versus -7.18% for the 149-call comparison set.

Key findings
  • 86 of 382 calls (22.5%, CI 18.6%-27.0%) answered YES to 'Leaning into the storm'.
  • Stress scores run higher (2.78 vs 2.24) and confidence lower (6.85 vs 7.43) than baseline, while candor (7.0) and specificity (7.55) sit near corpus norms.
  • These teams lower guidance at 24.4% versus 10.2% for the broader set, and the 'Results Worse Than Direction' theme shows up 1.62x more often.
  • The 34-call returns sample shows a median one-year return of -18.98% versus -7.18% for the comparison set, with a 32.4% beat rate against 40.3%.

1Introduction

When markets turn rough, some management teams spend the call managing the narrative downward with a shrug, while others walk straight into the bad news and describe it in detail. That choice is one of the most telling signals an earnings-call listener can watch for, because it reveals how a leadership team metabolizes adversity — and whether the storm they describe is operational reality or a pre-positioning exercise. Calls where the answer is 'Leaning into the storm' cluster around visible turbulence: lowered guidance, worse-than-advertised results, and stress that reads through the transcript. This study profiles those 86 calls out of 382, examining their language profile, guidance behavior, recurring themes, frequency over time, and the one-year returns that followed.

2Data & methodology

The corpus comprises 382 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "Leaning into the storm" (n = 86; 22.5% of the reference set, 95% Wilson interval 18.6%–27.0%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The behavioral profile is distinctive: stress is elevated at 2.78 versus 2.24 baseline and confidence is depressed at 6.85 versus 7.43, yet candor (7.0 vs 6.91) and specificity (7.55 vs 7.68) are essentially flat — leaning in costs confidence, not detail. Guidance behavior skews defensive: 24.4% lowered guidance versus 10.2% baseline, with only 9.3% raising. Thematically, 'Results Worse Than Direction' appears 1.62x the base rate (43.2% vs 43.2%... 43.2% share), and 'The Hidden Segment' runs 1.5x. The pattern surged to 15% of calls in 2023 after peaking at 11% in 2022, then fell to 5% in 2024. The 34-call returns sample shows a median one-year return of -18.98% versus -7.18% for the 149-call comparison set, with a 32.4% beat rate against 40.3%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor7.006.91+0.09
Evasion2.732.68+0.06
Specificity7.557.68-0.13
Stress2.782.24+0.54
Promotion4.995.13-0.14
Confidence6.857.43-0.58
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised9.3%22.5%
Maintained50.0%55.2%
Lowered24.4%10.2%
Withdrawn1.2%0.8%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Results Worse Than Direction1.62×69.8%43.2%
The Hidden Segment1.50×31.4%20.9%
The Question Left Hanging1.27×52.3%41.1%
Founder-Led Companies1.25×27.9%22.3%
Consolidation Among Peers0.68×14.0%20.4%
Skeptic Reassured0.73×54.7%74.6%
Volume About to Step Up0.75×19.8%26.4%
20150.03%
20160.06%
20170.05%
20180.03%
20190.00%
20200.00%
20210.04%
20220.11%
20230.15%
20240.05%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-19.0%-7.2%
Interquartile range-39.1% to +7.8%
Share beating SPY32.4% (95% CI 19%–49%)40.3%
Observations34149
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
ASOQ1 20242024-06-11C+
WRBYQ1 20242024-05-09A
HCKTQ1 20242024-05-08C
TNETQ1 20242024-04-26C
ASBQ1 20242024-04-25A
SNVQ1 20242024-04-18B
CLGNQ4 20232024-04-04F
HUYAQ4 20232024-03-19C

4Discussion

A careful reader should conclude that these calls are, on average, associated with rough terrain: lower confidence, more stress, more guidance cuts, and weaker subsequent one-year returns in the sampled subset. What should not be concluded is that the 'leaning in' behavior itself caused the weaker outcomes, that the pattern predicts returns for any individual stock, or that the 22.5% frequency says anything about timing. The 2023 spike coincided with a broadly difficult macro year, which likely inflates the count. The comparison medians describe groups, not forecasts, and the sample of calls with usable return data is small relative to the full set of 86.

5Limitations

All fields in this study are AI-read from transcripts and carry measurement noise; stress, candor, and theme tags are LLM judgments, not ground truth. The returns sample covers 22,449 calls and skews toward liquid names, so the 34-call subset may not generalize. Our own forward tests falsified directional prediction — the observed return gap should not be read as an exploitable signal. Finally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of these labels. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “The Storm Chasers: Earnings Calls That Lean Into Turbulence.” Artul.ai Earnings-Call Research Library, Study No. 47. https://artul.ai/research/hypothesis-leaning-into-the-storm

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecDon't Stick to the Script: Earnings Calls Where AnswHave Your Cake and Expand the Base Too: Repeat GrowtPressed Harder, Answered Straight: A Profile of 183
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.