Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 61Hypotheses TestedUpdated 2026-08-28

No One Saw This Coming: What 142 'Unpredictable' Quarters Sound Like

By Artul.ai Research Group · n = 142 earnings calls · First published 2026-08-28
Abstract

This study examines earnings calls that answered YES to the hypothesis "This quarter could not have been described last quarter" — calls whose content genuinely exceeded what was knowable one quarter earlier. Of 318 calls scored between 2015 and 2024, 142 (44.65%) qualified. Compared with other calls, these quarters show slightly lower candor (6.78 vs 6.85) but higher specificity (7.72 vs 7.62), lower stress (1.98 vs 2.40), and higher confidence (7.64 vs 7.29). Guidance behavior diverges sharply: 28.87% raised guidance versus 21.07% elsewhere, and only 7.75% lowered it versus 15.72%. Forward 90-day returns for a 47-call subset had a median of -7.18% versus -10.54% for the base. The study is descriptive and offers no predictive claim.

Key findings
  • 142 of 318 calls (44.65%, 95% CI 39.29%-50.15%) answered YES to the hypothesis that the quarter could not have been described last quarter.
  • These calls show lower stress (1.98 vs 2.40) and higher confidence (7.64 vs 7.29) than other calls, with specificity up (7.72 vs 7.62) and candor down slightly (6.78 vs 6.85).
  • Guidance was raised on 28.87% of these calls versus 21.07% of others, and lowered on only 7.75% versus 15.72%; none withdrew guidance (0.0%).
  • Among 47 calls with forward returns, the median 90-day return was -7.18% versus -10.54% for the base, with 42.55% beating the market versus 41.32% (CI 29.51%-56.72%).

1Introduction

Most earnings-call language is pre-scripted: themes, metrics, and even surprises can be previewed in the prior quarter's outlook. A small set of calls breaks that pattern — management describes a quarter that, by construction, could not have been telegraphed. For anyone who parses calls for a living, these moments are the interesting edge cases: do speakers become more guarded when they are off-script, or more expansive? Do they raise or cut guidance more often? Using AI-scored behavioral fields on calls from 2015 through 2024, this study characterizes the 142 calls (of 318) that answered YES to the hypothesis "This quarter could not have been described last quarter," and compares their language, guidance actions, and subsequent returns with the rest.

2Data & methodology

The corpus comprises 318 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "This quarter could not have been described last quarter" (n = 142; 44.7% of the reference set, 95% Wilson interval 39.3%–50.1%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The behavioral profile of these genuinely surprising quarters is distinctive: stress is lower (1.98 vs 2.40) and confidence higher (7.64 vs 7.29) than on other calls, while specificity edges up (7.72 vs 7.62) and candor slips slightly (6.78 vs 6.85). Promotion rises the most (5.61 vs 5.19, +0.42). Guidance actions diverge notably: 28.87% raised guidance versus 21.07% elsewhere, and 7.75% lowered it versus 15.72%, with no withdrawals. Over-indexed phrases include "Early Products Growing Fast" (lift 1.36) and "A Tiny Fraction of the Market" (1.32). Returns offer no clean edge: the 47-call subset's median 90-day return of -7.18% sits above the base's -10.54%, but the beat rate, 42.55%, is statistically indistinguishable from the base's 41.32%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.786.85-0.06
Evasion2.642.73-0.09
Specificity7.727.62+0.10
Stress1.982.40-0.42
Promotion5.615.19+0.42
Confidence7.647.29+0.35
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised28.9%21.1%
Maintained52.1%50.3%
Lowered7.7%15.7%
Withdrawn0.0%1.3%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Early Products Growing Fast1.36×54.2%39.9%
A Tiny Fraction of the Market1.32×43.7%33.0%
Founder-Led Companies1.30×30.3%23.3%
Volume About to Step Up1.29×32.4%25.2%
20150.03%
20160.08%
20170.03%
20180.12%
20190.02%
20200.00%
20210.11%
20220.20%
20230.16%
20240.09%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-7.2%-10.5%
Interquartile range-34.4% to +12.5%
Share beating SPY42.6% (95% CI 30%–57%)41.3%
Observations47121
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
SFIXQ3 20242024-06-04C+
CRGOQ1 20242024-05-20C+
WRBYQ1 20242024-05-09A
HCKTQ1 20242024-05-08C
LINCQ1 20242024-05-06B+
AESQ1 20242024-05-03C+
PPCQ1 20242024-05-03A
CTRAQ1 20242024-05-03A

4Discussion

The honest reading is descriptive. Calls answering YES to this hypothesis tend to coincide with confident, promotion-heavy language and more frequent guidance raises — an association, not a mechanism. The returns comparison is deliberately underwhelming: a higher median return against the base does not imply these calls outperform, because the 47-call returns sample is small, the confidence interval on the beat rate (29.51%-56.72%) spans the base rate, and medians can differ for many reasons. A careful reader should treat the phrase and guidance patterns as characterization of a corpus, not as signals. Nothing here supports predicting returns, selecting stocks, or assuming the labeled surprise is economically meaningful.

5Limitations

The behavioral fields (candor, evasion, specificity, stress, promotion, confidence) are AI-assigned and noisy; they should not be read as ground truth. The returns sample covers 47 calls out of 22,449 in the underlying corpus and is skewed toward liquid, widely covered names. Our own forward tests of this labeling approach falsified any directional prediction — the apparent return differences do not survive honest out-of-sample evaluation. Additionally, LLMs partially remember famous stocks' histories, which can contaminate labels and any backtest-looking statistic. All figures here are descriptive summaries of a scored corpus, not evidence of a tradable effect. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “No One Saw This Coming: What 142 'Unpredictable' Quarters Sound Like.” Artul.ai Earnings-Call Research Library, Study No. 61. https://artul.ai/research/hypothesis-this-quarter-could-not-have-been-described-last-quarter

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecDon't Stick to the Script: Earnings Calls Where AnswHave Your Cake and Expand the Base Too: Repeat GrowtPressed Harder, Answered Straight: A Profile of 183
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.