Research › Business Verdicts
Artul.ai Research LibraryStudy No. 4Business VerdictsUpdated 2026-08-28

The Guidance Goes Quiet: Language and Outcomes of Withdrawal Calls

By Artul.ai Research Group · n = 4,393 earnings calls · First published 2026-08-28
Abstract

This study examines 4,393 earnings calls where guidance was read as withdrawn — about 2.7% of a 165,182-call corpus spanning 1990 to 2026. On these calls, management language shifts in a consistent direction: stress runs 0.92 points above the corpus baseline (3.35 vs 2.43), evasion is 0.59 points higher (3.28 vs 2.70), and confidence drops 0.83 points (6.39 vs 7.21), while promotion falls 0.42 and specificity slips 0.25. Candor is modestly elevated at 7.09 vs 6.86. Post-call returns are available for 127 such calls: the median is -15.9%, versus -7.2% for the 22,449-call base sample, and 35.4% beat, versus 39.5% at baseline. The 2020 spike — 15.9% of that year's calls — shows how withdrawal clusters in crisis periods.

Key findings
  • Withdrawn-guidance calls make up 2.7% of the corpus (4,393 of 165,182 calls).
  • Stress language is 0.92 points higher than baseline on these calls (3.35 vs 2.43), the largest profile gap.
  • Confidence is 0.83 points lower (6.39 vs 7.21) and promotion is 0.42 points lower (4.63 vs 5.05) than the corpus baseline.
  • The median post-call return on 127 withdrawal calls is -15.9%, versus -7.2% for the 22,449-call base sample.
  • Withdrawal peaked at 15.9% of calls in 2020, far above every other year in the trend.

1Introduction

Few sentences on an earnings call land harder than an announcement that guidance is being withdrawn. It removes the one number Wall Street anchors on, and it typically arrives amid operational uncertainty rather than in spite of it. For analysts, the question is whether the surrounding call — the tone, the hedging, the language patterns — carries a recognizable signature, and whether the aftermath looks systematically different from a typical call. This study examines 4,393 calls where guidance was read as withdrawn, drawn from a corpus of 165,182 calls covering 1990 through 2026, comparing their language profiles, post-call returns, and thematic content against the rest of the corpus.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where guidance was read as withdrawn (n = 4,393; 2.7% of the reference set, 95% Wilson interval 2.6%–2.7%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The language profile of withdrawal calls tilts toward difficulty: stress is up 0.92 points (3.35 vs 2.43) and evasion up 0.59 (3.28 vs 2.70), while confidence is down 0.83 (6.39 vs 7.21) and promotion down 0.42 (4.63 vs 5.05). Notably, candor is slightly higher (7.09 vs 6.86) — these calls read as strained but not evasive-first. Themes overrepresented include 'The Question Left Hanging' (1.57x) and 'Results Worse Than Direction' (1.47x). Post-call outcomes skew weak: the median return is -15.9% versus -7.2% at baseline, with 35.4% beating versus 39.5% at baseline — a wide interquartile range (-39.2% to 7.6%) cautions against treating withdrawal as a uniform event.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor7.096.86+0.23
Evasion3.282.70+0.59
Specificity7.317.56-0.25
Stress3.352.43+0.92
Promotion4.635.05-0.42
Confidence6.397.21-0.83
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised0.0%21.1%
Maintained0.0%48.8%
Lowered0.0%11.6%
Withdrawn100.0%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Underused Fixed Costs1.62×67.4%41.6%
The Question Left Hanging1.57×75.1%48.0%
Results Worse Than Direction1.47×75.0%51.1%
Guidance Worth Underwriting0.12×8.3%71.5%
Pricing Recovering0.58×12.4%21.5%
Calls That Read Rehearsed0.69×26.8%38.7%
Deferred Revenue Growing0.71×6.3%8.9%
20151.49%
20160.95%
20170.65%
20180.69%
20191.41%
202015.92%
20211.90%
20221.35%
20230.94%
20240.74%
20251.85%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-15.9%-7.2%
Interquartile range-39.2% to +7.6%
Share beating SPY35.4% (95% CI 28%–44%)39.5%
Observations12722,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
WZZAFQ1 20262025-07-25F
VICRQ2 20252025-07-22F
AEHRQ4 20252025-07-08C
AOUTQ4 20252025-06-26C
FDXQ4 20252025-06-24C
LVROQ2 20252025-06-20F
VNCEQ1 20252025-06-17F
JILLQ1 20252025-06-11F

4Discussion

A careful reader should conclude that withdrawn-guidance calls carry a distinct observable profile: more stress and evasion, less confidence and promotion, weaker median post-call returns, and heavy clustering in 2020. What should not be concluded is that withdrawal itself causes poor outcomes, that the language gap implies intent or deception, or that any of these statistics predicts returns — the distributions overlap substantially, and the beat-rate difference (35.4% vs 39.5%) is modest. The elevated candor score suggests withdrawal often accompanies frank acknowledgment, not concealment. These are descriptive patterns measured after the fact; they describe what withdrawal calls look like, not what to trade on.

5Limitations

The language fields are AI-read and noisy, so profile deltas like the 0.92-point stress gap should be read as tendencies, not precise measurements. The returns sample is 22,449 calls skewed toward liquid names, and only 127 withdrawal calls have returns attached — a small, selection-prone subset. Our own forward tests falsified directional prediction, and LLMs partially remember famous stocks' histories, contaminating any backtest. The 2020 spike reflects a unique macro episode rather than a repeatable pattern. Nothing here establishes causation, prediction, or a trading edge. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “The Guidance Goes Quiet: Language and Outcomes of Withdrawal Calls.” Artul.ai Earnings-Call Research Library, Study No. 4. https://artul.ai/research/after-guidance-is-withdrawn-earnings-calls

Related studies

Steady As She Goes: What Happens When Guidance HoldsMargins Are Expanding, and So Is the Confidence: ReaCapex Up, Stress Down: Reading 64,183 Calls Where SpGuidance Is a Vibe: Calls Where Demand Reads as AcceBacklog Is Growing, and So Is the Confidence: EarninMargins Are Fine, Thanks for Asking: What Contractio
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.