Research › Management Behavior
Artul.ai Research LibraryStudy No. 70Management BehaviorUpdated 2026-08-28

All Stress and No Confidence: The Language of High-Stress Earnings Calls

By Artul.ai Research Group · n = 2,893 earnings calls · First published 2026-08-28
Abstract

This study examines earnings calls scoring 6 or higher on Artul.ai's 0-9 stress meter, a group covering just 1.75% of the 165,182-call corpus from 1990 to 2026 (n = 2,893). Stressed calls show sharply different language profiles than the baseline: evasion runs 4.25 versus 2.70, while confidence sits at 5.64 versus 7.21 and specificity at 6.84 versus 7.56. Guidance behavior diverges too: guidance was lowered on 32.5% of stressed calls versus 11.6% of baseline calls, and withdrawn on 10.5% versus 2.7%. Post-call returns in a 232-call subsample have a median of -0.118 versus -0.072 for the broader sample.

Key findings
  • Stressed calls show 55% higher evasion than the baseline (4.25 vs 2.70) and confidence 1.57 points lower (5.64 vs 7.21).
  • Guidance was lowered on 32.5% of stressed calls versus 11.6% of baseline calls, and withdrawn on 10.5% versus 2.7%.
  • The 'Scale-Dependent Advantage Claims' tag appears 3.94x more often than expected on stressed calls; 'Skeptic Reassured' appears 0.01x as often.
  • In the 232-call returns subsample, the median post-call return is -0.118 versus -0.072 in the 22,449-call baseline, with 32.3% beating versus 39.5%.
  • The stress-meter trend declines from 3.07 in 2015 to 1.45 in 2025, with a 2022 spike to 2.01.

1Introduction

Most earnings calls are routine: rehearsed scripts, friendly analysts, guidance nudged a point or two. But a small fraction of calls register unmistakable strain on Artul.ai's stress meter, and those calls are where the interesting communication happens. Do stressed management teams talk differently, guide differently, and see different outcomes than their calmer peers? This study examines 2,893 calls scoring 6 or higher on the 0-9 stress meter out of a 165,182-call corpus spanning 1990 to 2026. We compare their language profiles, guidance actions, signature phrases, and post-call returns against the full baseline, and trace how the stress-meter reading has drifted over the past decade.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 stress meter (n = 2,893; 1.8% of the reference set, 95% Wilson interval 1.7%–1.8%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The language profile separates stressed calls cleanly from the baseline: stress reads 6.21 versus 2.43, evasion 4.25 versus 2.70, and confidence 5.64 versus 7.21, with candor (6.46 vs 6.86) and specificity (6.84 vs 7.56) also lower. Guidance tells a consistent story: lowered guidance appears on 32.5% of stressed calls versus 11.6% of baseline, and withdrawn guidance on 10.5% versus 2.7%, while raised guidance is rarer (3.8% vs 21.1%). Tag lifts point the same direction: 'Scale-Dependent Advantage Claims' runs 3.94x expected, while reassurance tags like 'Skeptic Reassured' (0.01x) and 'Calls That Resolve Doubts' (0.10x) are nearly absent. In the 232-call returns subsample, the median return is -0.118 versus -0.072 baseline, and 32.3% beat versus 39.5%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.466.86-0.40
Evasion4.252.70+1.55
Specificity6.847.56-0.72
Stress6.212.43+3.78
Promotion4.835.05-0.22
Confidence5.647.21-1.57
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised3.8%21.1%
Maintained24.4%48.8%
Lowered32.5%11.6%
Withdrawn10.5%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims3.94×43.7%11.1%
The Question Left Hanging2.05×98.5%48.0%
The Finished-Story Tell1.85×8.1%4.4%
Underused Fixed Costs1.83×76.2%41.6%
Results Worse Than Direction1.64×83.8%51.1%
Skeptic Reassured0.01×0.9%66.4%
Calls That Resolve Doubts0.10×7.8%79.5%
Guidance Worth Underwriting0.15×10.7%71.5%
Confidence Proportionate to Evidence0.38×33.6%87.6%
Deferred Revenue Growing0.62×5.5%8.9%
20153.07%
20162.49%
20172.14%
20181.94%
20191.73%
20201.24%
20210.89%
20222.01%
20231.81%
20241.54%
20251.45%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-11.8%-7.2%
Interquartile range-43.7% to +10.8%
Share beating SPY32.3% (95% CI 27%–39%)39.5%
Observations23222,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
DOWQ2 20252025-07-24D
CMAQ2 20252025-07-18D
EDUCQ4 20252025-07-07F
CJREFQ3 20252025-06-26F
LVROQ2 20252025-06-20F
GRRRQ1 20252025-06-18F
MCHOYQ4 20252025-06-12F
NROMQ1 20252025-06-10F

4Discussion

A careful reader should conclude that high-stress calls are genuinely distinct: management teams on these calls evade more, project less confidence, and cut or withdraw guidance far more often than typical calls. What one should not conclude is that the stress meter causes anything, or that a high stress score predicts returns or beats. The returns differences are descriptive only, the subsample is small relative to the corpus, and the annual trend shows the meter's reading has been drifting lower over time, which complicates comparisons across years.

5Limitations

The language scores and tag lifts come from AI-read fields and are noisy, so individual call classifications carry error. The returns analysis covers 232 high-stress calls against a baseline of 22,449 calls skewed toward liquid names, so comparisons may reflect listing differences rather than stress. Our own forward tests falsified directional prediction from these signals. Finally, LLMs partially remember famous stocks' histories, which contaminates any backtest or lift computed on well-known names and should temper all conclusions here. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “All Stress and No Confidence: The Language of High-Stress Earnings Calls.” Artul.ai Earnings-Call Research Library, Study No. 70. https://artul.ai/research/management-under-visible-stress

Related studies

Specificity Is Next to Godliness: 161,829 Earnings CThe Truth Won't Set You Free: High Candor Is the NorNine Out of Ten Calls Are 'Confident' Now: Inside thStraight Talk, Stress Less: Low-Evasion Earnings CalThe Long Way to Say Less: Inside High-Complexity EarPromoted to the Front Page: High-Score Calls Across
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.