Research › Signal Combinations
Artul.ai Research LibraryStudy No. 25Signal CombinationsUpdated 2026-08-28

Evasion Meets Stress: A Census of the Most Guarded Earnings Calls

By Artul.ai Research Group · n = 657 earnings calls · First published 2026-08-28
Abstract

We examine a rare slice of the Artul.ai corpus of 165,182 earnings calls (1990-2026): the 657 calls (0.40%) where both evasion and stress scored 6/9 or higher. These calls read very differently from the corpus: candor averages 5.49 vs 6.86, specificity 6.03 vs 7.56, and confidence 5.50 vs 7.21, while evasion (6.55 vs 2.70) and stress (6.41 vs 2.43) dominate. Guidance behavior diverges sharply: 27.1% lowered and 14.0% withdrew guidance versus 11.6% and 2.7% corpus-wide. In the 45-call returns sample, the median return was -13.0% versus -7.2% for the 22,449-call base, and 28.9% beat versus 39.5% baseline. High-evasion, high-stress calls mark a distinctive, strained disclosure posture.

Key findings
  • High-evasion, high-stress calls make up 657 of 165,182 calls (0.40%), with a confidence score of 5.50 versus 7.21 corpus-wide.
  • Guidance was withdrawn on 14.0% of these calls versus 2.7% of all calls, and lowered on 19.5% versus 11.6%.
  • The 'Skeptic Reassured' narrative is absent from this group (lift 0.0 against a 0.15% base rate), while 'Scale-Dependent Advantage Claims' shows the strongest lift at 4.0x.
  • Among 45 calls with return data, the median return was -13.0% versus -7.2% for the base sample, and 28.9% beat versus 39.5% baseline.
  • Flagged-call rates fell steadily from 0.78 per 100 calls in 2015 to 0.27 in 2024.

1Introduction

Most earnings calls are forgettable; a small fraction feel like watching someone negotiate with a hostage-taker who is themselves. Calls that pair elevated evasion with elevated stress are exactly that: management ducking questions while audibly under strain. For anyone who parses earnings calls for a living, these moments are where soft language and hard consequences meet, and where guidance actions - withdrawals, cuts, hedges - cluster most visibly. Because both signals must clear a high bar simultaneously, the pattern is rare enough to study as a group rather than as anecdotes. This study profiles the 657 such calls in our corpus, comparing their language, guidance behavior, narrative fingerprints, frequency over time, and subsequent returns against the full 165,182-call base.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where evasion scored 6/9 or higher AND stress scored 6/9 or higher (n = 657; 0.4% of the reference set, 95% Wilson interval 0.4%–0.4%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The behavioral gap is stark: on these 657 calls, evasion averages 6.55 versus 2.70 corpus-wide and stress 6.41 versus 2.43, while confidence sits 1.71 points below the base. Guidance tells the same story - 14.0% withdrawn (vs 2.7%) and 19.5% lowered (vs 11.6%), with only 4.0% raised (vs 21.1%). Narrative lifts concentrate on defensive scripts: 'Scale-Dependent Advantage Claims' at 4.0x and 'The Question Left Hanging' at 2.09x, while reassurance narratives like 'Skeptic Reassured' vanish entirely. Frequency has declined from 0.78 flagged calls per 100 in 2015 to 0.27 in 2024. In the returns sample (n=45), the median outcome was -13.0% versus -7.2% for 22,449 base calls, and 28.9% beat versus 39.5% - a descriptive gap, not a signal.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor5.496.86-1.38
Evasion6.552.70+3.85
Specificity6.037.56-1.53
Stress6.412.43+3.98
Promotion4.885.05-0.17
Confidence5.507.21-1.71
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised4.0%21.1%
Maintained27.1%48.8%
Lowered19.5%11.6%
Withdrawn14.0%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims4.00×44.3%11.1%
The Question Left Hanging2.09×100.0%48.0%
The Finished-Story Tell1.98×8.7%4.4%
Underused Fixed Costs1.67×69.4%41.6%
Results Worse Than Direction1.57×80.2%51.1%
Skeptic Reassured0.00×0.2%66.4%
Calls That Resolve Doubts0.01×0.8%79.5%
Guidance Worth Underwriting0.10×6.8%71.5%
Confidence Proportionate to Evidence0.21×18.1%87.6%
Pricing Recovering0.55×11.9%21.5%
20150.78%
20160.70%
20170.59%
20180.51%
20190.40%
20200.27%
20210.28%
20220.34%
20230.30%
20240.27%
20250.18%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-13.0%-7.2%
Interquartile range-41.6% to +13.9%
Share beating SPY28.9% (95% CI 18%–43%)39.5%
Observations4522,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
PRPHQ1 20252025-05-20F
LAZRQ1 20252025-05-14F
SKMQ1 20252025-05-12F
PMMAFQ1 20252025-05-10F
CLQDFQ1 20252025-05-09F
HAINQ3 20252025-05-07F
DSNYQ2 20252025-04-14F
REKRQ4 20242025-03-31F

4Discussion

A careful reader should conclude that calls combining high evasion and high stress are linguistically distinctive, associated with defensive guidance actions, and, in this sample, followed by weaker outcomes on average. They should not conclude that the language caused the outcomes, that spotting such calls predicts returns, or that the pattern is tradeable. The returns comparison involves only 45 flagged calls against a 22,449-call base skewed toward liquid names, and wide quartiles (-41.6% to +13.9%) show enormous dispersion. Our own forward tests falsified directional prediction. Treat this as a description of a communication pattern under pressure, not a forecasting tool.

5Limitations

Evasion, stress, and companion scores are AI-read fields and inherit LLM noise and miscalibration. The returns sample covers only 45 flagged calls against a 22,449-call base concentrated in liquid names, limiting representativeness. Our own forward tests falsified directional prediction from these signals. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest that reuses historical calls; the descriptive gaps reported here may partly reflect that leakage rather than genuine market patterns. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Evasion Meets Stress: A Census of the Most Guarded Earnings Calls.” Artul.ai Earnings-Call Research Library, Study No. 25. https://artul.ai/research/evasive-and-under-stress

Related studies

Small Pond, Big Fish: Calls Claiming Fast Growth in Trust the Raise: When Skeptics Get Reassured and GuiThe Founder Is Present: Candor and Confidence on EarLoud Backlogs and Busy Phones: Calls Where the ModelFounder Knows Best: High-Promotion Founder-Led EarniThe Quiet Hand on the Microphone: CFO Dominance Meet
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.