Evasion Meets Stress: A Census of the Most Guarded Earnings Calls
We examine a rare slice of the Artul.ai corpus of 165,182 earnings calls (1990-2026): the 657 calls (0.40%) where both evasion and stress scored 6/9 or higher. These calls read very differently from the corpus: candor averages 5.49 vs 6.86, specificity 6.03 vs 7.56, and confidence 5.50 vs 7.21, while evasion (6.55 vs 2.70) and stress (6.41 vs 2.43) dominate. Guidance behavior diverges sharply: 27.1% lowered and 14.0% withdrew guidance versus 11.6% and 2.7% corpus-wide. In the 45-call returns sample, the median return was -13.0% versus -7.2% for the 22,449-call base, and 28.9% beat versus 39.5% baseline. High-evasion, high-stress calls mark a distinctive, strained disclosure posture.
- High-evasion, high-stress calls make up 657 of 165,182 calls (0.40%), with a confidence score of 5.50 versus 7.21 corpus-wide.
- Guidance was withdrawn on 14.0% of these calls versus 2.7% of all calls, and lowered on 19.5% versus 11.6%.
- The 'Skeptic Reassured' narrative is absent from this group (lift 0.0 against a 0.15% base rate), while 'Scale-Dependent Advantage Claims' shows the strongest lift at 4.0x.
- Among 45 calls with return data, the median return was -13.0% versus -7.2% for the base sample, and 28.9% beat versus 39.5% baseline.
- Flagged-call rates fell steadily from 0.78 per 100 calls in 2015 to 0.27 in 2024.
1Introduction
Most earnings calls are forgettable; a small fraction feel like watching someone negotiate with a hostage-taker who is themselves. Calls that pair elevated evasion with elevated stress are exactly that: management ducking questions while audibly under strain. For anyone who parses earnings calls for a living, these moments are where soft language and hard consequences meet, and where guidance actions - withdrawals, cuts, hedges - cluster most visibly. Because both signals must clear a high bar simultaneously, the pattern is rare enough to study as a group rather than as anecdotes. This study profiles the 657 such calls in our corpus, comparing their language, guidance behavior, narrative fingerprints, frequency over time, and subsequent returns against the full 165,182-call base.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where evasion scored 6/9 or higher AND stress scored 6/9 or higher (n = 657; 0.4% of the reference set, 95% Wilson interval 0.4%–0.4%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral gap is stark: on these 657 calls, evasion averages 6.55 versus 2.70 corpus-wide and stress 6.41 versus 2.43, while confidence sits 1.71 points below the base. Guidance tells the same story - 14.0% withdrawn (vs 2.7%) and 19.5% lowered (vs 11.6%), with only 4.0% raised (vs 21.1%). Narrative lifts concentrate on defensive scripts: 'Scale-Dependent Advantage Claims' at 4.0x and 'The Question Left Hanging' at 2.09x, while reassurance narratives like 'Skeptic Reassured' vanish entirely. Frequency has declined from 0.78 flagged calls per 100 in 2015 to 0.27 in 2024. In the returns sample (n=45), the median outcome was -13.0% versus -7.2% for 22,449 base calls, and 28.9% beat versus 39.5% - a descriptive gap, not a signal.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 5.49 | 6.86 | -1.38 |
| Evasion | 6.55 | 2.70 | +3.85 |
| Specificity | 6.03 | 7.56 | -1.53 |
| Stress | 6.41 | 2.43 | +3.98 |
| Promotion | 4.88 | 5.05 | -0.17 |
| Confidence | 5.50 | 7.21 | -1.71 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 4.0% | 21.1% |
| Maintained | 27.1% | 48.8% |
| Lowered | 19.5% | 11.6% |
| Withdrawn | 14.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 4.00× | 44.3% | 11.1% |
| The Question Left Hanging | 2.09× | 100.0% | 48.0% |
| The Finished-Story Tell | 1.98× | 8.7% | 4.4% |
| Underused Fixed Costs | 1.67× | 69.4% | 41.6% |
| Results Worse Than Direction | 1.57× | 80.2% | 51.1% |
| Skeptic Reassured | 0.00× | 0.2% | 66.4% |
| Calls That Resolve Doubts | 0.01× | 0.8% | 79.5% |
| Guidance Worth Underwriting | 0.10× | 6.8% | 71.5% |
| Confidence Proportionate to Evidence | 0.21× | 18.1% | 87.6% |
| Pricing Recovering | 0.55× | 11.9% | 21.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -13.0% | -7.2% |
| Interquartile range | -41.6% to +13.9% | — |
| Share beating SPY | 28.9% (95% CI 18%–43%) | 39.5% |
| Observations | 45 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| PRPH | Q1 2025 | 2025-05-20 | F |
| LAZR | Q1 2025 | 2025-05-14 | F |
| SKM | Q1 2025 | 2025-05-12 | F |
| PMMAF | Q1 2025 | 2025-05-10 | F |
| CLQDF | Q1 2025 | 2025-05-09 | F |
| HAIN | Q3 2025 | 2025-05-07 | F |
| DSNY | Q2 2025 | 2025-04-14 | F |
| REKR | Q4 2024 | 2025-03-31 | F |
4Discussion
A careful reader should conclude that calls combining high evasion and high stress are linguistically distinctive, associated with defensive guidance actions, and, in this sample, followed by weaker outcomes on average. They should not conclude that the language caused the outcomes, that spotting such calls predicts returns, or that the pattern is tradeable. The returns comparison involves only 45 flagged calls against a 22,449-call base skewed toward liquid names, and wide quartiles (-41.6% to +13.9%) show enormous dispersion. Our own forward tests falsified directional prediction. Treat this as a description of a communication pattern under pressure, not a forecasting tool.
5Limitations
Evasion, stress, and companion scores are AI-read fields and inherit LLM noise and miscalibration. The returns sample covers only 45 flagged calls against a 22,449-call base concentrated in liquid names, limiting representativeness. Our own forward tests falsified directional prediction from these signals. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest that reuses historical calls; the descriptive gaps reported here may partly reflect that leakage rather than genuine market patterns. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.