Research › Management Behavior
Artul.ai Research LibraryStudy No. 69Management BehaviorUpdated 2026-08-28

Straight Talk, Stress Less: Low-Evasion Earnings Calls, 1990–2026

By Artul.ai Research Group · n = 86,188 earnings calls · First published 2026-08-28
Abstract

We examined 86,188 earnings calls—52.2% of a 165,182-call corpus spanning 1990 to 2026—that scored 2 or lower on a 0–9 evasion meter. These candid calls ran higher on candor (7.12 vs 6.86), specificity (7.85 vs 7.56), and confidence (7.36 vs 7.21), while showing less evasion (1.79 vs 2.70) and stress (2.02 vs 2.43) than the base. Guidance profiles leaned more positive: 24.1% raised guidance versus 21.1% in the base, with only 1.7% withdrawn versus 2.7%. In the matched sample of 10,966 calls with returns, the median return was -0.061 versus -0.072 for the base, and 40.6% beat versus 39.5%.

Key findings
  • Low-evasion calls made up 52.2% of the 165,182-call corpus (86,188 calls, 95% CI 51.9%–52.4%).
  • These calls scored higher on candor (7.12 vs 6.86) and specificity (7.85 vs 7.56) and lower on stress (2.02 vs 2.43) than the base.
  • Guidance was raised after 24.1% of low-evasion calls versus 21.1% of base calls, and withdrawn after just 1.7% versus 2.7%.
  • Among 10,966 calls with post-call returns, the median return was -0.061 versus -0.072 for the 22,449-call base, with 40.6% beating versus 39.5%.

1Introduction

Evasion on earnings calls is easy to sense and hard to quantify: dodged questions, vague scale claims, answers that end nowhere near where they started. Because call transcripts are now read by analysts, journalists, and machines alike, systematic differences in how directly management speaks are worth measuring at scale. Low-evasion calls are the natural contrast group—management answering what was asked. This study examines 86,188 calls scoring 2 or lower on a 0–9 evasion meter, drawn from a corpus of 165,182 calls covering 1990 through 2026, and compares their language profiles, guidance actions, and post-call returns against the rest of the corpus.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 2 or lower on the 0–9 evasion meter (n = 86,188; 52.2% of the reference set, 95% Wilson interval 51.9%–52.4%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Low-evasion calls differ from the base on every language dimension measured: candor 7.12 vs 6.86, specificity 7.85 vs 7.56, confidence 7.36 vs 7.21, with evasion itself at 1.79 vs 2.70 and stress at 2.02 vs 2.43. The only overrepresented evasion topics are 'The Question Left Hanging' (55% over the base's 26.5% share) and 'Scale-Dependent Advantage Claims' (6.7% vs 11.1%). Guidance skews favorable: 24.1% raised versus 21.1% in the base, and 1.7% withdrawn versus 2.7%. Returns are modestly less negative—median -0.061 vs -0.072—and 40.6% beat versus 39.5%. The annual share rose from 48.0% in 2015 to 57.95% in 2025.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor7.126.86+0.26
Evasion1.792.70-0.91
Specificity7.857.56+0.29
Stress2.022.43-0.41
Promotion4.895.05-0.16
Confidence7.367.21+0.15
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised24.1%21.1%
Maintained50.1%48.8%
Lowered10.4%11.6%
Withdrawn1.7%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
The Question Left Hanging0.55×26.5%48.0%
Scale-Dependent Advantage Claims0.60×6.7%11.1%
201548.03%
201649.35%
201750.49%
201850.82%
201950.13%
202050.60%
202153.90%
202253.91%
202354.34%
202454.65%
202557.95%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-6.1%-7.2%
Interquartile range-23.3% to +11.6%
Share beating SPY40.6% (95% CI 40%–42%)39.5%
Observations10,96622,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
SBFGQ2 20252025-07-25A
USCBQ2 20252025-07-25B+
NWGQ2 20252025-07-25B+
FFICQ2 20252025-07-25B+
GBCIQ2 20252025-07-25A
MOG.AQ3 20252025-07-25B+
FLGQ2 20252025-07-25B
FRSTQ2 20252025-07-25A

4Discussion

A careful reader should conclude that calls where management avoids dodging are, in this dataset, associated with slightly stronger language profiles, somewhat more favorable guidance actions, and marginally less negative median post-call returns. These are descriptive associations, not proof that candor causes better outcomes—firms that answer directly may simply differ in other ways from firms that do not. The gaps are also small in practical terms, and the returns comparison covers only a subset of calls. Nothing here should be read as a signal for predicting returns or picking trades.

5Limitations

Evasion, candor, and related scores are AI-read fields and inherently noisy; misclassification is expected at scale. The returns sample covers 22,449 base calls and is skewed toward liquid names, limiting generalizability. Our own forward tests falsified directional prediction from these features, so no predictive claim is made. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest built on model-scored transcripts, including ours. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Straight Talk, Stress Less: Low-Evasion Earnings Calls, 1990–2026.” Artul.ai Earnings-Call Research Library, Study No. 69. https://artul.ai/research/management-that-doesn-t-dodge

Related studies

Specificity Is Next to Godliness: 161,829 Earnings CThe Truth Won't Set You Free: High Candor Is the NorNine Out of Ten Calls Are 'Confident' Now: Inside thThe Long Way to Say Less: Inside High-Complexity EarPromoted to the Front Page: High-Score Calls Across Bad News Comes in Sixes: What High-Uncertainty Calls
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.