Research › 20-Year Trends
Artul.ai Research LibraryStudy No. 720-Year TrendsUpdated 2026-08-28

Give Them Nothing: The Calls Where Critics Feast

By Artul.ai Research Group · n = 145,051 earnings calls · First published 2026-08-28
Abstract

We examine 145,051 earnings calls from 1990 to 2026 where Artul.ai's model answered YES to "Critic Ammunition" — material a critic could seize on. These calls make up 87.81% of the 165,182-call corpus (95% CI: 87.65%–87.97%). Relative to other calls, they show slightly lower candor (6.83 vs 6.86), higher evasion (2.79 vs 2.70), and higher stress (2.58 vs 2.43). Post-call returns for a 19,045-call subset have a median of -0.08 and a 38.38% beat rate versus 39.47% for the 22,449-call baseline. Guidance is lowered in 12.96% of these calls versus 11.56% otherwise. This is descriptive, not predictive.

Key findings
  • 87.81% of the 165,182-call corpus (145,051 calls) was flagged as containing critic ammunition.
  • Flagged calls show higher evasion (2.79 vs 2.70) and stress (2.58 vs 2.43) than the rest of the corpus.
  • Median post-call return in the flagged returns sample is -0.08 vs -0.07 for the 22,449-call baseline.
  • Guidance is lowered in 12.96% of flagged calls versus 11.56% of the remainder.

1Introduction

Every earnings call hands skeptics something: a hedged answer, a soft metric, an awkward pause. Whether that material is common — and what else it travels with — is a fair empirical question for anyone who reads transcripts closely. Using Artul.ai's "Critic Ammunition" field, we identify calls where the model answered YES and track them by year across a 1990–2026 corpus of 165,182 calls. The study describes how frequent the flag is, how tone scores and guidance behavior differ, and how returns distribute — without claiming any predictive value.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "Critic Ammunition", tracked by year (n = 145,051; 87.8% of the reference set, 95% Wilson interval 87.7%–88.0%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The flag is nearly universal: 87.81% of calls (n=145,051) qualify, with a 95% CI of 87.65% to 87.97%. Annual rates range from 84.29% (2021) to 90.98% (2015). Flagged calls skew slightly toward defensiveness: evasion 2.79 vs 2.70, stress 2.58 vs 2.43, candor 6.83 vs 6.86, confidence 7.13 vs 7.21. Guidance is lowered in 12.96% of flagged calls versus 11.56% otherwise, while raised-guidance rates sit at 18.59% vs 21.05%. In the 19,045-call returns subset, the median return is -0.08 vs -0.07 in the 22,449-call baseline, with beat rates of 38.38% vs 39.47%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.836.86-0.03
Evasion2.792.70+0.10
Specificity7.517.56-0.04
Stress2.582.43+0.15
Promotion5.085.05+0.03
Confidence7.137.21-0.09
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised18.6%21.1%
Maintained48.5%48.8%
Lowered13.0%11.6%
Withdrawn2.9%2.7%
201590.98%
201687.51%
201784.33%
201886.08%
201988.47%
202088.06%
202184.29%
202289.99%
202390.50%
202489.53%
202587.91%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 3. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-8.2%-7.2%
Interquartile range-27.5% to +11.2%
Share beating SPY38.4% (95% CI 38%–39%)39.5%
Observations19,04522,449
Table 4. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
DOCQ2 20252025-07-25C
USCBQ2 20252025-07-25B+
HCAQ2 20252025-07-25C
AONQ2 20252025-07-25C
CNCQ2 20252025-07-25F
NWGQ2 20252025-07-25B+
BFHQ2 20252025-07-25B
FFICQ2 20252025-07-25B+

4Discussion

A careful reader should conclude only that critic-ammunition flags co-occur with modestly more evasive, more stressed language and slightly weaker guidance and return outcomes. These are small descriptive gaps, not evidence that the flag causes anything or forecasts returns. The flag's near-universality (87.81%) limits its discriminating power: it may simply describe the nature of earnings calls generally. No causal claim, timing signal, or trading edge is supported by this data.

5Limitations

AI-read fields like Critic Ammunition are noisy and unvalidated against ground truth. The returns sample covers only 22,449 calls and is skewed toward liquid names, so return comparisons may not generalize. Our own forward tests falsified directional prediction, so nothing here should be read as an edge. Additionally, LLMs partially remember famous stocks' histories, contaminating any backtest that mixes model outputs with known outcomes. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Give Them Nothing: The Calls Where Critics Feast.” Artul.ai Earnings-Call Research Library, Study No. 7. https://artul.ai/research/are-calls-getting-easier-to-criticize

Related studies

Yes Means Yes: Inside the Earnings Calls Where the SThe Question Was Answered. Just Not in This Room: HaRehearsal Season: The Earnings Calls That Sound a LiReady, Set, Step Up: The Calls Where Volume Was SuppLike Father, Like Firm: Founder-Led Earnings Calls, Speak Softly and Carry the Spreadsheet: CFO-Dominate
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.