Research › Management Behavior
Artul.ai Research LibraryStudy No. 29Management BehaviorUpdated 2026-08-28

Reading the Silence: Inside the Rarest Candor Scores on Earnings Calls

By Artul.ai Research Group · n = 92 earnings calls · First published 2026-08-28
Abstract

This study examines the rarest tail of Artul.ai's candor meter: earnings calls scoring 2 or lower on the 0-9 scale. Across a corpus of 165,182 calls spanning 1990 to 2026, only 92 calls qualify, a share of 0.057% (95% CI 0.045% to 0.068%). These calls show large behavioral gaps versus typical calls: specificity runs 7.56 points below baseline, confidence 7.21 below, and candor itself 6.86 below. None of the 92 calls resolved skeptic doubts, underwrote guidance, or reassured skeptics. Strikingly, the pattern is concentrated in recent years: the 2024 rate of 0.15% is roughly 15 times the 0.01% typical of 2016 through 2024, and 2025 reaches 0.93%.

Key findings
  • Only 92 of 165,182 calls (0.057%, 95% CI 0.045% to 0.068%) score 2 or lower on the 0-9 candor meter.
  • These calls show specificity 7.56 points below baseline, confidence 7.21 below, and candor 6.86 below, with promotion only 5.05 below.
  • The 'The Question Left Hanging' theme appears with a lift of 2.09 over expectations, and 'Calls That Read Rehearsed' with a lift of 1.68.
  • The annual share rose from 0.01% in most years from 2016 to 2024 to 0.15% in 2024 and 0.93% in 2025.

1Introduction

Most earnings-call research focuses on the typical call. The tails can be more revealing. On the 0-9 candor meter, a score of 2 or lower marks calls where managers appear least forthcoming: low candor, low specificity, low confidence. These calls are extraordinarily rare in the historical record, which makes any change in their frequency notable, and the recent uptick in low-candor scoring is the kind of pattern a careful listener would want documented before drawing conclusions. This study profiles those calls: how they differ on behavioral dimensions, which conversational themes they over- or under-represent, and how their frequency has moved from 2015 through 2025.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 2 or lower on the 0–9 candor meter (n = 92; 0.1% of the reference set, 95% Wilson interval 0.0%–0.1%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The behavioral profile is stark. Relative to typical calls, the 92 low-candor calls run 7.56 points lower on specificity, 7.21 lower on confidence, 6.86 lower on candor, and 3.83 lower on promotion, while stress is only 1.77 lower and evasion 2.25 lower, suggesting these calls read as subdued rather than defensive. Thematically, 'The Question Left Hanging' appears 2.09 times more often than expected and 'Calls That Read Rehearsed' 1.68 times; founder-led companies show a lift of 1.46. Themes tied to clarity, 'Calls That Resolve Doubts', 'Guidance Worth Underwriting', 'Skeptic Reassured', 'The Finished-Story Tell', and 'The Hidden Segment', never appear. Guidance action rates are tiny: 2.17% raised and 1.09% maintained, with none lowered or withdrawn.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor0.126.86-6.74
Evasion0.452.70-2.25
Specificity0.347.56-7.22
Stress0.662.43-1.77
Promotion1.225.05-3.83
Confidence2.057.21-5.16
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised2.2%21.1%
Maintained1.1%48.8%
Lowered0.0%11.6%
Withdrawn0.0%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
The Question Left Hanging2.09×100.0%48.0%
Calls That Read Rehearsed1.68×65.2%38.7%
Founder-Led Companies1.46×29.3%20.0%
Calls That Resolve Doubts0.00×0.0%79.5%
Guidance Worth Underwriting0.00×0.0%71.5%
Skeptic Reassured0.00×0.0%66.4%
The Finished-Story Tell0.00×0.0%4.4%
The Hidden Segment0.00×0.0%21.1%
20150.00%
20160.01%
20170.01%
20180.01%
20190.01%
20200.01%
20210.00%
20220.01%
20230.01%
20240.15%
20250.93%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
CRDOFQ1 20252025-05-29F
AMWDQ4 20252025-05-29F
WRDQ1 20252025-05-21F
XPEVQ1 20252025-05-21F
DDDQ1 20252025-05-13F
BLZEQ1 20252025-05-11F
NOMDQ1 20252025-05-10F
JYNTQ1 20252025-05-10F

4Discussion

A careful reader should treat this as a description of a rare conversational profile, not a verdict on any company. The 92 calls are defined by the candor meter itself, so their behavioral gaps partly reflect the scoring method. The rise in frequency from 0.01% in most of 2016 through 2024 to 0.15% in 2024 and 0.93% in 2025 is a real observation in this corpus, but the 2025 figure rests on 6,012 scored calls, a smaller base, and nothing here establishes why the rate moved. No causal claim and no investment implication should be drawn from these statistics.

5Limitations

Behavioral scores and thematic tags are produced by AI reading of transcripts and are noisy; the candor meter is a model output, not ground truth. The returns sample covers 22,449 calls and is skewed toward liquid names, so it does not represent the full universe. Our own forward tests falsified directional prediction from these signals, and LLMs partially remember famous stocks' histories, contaminating any backtest. The 92-call sample is small, and 2025 covers only part of the year. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Reading the Silence: Inside the Rarest Candor Scores on Earnings Calls.” Artul.ai Earnings-Call Research Library, Study No. 29. https://artul.ai/research/guarded-management

Related studies

Specificity Is Next to Godliness: 161,829 Earnings CThe Truth Won't Set You Free: High Candor Is the NorNine Out of Ten Calls Are 'Confident' Now: Inside thStraight Talk, Stress Less: Low-Evasion Earnings CalThe Long Way to Say Less: Inside High-Complexity EarPromoted to the Front Page: High-Score Calls Across
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.