Reading the Silence: Inside the Rarest Candor Scores on Earnings Calls
This study examines the rarest tail of Artul.ai's candor meter: earnings calls scoring 2 or lower on the 0-9 scale. Across a corpus of 165,182 calls spanning 1990 to 2026, only 92 calls qualify, a share of 0.057% (95% CI 0.045% to 0.068%). These calls show large behavioral gaps versus typical calls: specificity runs 7.56 points below baseline, confidence 7.21 below, and candor itself 6.86 below. None of the 92 calls resolved skeptic doubts, underwrote guidance, or reassured skeptics. Strikingly, the pattern is concentrated in recent years: the 2024 rate of 0.15% is roughly 15 times the 0.01% typical of 2016 through 2024, and 2025 reaches 0.93%.
- Only 92 of 165,182 calls (0.057%, 95% CI 0.045% to 0.068%) score 2 or lower on the 0-9 candor meter.
- These calls show specificity 7.56 points below baseline, confidence 7.21 below, and candor 6.86 below, with promotion only 5.05 below.
- The 'The Question Left Hanging' theme appears with a lift of 2.09 over expectations, and 'Calls That Read Rehearsed' with a lift of 1.68.
- The annual share rose from 0.01% in most years from 2016 to 2024 to 0.15% in 2024 and 0.93% in 2025.
1Introduction
Most earnings-call research focuses on the typical call. The tails can be more revealing. On the 0-9 candor meter, a score of 2 or lower marks calls where managers appear least forthcoming: low candor, low specificity, low confidence. These calls are extraordinarily rare in the historical record, which makes any change in their frequency notable, and the recent uptick in low-candor scoring is the kind of pattern a careful listener would want documented before drawing conclusions. This study profiles those calls: how they differ on behavioral dimensions, which conversational themes they over- or under-represent, and how their frequency has moved from 2015 through 2025.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 2 or lower on the 0–9 candor meter (n = 92; 0.1% of the reference set, 95% Wilson interval 0.0%–0.1%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile is stark. Relative to typical calls, the 92 low-candor calls run 7.56 points lower on specificity, 7.21 lower on confidence, 6.86 lower on candor, and 3.83 lower on promotion, while stress is only 1.77 lower and evasion 2.25 lower, suggesting these calls read as subdued rather than defensive. Thematically, 'The Question Left Hanging' appears 2.09 times more often than expected and 'Calls That Read Rehearsed' 1.68 times; founder-led companies show a lift of 1.46. Themes tied to clarity, 'Calls That Resolve Doubts', 'Guidance Worth Underwriting', 'Skeptic Reassured', 'The Finished-Story Tell', and 'The Hidden Segment', never appear. Guidance action rates are tiny: 2.17% raised and 1.09% maintained, with none lowered or withdrawn.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 0.12 | 6.86 | -6.74 |
| Evasion | 0.45 | 2.70 | -2.25 |
| Specificity | 0.34 | 7.56 | -7.22 |
| Stress | 0.66 | 2.43 | -1.77 |
| Promotion | 1.22 | 5.05 | -3.83 |
| Confidence | 2.05 | 7.21 | -5.16 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 2.2% | 21.1% |
| Maintained | 1.1% | 48.8% |
| Lowered | 0.0% | 11.6% |
| Withdrawn | 0.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| The Question Left Hanging | 2.09× | 100.0% | 48.0% |
| Calls That Read Rehearsed | 1.68× | 65.2% | 38.7% |
| Founder-Led Companies | 1.46× | 29.3% | 20.0% |
| Calls That Resolve Doubts | 0.00× | 0.0% | 79.5% |
| Guidance Worth Underwriting | 0.00× | 0.0% | 71.5% |
| Skeptic Reassured | 0.00× | 0.0% | 66.4% |
| The Finished-Story Tell | 0.00× | 0.0% | 4.4% |
| The Hidden Segment | 0.00× | 0.0% | 21.1% |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| CRDOF | Q1 2025 | 2025-05-29 | F |
| AMWD | Q4 2025 | 2025-05-29 | F |
| WRD | Q1 2025 | 2025-05-21 | F |
| XPEV | Q1 2025 | 2025-05-21 | F |
| DDD | Q1 2025 | 2025-05-13 | F |
| BLZE | Q1 2025 | 2025-05-11 | F |
| NOMD | Q1 2025 | 2025-05-10 | F |
| JYNT | Q1 2025 | 2025-05-10 | F |
4Discussion
A careful reader should treat this as a description of a rare conversational profile, not a verdict on any company. The 92 calls are defined by the candor meter itself, so their behavioral gaps partly reflect the scoring method. The rise in frequency from 0.01% in most of 2016 through 2024 to 0.15% in 2024 and 0.93% in 2025 is a real observation in this corpus, but the 2025 figure rests on 6,012 scored calls, a smaller base, and nothing here establishes why the rate moved. No causal claim and no investment implication should be drawn from these statistics.
5Limitations
Behavioral scores and thematic tags are produced by AI reading of transcripts and are noisy; the candor meter is a model output, not ground truth. The returns sample covers 22,449 calls and is skewed toward liquid names, so it does not represent the full universe. Our own forward tests falsified directional prediction from these signals, and LLMs partially remember famous stocks' histories, contaminating any backtest. The 92-call sample is small, and 2025 covers only part of the year. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.