The Truth Won't Set You Free: High Candor Is the Norm on Earnings Calls
Does calling a company 'candid' tell you anything? Artul.ai scored 165,182 earnings-call transcripts from 1990 through 2026 on a 0–9 candor meter and studied the 158,979 calls scoring 6 or higher. The first surprise is how unexclusive the club is: 96.2% of calls qualify (confidence interval 96.2–96.3). High-candor calls barely differ from the rest — specificity runs 7.62 versus 7.56, evasion 2.66 versus 2.70 — and guidance behavior is nearly identical (21.4% raise guidance versus 21.1% baseline; 2.7% withdraw it in both groups). Post-call returns show no separation either: a 39.55% beat rate versus 39.47%, with medians of -0.0707 versus -0.0716. The one rupture is temporal: the qualifying share held between 96.06% and 97.5% from 2015 through 2023, slipped to 95.49% in 2024, and fell to 89.45% across 2025's 6,012 calls.
- 96.2% of the 165,182 scored calls clear the 6-or-higher candor bar — 158,979 transcripts, with a confidence interval of 96.2 to 96.3.
- High-candor calls look like everyone else: specificity 7.62 versus 7.56, evasion 2.66 versus 2.70, stress 2.40 versus 2.43, and confidence 7.22 versus 7.21.
- The qualifying share held between 96.06% and 97.5% from 2015 through 2023, slipped to 95.49% in 2024, and fell to 89.45% across 2025's 6,012 calls.
- Returns show no separation: a 39.55% beat rate versus 39.47% for the 22,449-call baseline, and a median of -0.0707 versus -0.0716.
1Introduction
Analysts, journalists, and screening tools increasingly treat candor scores as a quality signal: read the transcript, score the straight talk, weight the name accordingly. That workflow presumes high-candor calls form a select group worth isolating. The numbers say otherwise. On a 0–9 candor meter applied to earnings calls from 1990 through 2026, a score of 6 or higher — the intuitive 'candid' tier — captures 96.2% of 165,182 transcripts. A filter that inclusive removes almost nobody, and the sliver it does exclude may be the more interesting population. This study profiles the 158,979 qualifying calls: their evasion, specificity, stress, promotion, and confidence scores; how often they raise, maintain, lower, or withdraw guidance; how the share moved from 2015 to 2025; and what their post-call returns look like against the full sample.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 candor meter (n = 158,979; 96.2% of the reference set, 95% Wilson interval 96.2%–96.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The gap between high-candor calls and the full corpus is measured in hundredths: candor 6.95 versus 6.86, specificity 7.62 versus 7.56, evasion 2.66 versus 2.70, stress 2.40 versus 2.43, promotion 5.01 versus 5.05, confidence 7.22 versus 7.21. Guidance behavior matches just as closely — 21.4% raise guidance versus 21.1% of all calls, 49.4% versus 48.8% maintain it, 11.9% versus 11.6% lower it, and 2.7% versus 2.7% withdraw it. The trend table holds the only drama: after years pinned between 96.06% and 97.5%, the qualifying share slipped to 95.49% in 2024 and dropped to 89.45% across 2025's 6,012 calls. Returns agree with the flat profile: 22,192 high-candor calls with returns post a 39.55% beat rate versus 39.47% for all 22,449 (confidence band 38.91–40.20), medians of -0.0707 versus -0.0716, and a mean of -0.0503.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.95 | 6.86 | +0.09 |
| Evasion | 2.66 | 2.70 | -0.04 |
| Specificity | 7.62 | 7.56 | +0.07 |
| Stress | 2.40 | 2.43 | -0.03 |
| Promotion | 5.01 | 5.05 | -0.04 |
| Confidence | 7.22 | 7.21 | +0.01 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 21.4% | 21.1% |
| Maintained | 49.4% | 48.8% |
| Lowered | 11.9% | 11.6% |
| Withdrawn | 2.7% | 2.7% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -7.1% | -7.2% |
| Interquartile range | -25.9% to +11.9% | — |
| Share beating SPY | 39.6% (95% CI 39%–40%) | 39.5% |
| Observations | 22,192 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
4Discussion
The safe read is descriptive. A 6-or-higher candor score is so common that it functions as a default label, not a discriminator; its near-identical guidance mix, profile deltas, and beat rates are consistent with a threshold that separates very little in this corpus. The 2025 figure deserves its own caution: 6,012 calls is a thin slice, and the 89.45% share may reflect an incomplete year or a shifting scoring environment rather than a behavioral collapse. Readers should not conclude that candor is meaningless, that scoring high causes better or worse outcomes, or that any of these gaps points to a tradeable pattern — none of that is tested here.
5Limitations
Every score here is machine-assigned from transcript text, and AI-read fields are noisy at the call level. The returns comparison covers 22,449 calls, a subset skewed toward liquid names, so it may not speak for the full 165,182-call corpus. Artul.ai's own forward tests falsified directional prediction from these signals. LLMs also partially remember famous stocks' histories, which can contaminate any backtest with training-data leakage. Treat these numbers as measurement, not advice. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.