Nine Out of Ten Calls Are 'Confident' Now: Inside the 6+ Confidence Club
This study examines earnings calls scoring 6 or higher on a 0-9 AI-assigned confidence meter. Of 165,182 calls in a 1990-2026 corpus, 157,716 (95.5%) clear the bar, with a 95% interval of 95.4%-95.6%. High-confidence calls skew slightly more specific (7.61 vs 7.56), more promotion-heavy (5.12 vs 5.05), less stressed (2.34 vs 2.43), and less evasive (2.67 vs 2.70) than the base corpus. They raise guidance more often (22.0% vs 21.1%) and lower it less often (10.2% vs 11.6%). Annual share of qualifying calls peaked at 98.08% in 2021 and slid to 90.05% in 2025. Among 22,449 calls with measured outcomes, the qualifying subset's 39.6% beat rate sits within sampling noise of the base's 39.5%.
- 95.5% of 165,182 calls scored 6 or higher on the 0-9 confidence meter (95% CI: 95.4%-95.6%).
- Qualifying calls are less stressed (2.34 vs 2.43), less evasive (2.67 vs 2.70), and more specific (7.61 vs 7.56) than the base corpus.
- They raise guidance 22.0% of the time vs 21.1% for the base, and lower it 10.2% vs 11.6%.
- Among 22,449 calls with post-call outcomes, qualifying calls beat 39.6% of the time vs 39.5% for the base - a negligible gap.
- The share of qualifying calls peaked at 98.08% in 2021 and fell to 90.05% in 2025.
1Introduction
Anyone who reads earnings calls for a living eventually wonders whether the 'confidence' labels their tools assign actually mean anything. If a meter says a call scores 6 or higher on a 0-9 scale, does that call look, sound, or resolve differently from the rest of the corpus? The answer matters for anyone using such scores to triage coverage, because a label that applies to 95.5% of calls is a weak filter on its face. This study examines 157,716 qualifying calls out of 165,182 in a 1990-2026 corpus, comparing their behavioral profiles, guidance actions, and measured outcomes against the full base.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 confidence meter (n = 157,716; 95.5% of the reference set, 95% Wilson interval 95.4%–95.6%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The qualifying group is measurably calmer and more forthcoming than the corpus: stress runs 2.34 vs 2.43, evasion 2.67 vs 2.70, specificity 7.61 vs 7.56, and confidence itself 7.33 vs 7.21. Guidance actions tilt the same way: 22.0% raise versus 21.1% in the base, and 10.2% lower versus 11.6%. But the trend series complicates the picture - qualifying share climbed from 92.59% in 2015 to a 98.08% peak in 2021, then slid to 95.5% in 2024 and 90.05% in 2025. And the outcome check is blunt: a 39.6% beat rate against 39.5% for the base, with median outcomes of -0.070 vs -0.072, shows the label separating tone, not results.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.87 | 6.86 | +0.01 |
| Evasion | 2.67 | 2.70 | -0.03 |
| Specificity | 7.61 | 7.56 | +0.05 |
| Stress | 2.34 | 2.43 | -0.09 |
| Promotion | 5.12 | 5.05 | +0.07 |
| Confidence | 7.33 | 7.21 | +0.12 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 22.0% | 21.1% |
| Maintained | 50.2% | 48.8% |
| Lowered | 10.2% | 11.6% |
| Withdrawn | 2.3% | 2.7% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -7.0% | -7.2% |
| Interquartile range | -25.8% to +11.9% | — |
| Share beating SPY | 39.6% (95% CI 39%–40%) | 39.5% |
| Observations | 21,927 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
4Discussion
A careful reader should conclude that the 6+ threshold mostly identifies calls with a certain tone: less stress, less evasion, more specificity, and somewhat friendlier guidance actions. That is a real, measurable pattern in the text. What the reader should not conclude is that the score flags better outcomes - the beat rate difference of 0.1 percentage points (39.6% vs 39.5%) is well within the reported interval of roughly 39.0%-40.3%, and medians are nearly identical. The score describes how calls sound, not how they turn out. No causal or predictive claim is supported by these comparisons.
5Limitations
The behavioral fields are assigned by AI models and are noisy; deltas of 0.01-0.12 points should be read cautiously. The returns sample covers 22,449 calls and is skewed toward liquid names, limiting generalizability. Our own forward tests of directional prediction using these scores were falsified - the score did not predict outcomes out of sample. Additionally, LLMs partially remember the history of famous stocks, which can contaminate any backtest built on these labels, since the labels may encode knowledge of what actually happened rather than information available at call time. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.