Yes Means Yes: Inside the Earnings Calls Where the Skeptic Blinks
Artul.ai's model rated 165,182 earnings calls from 1990 to 2026 on whether the skeptical analyst voice was ultimately reassured; it answered YES on 109,699 calls, a 66.4% share (95% CI 66.2% to 66.6%). This paper profiles those calls. Reassured calls show higher candor (7.03 vs 6.86), lower evasion (2.44 vs 2.70), and lower stress (1.93 vs 2.43) than the base. Guidance on reassured calls was lowered 6.3% of the time versus 11.6% overall, and raised 28.4% versus 21.1%. Four rhetorical patterns appear disproportionately often, including 'Results Worse Than Direction' with a 0.69 lift. Post-call returns were less bad: median -5.7% versus -7.2%.
- The model answered YES to 'Skeptic Reassured' on 66.4% of 165,182 calls (CI 66.2%-66.6%), covering 109,699 calls.
- Reassured calls score higher on confidence (7.55 vs 7.21, +0.34) and lower on stress (1.93 vs 2.43, -0.49) than the overall corpus.
- Guidance was lowered on only 6.3% of reassured calls versus 11.6% of all calls, and raised on 28.4% versus 21.1%.
- The rhetorical pattern 'Results Worse Than Direction' shows a 0.69 lift on reassured calls, and post-call median returns were -5.7% versus -7.2% for the base (beat rate 40.9% vs 39.5%).
1Introduction
Anyone who listens to earnings calls knows the arc: an analyst presses, management deflects, and eventually the room either buys the answer or does not. Artul.ai's 'Skeptic Reassured' flag captures that moment of resolution, and it resolves in the caller's favor about two-thirds of the time. Whether reassurance is a marker of genuine transparency or simply of management skill is exactly the kind of question a large corpus can at least describe. This study examines the 109,699 calls, out of 165,182 spanning 1990 to 2026, where the model answered YES, profiling their language, guidance behavior, recurring rhetorical patterns, and post-call return distributions.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "Skeptic Reassured", tracked by year (n = 109,699; 66.4% of the reference set, 95% Wilson interval 66.2%–66.6%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Reassured calls read differently: candor runs 7.03 versus 6.86, evasion 2.44 versus 2.70, specificity 7.81 versus 7.56, and confidence 7.55 versus 7.21, while stress drops from 2.43 to 1.93. Guidance skews favorable: lowered on 6.3% of reassured calls against 11.6% overall, raised on 28.4% against 21.1%. Overused rhetoric on reassured calls includes 'The Question Left Hanging' (lift 0.54) and 'Results Worse Than Direction' (0.69). The reassured share peaked at 71.76% in 2020 and 72.96% in 2021 before settling near 64-66% from 2022 onward. Post-call returns were less negative: median -5.7% versus -7.2%, mean -3.6%, with 40.9% beating expectations versus 39.5%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 7.03 | 6.86 | +0.16 |
| Evasion | 2.44 | 2.70 | -0.26 |
| Specificity | 7.81 | 7.56 | +0.25 |
| Stress | 1.93 | 2.43 | -0.49 |
| Promotion | 4.98 | 5.05 | -0.07 |
| Confidence | 7.55 | 7.21 | +0.34 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 28.4% | 21.1% |
| Maintained | 53.9% | 48.8% |
| Lowered | 6.3% | 11.6% |
| Withdrawn | 2.1% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 0.09× | 1.0% | 11.1% |
| The Question Left Hanging | 0.54× | 25.9% | 48.0% |
| Results Worse Than Direction | 0.69× | 35.4% | 51.1% |
| Underused Fixed Costs | 0.74× | 31.0% | 41.6% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -5.7% | -7.2% |
| Interquartile range | -23.3% to +12.3% | — |
| Share beating SPY | 40.9% (95% CI 40%–42%) | 39.5% |
| Observations | 18,030 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| OMF | Q2 2025 | 2025-07-25 | A |
4Discussion
A careful reader should treat these numbers as descriptive. Calls the model rates as reassuring do co-occur with better language scores, friendlier guidance moves, milder post-call returns, and a distinctive set of rhetorical tics. That is association, not proof that reassurance causes anything, and not evidence that spotting reassurance gives anyone an edge in trading. The rhetorical lifts in particular suggest the model's sense of 'reassured' partly tracks management style rather than outcomes. Read this as a map of how calls that end in reassurance tend to sound, and treat every cross-tabulation as a hypothesis for further study.
5Limitations
The underlying fields are AI-generated judgments and carry real noise; model ratings of candor, evasion, and reassurance are not ground truth. The returns analysis covers only 22,449 calls, skewed toward liquid names, and the reassured subset just 18,030, so survivorship and coverage bias are plausible. Our own forward tests falsified directional prediction, and LLMs partially remember famous stocks' histories, contaminating any backtest. The rhetorical lifts describe co-occurrence, not causation, and the 2015 start of the trend series reflects data availability rather than any structural change in earnings calls. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.