Promoted to the Front Page: High-Score Calls Across 165,182 Earnings Transcripts
This study asks what distinguishes earnings calls that score 6 or higher on Artul.ai's 0-9 promotion meter, a measure of upbeat language. Of 165,182 calls spanning 1990-2026, 61,671 (37.3%) meet the threshold. High-scoring calls show higher confidence (7.68 vs 7.21) and higher promotion language by construction, alongside slightly lower candor (6.57 vs 6.86) and higher evasion (2.86 vs 2.70). Guidance tells a sharper story: 26.1% of high-scoring calls raised guidance versus 21.1% of others, while only 7.1% lowered it versus 11.6%. Yet forward returns were slightly worse: median -9.3% versus -7.2%, with 38.2% beating versus 39.5%.
- 37.3% of 165,182 calls (61,671) score 6 or higher on the 0-9 promotion meter.
- High-scoring calls show higher confidence (7.68 vs 7.21) but lower candor (6.57 vs 6.86) and higher evasion (2.86 vs 2.70).
- 26.1% of high-scoring calls raised guidance versus 21.1% of other calls, while 7.1% lowered it versus 11.6%.
- Among 7,706 high-scoring calls with return data, the median forward return was -9.3%, versus -7.2% for the 22,449-call base.
1Introduction
Anyone who follows earnings calls knows the tone: the confident CEO, the soaring superlatives, the carefully optimistic outlook. Promotional language is the most visible signal on a call, and the easiest to imitate, which makes it worth measuring carefully. If upbeat calls reliably accompanied better outcomes, tone would be a shortcut; if they do not, tone is mostly theater. Using Artul.ai's 0-9 promotion meter, we identify calls scoring 6 or higher and compare their language, guidance behavior, and forward returns against the rest of the 165,182-call corpus spanning 1990 to 2026. This paper reports what we find.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 promotion meter (n = 61,671; 37.3% of the reference set, 95% Wilson interval 37.1%–37.6%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
High-scoring calls are more confident (7.68 vs 7.21) and more evasive (2.86 vs 2.70), with slightly lower candor (6.57 vs 6.86) and specificity (7.47 vs 7.56). Guidance skews positive: 26.1% raised versus 21.1% for the base, and 7.1% lowered versus 11.6%. The most overrepresented phrases are 'Scale-Dependent Advantage Claims' (1.8x lift) and 'A Tiny Fraction of the Market' (1.66x); 'When the CFO Dominates' is underrepresented at 0.53x. Annually, high-scoring shares hovered near 33-35% through 2020, jumped to 48.4% in 2021, then eased back to 39.2% in 2024. Forward returns ran slightly below the base: median -9.3% vs -7.2%, with 38.2% beating versus 39.5%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.57 | 6.86 | -0.30 |
| Evasion | 2.86 | 2.70 | +0.16 |
| Specificity | 7.47 | 7.56 | -0.09 |
| Stress | 2.35 | 2.43 | -0.08 |
| Promotion | 6.44 | 5.05 | +1.38 |
| Confidence | 7.68 | 7.21 | +0.47 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 26.1% | 21.1% |
| Maintained | 45.1% | 48.8% |
| Lowered | 7.1% | 11.6% |
| Withdrawn | 1.7% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 1.80× | 20.0% | 11.1% |
| A Tiny Fraction of the Market | 1.66× | 49.9% | 30.0% |
| Founder-Led Companies | 1.51× | 30.3% | 20.0% |
| Early Products Growing Fast | 1.45× | 55.9% | 38.5% |
| Deferred Revenue Growing | 1.45× | 12.8% | 8.9% |
| When the CFO Dominates | 0.53× | 7.5% | 14.1% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -9.3% | -7.2% |
| Interquartile range | -30.9% to +12.9% | — |
| Share beating SPY | 38.2% (95% CI 37%–39%) | 39.5% |
| Observations | 7,706 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| AON | Q2 2025 | 2025-07-25 | C |
| FLG | Q2 2025 | 2025-07-25 | B |
| FRST | Q2 2025 | 2025-07-25 | A |
| CHTR | Q2 2025 | 2025-07-25 | C+ |
| BAH | Q1 2026 | 2025-07-25 | C+ |
| ASPS | Q2 2025 | 2025-07-25 | D |
| AAFRF | Q1 2026 | 2025-07-25 | B |
4Discussion
The pattern a careful reader should take away is a mismatch, not a signal: calls heavy on promotion talk more about raising guidance and sound more confident, yet their subsequent returns are, if anything, slightly weaker than the base. Promotion language coexists with lower candor and higher evasion in this data. What should not be concluded is that upbeat tone causes poor returns or predicts them; these are descriptive averages across thousands of calls, and the differences are modest. Correlation in tone data is not a trading rule, and the guidance association does not establish that promotional calls are dishonest.
5Limitations
The promotion meter and related scores are AI-read fields and inherently noisy, so small deltas should not be overread. The returns comparison rests on 7,706 high-scoring calls within a 22,449-call sample skewed toward liquid names, limiting representativeness. Our own forward tests of call-based signals falsified directional prediction, so nothing here should be treated as an edge. Finally, large language models partially remember famous stocks' histories, which can contaminate any backtest of language scores against subsequent returns. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.