The Long Way to Say Less: Inside High-Complexity Earnings Calls
This study examines earnings calls scoring 6 or higher on Artul.ai's 0-9 complexity meter, a group spanning 64,523 of 165,182 calls (39.1%) from 1990 to 2026. High-complexity calls sound more evasive (evasion 3.04 vs. 2.70) and more stressed (2.75 vs. 2.43) than the corpus baseline, while showing lower candor (6.75 vs. 6.86), specificity (7.47 vs. 7.56), and confidence (7.07 vs. 7.21). Guidance posture skews negative: 16.2% of these calls raised guidance versus 21.1% of all calls. Their excess returns skew more negative than baseline (median -0.0894 vs. -0.0716), and only 37.0% beat, versus 39.5% overall. The complexity share of calls has also fallen steadily, from 44.6% in 2015 to 32.8% in 2025, suggesting disclosure styles have shifted toward simpler talk.
- High-complexity calls represent 39.1% of the corpus (64,523 of 165,182 calls), with a 95% interval of 38.8% to 39.3%.
- Compared with all calls, these show higher evasion (3.04 vs. 2.70) and stress (2.75 vs. 2.43) but lower candor (6.75 vs. 6.86) and confidence (7.07 vs. 7.21).
- Guidance was raised on 16.2% of high-complexity calls versus 21.1% of all calls, while lowered guidance appeared at 12.7% versus 11.6%.
- The annual complexity share declined from 44.6% of calls in 2015 to 32.8% in 2025.
1Introduction
Anyone who listens to earnings calls for a living develops a feel for the difference between a complicated business story and a deliberately tangled one. Complexity in management language is one of the loudest quiet signals in the transcript: it can flag stress, hedging, or genuine strategic nuance, and readers routinely over- or under-weight it. Because transcripts now number in the hundreds of thousands, we can measure how the wordiest, most knotted calls actually differ from the average. This study examines the 64,523 calls scoring 6 or higher on a 0-9 complexity meter, profiling their language, guidance behavior, and market-reaction statistics against the full corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 complexity meter (n = 64,523; 39.1% of the reference set, 95% Wilson interval 38.8%–39.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile paints a consistent picture: high-complexity calls run higher on evasion (3.04 vs. 2.70) and stress (2.75 vs. 2.43), and lower on candor (6.75 vs. 6.86), specificity (7.47 vs. 7.56), and confidence (7.07 vs. 7.21). Guidance skews cautious: 16.2% raised versus 21.1% corpus-wide, and 12.7% lowered versus 11.6%. Only 37.0% of these calls beat, against 39.5% of all calls, with a median excess return of -0.0894 versus -0.0716. Two topics appear more often than expected: Scale-Dependent Advantage Claims (1.58x observed lift) and The Question Left Hanging (1.31x). Meanwhile, the complexity share fell from 44.6% in 2015 to 32.8% in 2025, a broad stylistic shift toward simpler talk.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.75 | 6.86 | -0.11 |
| Evasion | 3.04 | 2.70 | +0.34 |
| Specificity | 7.47 | 7.56 | -0.08 |
| Stress | 2.75 | 2.43 | +0.32 |
| Promotion | 5.10 | 5.05 | +0.05 |
| Confidence | 7.07 | 7.21 | -0.14 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 16.2% | 21.1% |
| Maintained | 49.7% | 48.8% |
| Lowered | 12.7% | 11.6% |
| Withdrawn | 2.3% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 1.58× | 17.5% | 11.1% |
| The Question Left Hanging | 1.31× | 62.8% | 48.0% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -8.9% | -7.2% |
| Interquartile range | -27.9% to +10.1% | — |
| Share beating SPY | 37.0% (95% CI 36%–38%) | 39.5% |
| Observations | 8,659 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| DOC | Q2 2025 | 2025-07-25 | C |
| HCA | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FLG | Q2 2025 | 2025-07-25 | B |
| TNET | Q2 2025 | 2025-07-25 | C+ |
| DBOEY | Q2 2025 | 2025-07-25 | B+ |
4Discussion
A careful reader should conclude that high-complexity language co-occurs with more evasive, more stressed delivery and somewhat weaker outcomes - not that complexity causes any of it. The differences are descriptive averages over decades; individual calls vary widely, and the beat-rate gap of roughly two and a half points is modest. Nothing here identifies a trading edge, and the trend toward simpler calls may reflect changes in disclosure norms, coaching, or how transcripts are produced rather than a change in corporate health.
5Limitations
Complexity, candor, and evasion scores come from AI-read fields and are noisy measurements of fuzzy constructs. The returns comparison rests on 22,449 calls with available returns, a sample skewed toward liquid names, so small-cap behavior is underrepresented. Our own forward tests falsified directional prediction: the beat-rate differences do not support an implementable strategy. Finally, language models partially remember the history of famous stocks, which can contaminate any backtest built on these scores, and the observed lifts describe association only. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.