The Question Left Hanging: A Study of Discovery-Tone Earnings Calls
This study examines 107 earnings calls (21% of a 499-call corpus spanning 2015 to 2024) whose language was flagged for a distinctive quality: a tone of discovery, where management seems to be working out problems in public rather than reciting prepared answers. Compared with the corpus baseline, these calls show higher confidence (7.68 vs 7.33), higher specificity (7.76 vs 7.65), and lower stress (2.09 vs 2.39). Guidance was raised on 39% of these calls versus 22% in the baseline. Yet subsequent returns were negative: a median of -5.4%, with 47% beating expectations.
- Discovery-tone calls made up 21.4% of the 499-call corpus (107 calls).
- Confidence scored 7.68 on these calls versus 7.33 in the baseline, the largest profile delta at +0.35.
- Guidance was raised on 39.3% of discovery-tone calls versus 21.8% of baseline calls.
- Among 49 discovery-tone calls with measured returns, the median return was -5.4% versus -10.0% for the base sample, with 46.9% beating expectations versus 38.4%.
1Introduction
Most earnings calls are performances: rehearsed scripts, planted questions, carefully sealed answers. Occasionally a call breaks form, and management sounds like it is genuinely discovering its own situation in real time, leaving loose ends dangling. That texture matters to anyone who reads transcripts closely, because it may signal an inflection point, a turnaround being worked through, or a story not yet polished. Artul.ai's field annotations capture this quality as the research hypothesis 'Tone of discovery'. This study examines the 107 calls (out of 499, spanning 2015 to 2024) that answered yes, and what distinguishes their language, guidance behavior, and subsequent outcomes.
2Data & methodology
The corpus comprises 499 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "Tone of discovery" (n = 107; 21.4% of the reference set, 95% Wilson interval 18.1%–25.3%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Discovery-tone calls read as more composed than the average: confidence runs 7.68 versus 7.33, specificity 7.76 versus 7.65, and stress 2.09 versus 2.39, while evasion is slightly lower at 2.61 versus 2.71. Guidance is a clear differentiator: raised guidance appears on 39.3% of these calls versus 21.8% in the baseline, and lowered guidance on just 5.6% versus 13.2%. Overrepresented topics include 'Pricing Recovering' (1.43x lift) and 'Early Products Growing Fast' (1.38x). The topic 'The Question Left Hanging' is underrepresented at 0.68x. Measured returns on 49 calls show a median of -5.4% versus -10.0% for the base sample, with beat rates of 46.9% versus 38.4%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.95 | 6.91 | +0.04 |
| Evasion | 2.61 | 2.71 | -0.10 |
| Specificity | 7.76 | 7.65 | +0.11 |
| Stress | 2.09 | 2.39 | -0.30 |
| Promotion | 5.17 | 5.11 | +0.06 |
| Confidence | 7.68 | 7.33 | +0.35 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 39.3% | 21.8% |
| Maintained | 51.4% | 53.1% |
| Lowered | 5.6% | 13.2% |
| Withdrawn | 0.9% | 1.0% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Pricing Recovering | 1.43× | 25.2% | 17.6% |
| Early Products Growing Fast | 1.38× | 53.3% | 38.7% |
| Volume About to Step Up | 1.27× | 31.8% | 25.1% |
| The Question Left Hanging | 0.68× | 29.0% | 42.9% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -5.4% | -10.0% |
| Interquartile range | -26.3% to +10.4% | — |
| Share beating SPY | 46.9% (95% CI 34%–61%) | 38.4% |
| Observations | 49 | 198 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| MCD | Q2 2024 | 2024-07-29 | D |
| SFIX | Q3 2024 | 2024-06-04 | C+ |
| CRGO | Q1 2024 | 2024-05-20 | C+ |
| ERO | Q1 2024 | 2024-05-10 | A |
| HCKT | Q1 2024 | 2024-05-08 | C |
| AZEK | Q2 2024 | 2024-05-08 | B+ |
| LINC | Q1 2024 | 2024-05-06 | B+ |
| NC | Q1 2024 | 2024-05-05 | C+ |
4Discussion
A careful reader should treat these findings as descriptive. Calls flagged for tone of discovery co-occur with more raised guidance, calmer language, and somewhat better measured outcomes than the baseline, but co-occurrence is not causation, and the differences in returns are modest. Nothing here identifies a trading edge, and the returns sample is small. What the data does support is a portrait: these calls tend to accompany companies in an upswing being narrated honestly, with management thinking aloud rather than dodging.
5Limitations
All fields are AI-read and therefore noisy; tone labels are judgments, not measurements. The returns analysis covers only 49 of 107 flagged calls, against a base of 198, drawn from a larger universe of 22,449 calls skewed toward liquid names, so survivorship and liquidity bias lurk throughout. Artul.ai's own forward tests falsified directional prediction, so no inferential weight should be placed on the return comparisons. Finally, LLMs partially remember famous stocks' histories, contaminating any backtest with hindsight that the models cannot fully bracket. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.
Companion page: every company matching this hypothesis is listed at the question’s own page.