Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 62Hypotheses TestedUpdated 2026-08-28

The Question Left Hanging: A Study of Discovery-Tone Earnings Calls

By Artul.ai Research Group · n = 107 earnings calls · First published 2026-08-28
Abstract

This study examines 107 earnings calls (21% of a 499-call corpus spanning 2015 to 2024) whose language was flagged for a distinctive quality: a tone of discovery, where management seems to be working out problems in public rather than reciting prepared answers. Compared with the corpus baseline, these calls show higher confidence (7.68 vs 7.33), higher specificity (7.76 vs 7.65), and lower stress (2.09 vs 2.39). Guidance was raised on 39% of these calls versus 22% in the baseline. Yet subsequent returns were negative: a median of -5.4%, with 47% beating expectations.

Key findings
  • Discovery-tone calls made up 21.4% of the 499-call corpus (107 calls).
  • Confidence scored 7.68 on these calls versus 7.33 in the baseline, the largest profile delta at +0.35.
  • Guidance was raised on 39.3% of discovery-tone calls versus 21.8% of baseline calls.
  • Among 49 discovery-tone calls with measured returns, the median return was -5.4% versus -10.0% for the base sample, with 46.9% beating expectations versus 38.4%.

1Introduction

Most earnings calls are performances: rehearsed scripts, planted questions, carefully sealed answers. Occasionally a call breaks form, and management sounds like it is genuinely discovering its own situation in real time, leaving loose ends dangling. That texture matters to anyone who reads transcripts closely, because it may signal an inflection point, a turnaround being worked through, or a story not yet polished. Artul.ai's field annotations capture this quality as the research hypothesis 'Tone of discovery'. This study examines the 107 calls (out of 499, spanning 2015 to 2024) that answered yes, and what distinguishes their language, guidance behavior, and subsequent outcomes.

2Data & methodology

The corpus comprises 499 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "Tone of discovery" (n = 107; 21.4% of the reference set, 95% Wilson interval 18.1%–25.3%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Discovery-tone calls read as more composed than the average: confidence runs 7.68 versus 7.33, specificity 7.76 versus 7.65, and stress 2.09 versus 2.39, while evasion is slightly lower at 2.61 versus 2.71. Guidance is a clear differentiator: raised guidance appears on 39.3% of these calls versus 21.8% in the baseline, and lowered guidance on just 5.6% versus 13.2%. Overrepresented topics include 'Pricing Recovering' (1.43x lift) and 'Early Products Growing Fast' (1.38x). The topic 'The Question Left Hanging' is underrepresented at 0.68x. Measured returns on 49 calls show a median of -5.4% versus -10.0% for the base sample, with beat rates of 46.9% versus 38.4%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.956.91+0.04
Evasion2.612.71-0.10
Specificity7.767.65+0.11
Stress2.092.39-0.30
Promotion5.175.11+0.06
Confidence7.687.33+0.35
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised39.3%21.8%
Maintained51.4%53.1%
Lowered5.6%13.2%
Withdrawn0.9%1.0%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Pricing Recovering1.43×25.2%17.6%
Early Products Growing Fast1.38×53.3%38.7%
Volume About to Step Up1.27×31.8%25.1%
The Question Left Hanging0.68×29.0%42.9%
20150.00%
20160.10%
20170.06%
20180.08%
20190.01%
20200.00%
20210.09%
20220.08%
20230.11%
20240.08%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-5.4%-10.0%
Interquartile range-26.3% to +10.4%
Share beating SPY46.9% (95% CI 34%–61%)38.4%
Observations49198
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
MCDQ2 20242024-07-29D
SFIXQ3 20242024-06-04C+
CRGOQ1 20242024-05-20C+
EROQ1 20242024-05-10A
HCKTQ1 20242024-05-08C
AZEKQ2 20242024-05-08B+
LINCQ1 20242024-05-06B+
NCQ1 20242024-05-05C+

4Discussion

A careful reader should treat these findings as descriptive. Calls flagged for tone of discovery co-occur with more raised guidance, calmer language, and somewhat better measured outcomes than the baseline, but co-occurrence is not causation, and the differences in returns are modest. Nothing here identifies a trading edge, and the returns sample is small. What the data does support is a portrait: these calls tend to accompany companies in an upswing being narrated honestly, with management thinking aloud rather than dodging.

5Limitations

All fields are AI-read and therefore noisy; tone labels are judgments, not measurements. The returns analysis covers only 49 of 107 flagged calls, against a base of 198, drawn from a larger universe of 22,449 calls skewed toward liquid names, so survivorship and liquidity bias lurk throughout. Artul.ai's own forward tests falsified directional prediction, so no inferential weight should be placed on the return comparisons. Finally, LLMs partially remember famous stocks' histories, contaminating any backtest with hindsight that the models cannot fully bracket. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “The Question Left Hanging: A Study of Discovery-Tone Earnings Calls.” Artul.ai Earnings-Call Research Library, Study No. 62. https://artul.ai/research/hypothesis-tone-of-discovery

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecDon't Stick to the Script: Earnings Calls Where AnswHave Your Cake and Expand the Base Too: Repeat GrowtPressed Harder, Answered Straight: A Profile of 183
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.