Research › Call Signals
Artul.ai Research LibraryStudy No. 14Call SignalsUpdated 2026-08-28

The Question Left Hanging, Fewer and Farther Between: A Study of Calls That Resolve Doubts

By Artul.ai Research Group · n = 131,335 earnings calls · First published 2026-08-28
Abstract

This study examines earnings calls where the model answered YES to the battery item "Calls That Resolve Doubts" across a corpus spanning 1990 to 2026. Of 165,182 calls, 131,335 (79.51%) met the definition. These calls show a notably different behavioral profile: specificity of 7.75 versus a 7.56 baseline, confidence of 7.42 versus 7.21, and lower stress (2.12 vs 2.43) and evasion (2.49 vs 2.70). They are also overrepresented among "The Question Left Hanging" (0.72 observed vs 0.35 expected). Guidance was raised on 24.90% of these calls versus 21.05% in the base. In the returns sample of 19,998 calls, median excess return was -0.064% versus -0.072% baseline.

Key findings
  • 79.51% of the 165,182-call corpus (131,335 calls) were flagged as resolving doubts, with a 95% CI of 79.31% to 79.70%.
  • These calls score higher on specificity (7.75 vs 7.56), confidence (7.42 vs 7.21), and candor (6.99 vs 6.86), and lower on stress (2.12 vs 2.43) and evasion (2.49 vs 2.70).
  • "The Question Left Hanging" is overrepresented at 0.72 observed versus 0.35 expected (lift 2.11x), the largest lift of any flagged item.
  • Guidance was raised on 24.90% of doubt-resolving calls versus 21.05% of the base, and median excess return was -0.064% versus -0.072%.

1Introduction

Anyone who listens to earnings calls knows the feeling of a call that settles the room: questions answered directly, uncertainty visibly reduced. Whether a call resolves doubts is one of the most basic qualities an analyst judges, yet it is rarely measured at scale. Because our corpus spans 1990 to 2026 and covers 165,182 calls, we can ask how common these calls are, what behavioral signatures accompany them, and how their guidance and returns compare to the rest. This study examines calls where the model answered YES to the battery item "Calls That Resolve Doubts" and profiles them against all other calls.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Calls That Resolve Doubts" (n = 131,335; 79.5% of the reference set, 95% Wilson interval 79.3%–79.7%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Doubt-resolving calls differ most sharply on stress and evasion, both lower than baseline (2.12 vs 2.43 and 2.49 vs 2.70), with specificity and confidence higher (7.75 vs 7.56; 7.42 vs 7.21). The overrepresentation table is striking: "The Question Left Hanging" appears at 0.72 versus 0.35 expected, suggesting calls can resolve most doubts while still leaving one open. Guidance actions lean positive (24.90% raised vs 21.05%; 8.93% lowered vs 11.56%). Returns show essentially no difference: median excess return of -0.064% versus -0.072%, with 40.31% beating versus 39.47% in the base. The trend series peaks at 83.34 in 2020 and falls to 74.25 in 2025.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.996.86+0.13
Evasion2.492.70-0.21
Specificity7.757.56+0.19
Stress2.122.43-0.31
Promotion5.015.05-0.04
Confidence7.427.21+0.21
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised24.9%21.1%
Maintained52.3%48.8%
Lowered8.9%11.6%
Withdrawn2.2%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims0.33×3.6%11.1%
The Question Left Hanging0.72×34.7%48.0%
201576.43%
201677.93%
201779.34%
201879.57%
201978.55%
202083.34%
202183.05%
202278.05%
202379.20%
202478.81%
202574.25%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-6.4%-7.2%
Interquartile range-24.4% to +12.2%
Share beating SPY40.3% (95% CI 40%–41%)39.5%
Observations19,99822,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
SBFGQ2 20252025-07-25A
DOCQ2 20252025-07-25C
USCBQ2 20252025-07-25B+
AONQ2 20252025-07-25C
NWGQ2 20252025-07-25B+
BFHQ2 20252025-07-25B
FFICQ2 20252025-07-25B+
OMFQ2 20252025-07-25A

4Discussion

A careful reader should conclude that calls resolving doubts are common (79.51% of the corpus) and carry a coherent behavioral fingerprint: more specific, more confident, less stressed, less evasive. They also leave more hanging questions on average, which is a curious but descriptive pattern. What one should not conclude is that resolving doubts causes better outcomes, that the return differences of -0.064% versus -0.072% represent any exploitable signal, or that the 2020-2021 peaks (83.34 and 83.05) tell us anything about future call behavior. All findings here are descriptive.

5Limitations

The battery items and profile scores are AI-read fields and are inherently noisy; the model may misjudge candor, stress, or evasion on any given call. The returns sample covers only 22,449 calls overall (19,998 here) and is skewed toward liquid names, so results may not generalize. Our own forward tests falsified directional prediction, so nothing here should be read as an edge. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of model judgments against realized outcomes. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “The Question Left Hanging, Fewer and Farther Between: A Study of Calls That Resolve Doubts.” Artul.ai Earnings-Call Research Library, Study No. 14. https://artul.ai/research/calls-that-resolve-doubts-earnings-calls

Related studies

Ammunition, Not Smoke: Inside the Calls Where CriticKnow What You Know: Calls Where Confidence Matched tThe Guidance Was Fine All Along: What Model-EndorsedTalking the Skeptic Down: 66% of Calls Leave the DouThe Admission Is Not the Confession: Calls Flagged aThe Question Left Hanging: 79,206 Calls That End on
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.