Research › Signal Combinations
Artul.ai Research LibraryStudy No. 23Signal CombinationsUpdated 2026-08-28

Dodging the Question, Formally: The Anatomy of an Unanswered Earnings Call

By Artul.ai Research Group · n = 3,629 earnings calls · First published 2026-08-28
Abstract

We examine 3,629 earnings calls (2.20% of a 165,182-call corpus spanning 1990-2026) where the model answered YES to "The Question Left Hanging" and evasion scored 6/9 or higher. These calls are behaviorally distinctive: evasion runs 6.23 vs. 2.70 baseline, while candor (5.75 vs. 6.86), specificity (6.25 vs. 7.56), and confidence (6.47 vs. 7.21) all fall short. Management guidance tells a matching story: 7.99% of these calls withdrew guidance versus 2.66% of the baseline. Subsequent one-quarter returns skew worse, with a median of -0.13 against -0.07 baseline, and only 31.9% beating expectations versus 39.5%. The share of such calls has also fallen steadily, from 3.93 in 2015 to 1.36 in 2025.

Key findings
  • Evasion scores average 6.23 on these calls versus 2.70 baseline, while candor (5.75 vs. 6.86) and specificity (6.25 vs. 7.56) sit well below typical.
  • 7.99% of flagged calls withdrew guidance compared with 2.66% of all calls, and only 8.49% raised it versus 21.05%.
  • Subsequent returns have a median of -0.13 versus -0.07 for the 22,449-call baseline, with 31.9% beating expectations (CI 27.3%-36.9%) against 39.5% overall.
  • The flagged share of calls declined from 3.93 in 2015 to 1.36 in 2025, with the strongest top-tag lift being "Scale-Dependent Advantage Claims" at 2.65x.

1Introduction

An earnings call transcript can be fluent, confident, and still never answer the question that was asked. Analysts learn to listen for that gap, but reading thousands of calls by hand is impractical, and the signals are easy to talk yourself out of: was the silence evasion, or just discipline? This study isolates the extreme end of that behavior: calls where our model explicitly registered a question left hanging and where measured evasion reached 6 or higher on a 9-point scale. It profiles how those calls differ in tone, guidance behavior, analyst reception, and subsequent outcomes, and asks whether the pattern is becoming more or less common over time.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "The Question Left Hanging" AND evasion scored 6/9 or higher (n = 3,629; 2.2% of the reference set, 95% Wilson interval 2.1%–2.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The behavioral profile is coherent: evasion averages 6.23 against a 2.70 baseline, while candor (5.75 vs. 6.86), specificity (6.25 vs. 7.56), and confidence (6.47 vs. 7.21) all run below typical, and stress is elevated at 4.13 vs. 2.43. Guidance behavior mirrors the tone: 7.99% of these calls withdrew guidance versus 2.66% of the corpus, and 11.93% lowered it versus 11.56%. Tag-level lifts are modest, led by "Scale-Dependent Advantage Claims" at 2.65x, while "Calls That Resolve Doubts" and "Skeptic Reassured" appear far less often than expected (0.25x and 0.26x). In the returns sample, the median subsequent return is -0.13 versus -0.07 baseline, and 31.9% beat expectations against 39.5%. The annual share has fallen from 3.93 in 2015 to 1.36 in 2025.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor5.756.86-1.11
Evasion6.232.70+3.53
Specificity6.257.56-1.31
Stress4.132.43+1.70
Promotion5.185.05+0.13
Confidence6.477.21-0.74
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised8.5%21.1%
Maintained38.6%48.8%
Lowered11.9%11.6%
Withdrawn8.0%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims2.65×29.4%11.1%
Calls That Read Rehearsed1.36×52.6%38.7%
Results Worse Than Direction1.32×67.6%51.1%
The Finished-Story Tell1.28×5.6%4.4%
Underused Fixed Costs1.27×52.8%41.6%
Calls That Resolve Doubts0.25×19.9%79.5%
Skeptic Reassured0.26×17.0%66.4%
Guidance Worth Underwriting0.41×29.1%71.5%
Confidence Proportionate to Evidence0.63×55.4%87.6%
Pricing Recovering0.66×14.1%21.5%
20153.93%
20163.01%
20172.85%
20182.64%
20192.54%
20202.20%
20211.81%
20221.70%
20231.58%
20241.53%
20251.36%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-13.3%-7.2%
Interquartile range-32.5% to +11.7%
Share beating SPY31.9% (95% CI 27%–37%)39.5%
Observations36022,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
HYMTFQ2 20252025-07-24F
UNPQ2 20252025-07-24D
WFGQ2 20252025-07-24F
NESRFQ4 20252025-07-24F
TSLAQ2 20252025-07-23F
THCQ2 20252025-07-23C+
VICRQ2 20252025-07-22F
SARTFQ2 20252025-07-22D

4Discussion

A careful reader should conclude that calls flagged for evasion and an unanswered question are measurably different in tone, guidance behavior, and analyst reception, and that reception is on the whole less favorable. What follows is association, not cause: nothing here shows that evasion itself produces weaker outcomes, and the overlapping returns CIs caution against over-reading the gap. The declining annual share is a descriptive trend, not a forecast. Treat this as a descriptive anatomy of a communication style, useful for close reading of individual calls, not as a mechanical screen or a signal to trade on.

5Limitations

The fields underlying these scores are produced by AI readers and are noisy; individual calls can be mis-scored even when aggregate patterns hold. The returns comparison covers only 360 flagged calls against a 22,449-call baseline skewed toward liquid names, so the sample is not representative of the full corpus. Our own forward tests on similar constructs falsified directional prediction, and no result here should be read as an edge. Finally, LLMs partially remember the published history of famous stocks, which can contaminate any backtest of model-scored transcripts. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Dodging the Question, Formally: The Anatomy of an Unanswered Earnings Call.” Artul.ai Earnings-Call Research Library, Study No. 23. https://artul.ai/research/dodged-and-left-hanging

Related studies

Small Pond, Big Fish: Calls Claiming Fast Growth in Trust the Raise: When Skeptics Get Reassured and GuiThe Founder Is Present: Candor and Confidence on EarLoud Backlogs and Busy Phones: Calls Where the ModelFounder Knows Best: High-Promotion Founder-Led EarniThe Quiet Hand on the Microphone: CFO Dominance Meet
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.