Research › Call Signals
Artul.ai Research LibraryStudy No. 78Call SignalsUpdated 2026-08-28

The Admission Is Not the Confession: Calls Flagged as 'Results Worse Than Direction'

By Artul.ai Research Group · n = 84,487 earnings calls · First published 2026-08-28
Abstract

This study examines 84,487 earnings calls (out of 165,182, or 51.15%) where the model answered YES to the battery item 'Results Worse Than Direction', across a 1990–2026 corpus. Flagged calls show lower candor (6.86 vs 6.86), higher stress (2.94 vs 2.43) and evasion (2.85 vs 2.70), and lower confidence (6.85 vs 7.21). Guidance behavior diverges sharply: 17.64% lowered guidance versus 11.56% in the base, and 10.74% raised it versus 21.05%. In the returns subsample (n = 9,167), the median return was -0.098 versus -0.072 in the base, with 37.08% beating versus 39.47%. Flagged calls over-index on 'Scale-Dependent Advantage Claims' (lift 1.70) and under-index on 'Skeptic Reassured' (0.69).

Key findings
  • 51.15% of 165,182 calls (84,487) were flagged YES on 'Results Worse Than Direction', with a 95% share interval of 50.91% to 51.39%.
  • Flagged calls show higher stress (2.94 vs 2.43) and evasion (2.85 vs 2.70), but nearly identical candor (6.86 vs 6.86).
  • Guidance diverges: 17.64% of flagged calls lowered guidance vs 11.56% of the base, while only 10.74% raised it vs 21.05%.
  • In the returns subsample, flagged calls had a median return of -0.098 vs -0.072 for the base and a beat rate of 37.08% (CI 36.10%–38.07%) vs 39.47%.

1Introduction

Earnings calls rarely announce bad news outright; the damage usually shows up in tone, hedging, and what management chooses to lower rather than raise. A model flag of 'Results Worse Than Direction' is one attempt to capture that moment systematically: the call where the numbers or the narrative read as worse than the trajectory the company had set. For anyone who reads calls closely, the question is whether such a flag tracks measurable differences in language, guidance behavior, and market reaction, or whether it mostly restates the obvious. This study compares 84,487 flagged calls against the 165,182-call corpus to describe those differences.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Results Worse Than Direction" (n = 84,487; 51.1% of the reference set, 95% Wilson interval 50.9%–51.4%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Flagged calls are not more candid than the base (6.86 vs 6.86) but read as more strained: stress runs 2.94 vs 2.43, evasion 2.85 vs 2.70, and confidence drops from 7.21 to 6.85. Guidance tells a sharper story: 17.64% of flagged calls lowered guidance versus 11.56% of the base, and only 10.74% raised it versus 21.05%. Rhetorically, flagged calls over-index on 'Scale-Dependent Advantage Claims' (lift 1.70), 'The Hidden Segment' (1.56), and 'Underused Fixed Costs' (1.43), while 'Skeptic Reassured' appears at 0.69 lift. In the returns subsample, the median flagged-call return was -0.098 versus -0.072, with a beat rate of 37.08% against 39.47%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.866.86-0.00
Evasion2.852.70+0.15
Specificity7.427.56-0.14
Stress2.942.43+0.52
Promotion5.045.05-0.01
Confidence6.857.21-0.37
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised10.7%21.1%
Maintained47.2%48.8%
Lowered17.6%11.6%
Withdrawn3.9%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims1.70×18.9%11.1%
The Hidden Segment1.56×32.9%21.1%
Underused Fixed Costs1.43×59.8%41.6%
The Question Left Hanging1.29×61.7%48.0%
Skeptic Reassured0.69×45.9%66.4%
201554.49%
201653.77%
201751.13%
201848.53%
201950.76%
202058.45%
202143.08%
202249.89%
202354.71%
202452.21%
202544.79%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-9.8%-7.2%
Interquartile range-30.1% to +10.5%
Share beating SPY37.1% (95% CI 36%–38%)39.5%
Observations9,16722,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
DOCQ2 20252025-07-25C
CNCQ2 20252025-07-25F
GBCIQ2 20252025-07-25A
UVEQ2 20252025-07-25C+
VRTSQ2 20252025-07-25C+
FLGQ2 20252025-07-25B
FRSTQ2 20252025-07-25A
HMDPFQ2 20252025-07-25B

4Discussion

A careful reader should treat these as descriptive co-movements, not causes or signals. Flagged calls do coincide with more guidance cuts, more stressed language, and somewhat weaker subsequent returns, but none of these gaps establishes that the flag predicts anything or that tone drove the outcomes. The returns difference (-0.098 vs -0.072 median) is modest and measured on a skewed subsample. The safest conclusion is that the flag marks a recognizable cluster of language and guidance behavior, and nothing more directional than that.

5Limitations

The underlying fields are AI-read and inherently noisy, so tone deltas of a few tenths should not be over-read. The returns analysis covers only 22,449 calls with outcomes, skewed toward liquid names, and our own forward tests falsified directional prediction from these features. LLMs also partially remember famous stocks' histories, which can contaminate any backtest by leaking hindsight into scores. All figures here are descriptive of the corpus and carry no causal or predictive claim. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “The Admission Is Not the Confession: Calls Flagged as 'Results Worse Than Direction'.” Artul.ai Earnings-Call Research Library, Study No. 78. https://artul.ai/research/results-worse-than-direction-earnings-calls

Related studies

Ammunition, Not Smoke: Inside the Calls Where CriticKnow What You Know: Calls Where Confidence Matched tThe Question Left Hanging, Fewer and Farther BetweenThe Guidance Was Fine All Along: What Model-EndorsedTalking the Skeptic Down: 66% of Calls Leave the DouThe Question Left Hanging: 79,206 Calls That End on
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.