Research › Call Signals
Artul.ai Research LibraryStudy No. 30Call SignalsUpdated 2026-08-28

The Guidance Was Fine All Along: What Model-Endorsed Calls Look Like

By Artul.ai Research Group · n = 118,164 earnings calls · First published 2026-08-28
Abstract

We examine 165,182 earnings-call transcripts from 1990 through 2026 and isolate the 118,164 calls (71.5%, 95% CI 71.3% to 71.8%) where the model answered YES to the battery item 'Guidance Worth Underwriting'. These calls skew more candid (7.0 vs 6.86), more specific (7.77 vs 7.56), and more confident (7.42 vs 7.21), while showing less stress (2.1 vs 2.43) and less evasion (2.5 vs 2.7). Guidance actions differ too: 26.8% of these calls raised guidance versus 21.1% in the base set. Post-call returns medians were -0.063 versus -0.072, and 40.5% beat versus 39.5% in the base sample.

Key findings
  • 71.5% of 165,182 calls (CI 71.3%-71.8%) received a YES on 'Guidance Worth Underwriting' (n=118,164).
  • YES calls show higher specificity (7.77 vs 7.56) and confidence (7.42 vs 7.21), and lower stress (2.1 vs 2.43) and evasion (2.5 vs 2.7).
  • 26.8% of YES calls raised guidance versus 21.1% of base calls, while only 0.31% withdrew guidance versus 2.66% in the base set.
  • In the matched returns sample (n=19,751), the YES-group median post-call return was -0.063 versus -0.072 for the base sample, with a beat rate of 40.5% versus 39.5%.

1Introduction

Anyone who follows earnings calls knows the ritual: management walks through the quarter, then offers a forward view the Street must decide whether to trust. The battery item 'Guidance Worth Underwriting' asks a simple question of each call - does the forward guidance hold up as something an analyst could actually lean on? Because the answer is available for 165,182 calls spanning 1990 through 2026, it offers a way to characterize what a credible-guidance call sounds like in language, what actions follow it, and how those calls behave afterward. This study describes the 118,164 YES calls and compares them with the full corpus on language profile, guidance actions, discourse markers, annual trends, and post-call returns.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Guidance Worth Underwriting" (n = 118,164; 71.5% of the reference set, 95% Wilson interval 71.3%–71.8%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The YES group reads differently before any outcome is known: candor runs 7.0 versus 6.86, specificity 7.77 versus 7.56, and confidence 7.42 versus 7.21, while stress (2.1 vs 2.43) and evasion (2.5 vs 2.7) run lower. Guidance actions line up with the label - 26.8% raised versus 21.1% in the base set, and withdrawals are rare (0.31% vs 2.66%). Two discourse markers lift: 'The Question Left Hanging' appears in 0.65% of YES calls versus 0.31% of base calls, and 'Scale-Dependent Advantage Claims' in 0.20% versus 0.11%. The annual share is mostly stable in the low-to-mid 70s but dips sharply to 58.58% in 2020, recovering to 73.23% in 2021. In the matched returns sample, medians were -0.063 versus -0.072 and beat rates 40.5% versus 39.5% - small, descriptive gaps.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor7.006.86+0.14
Evasion2.502.70-0.19
Specificity7.777.56+0.21
Stress2.102.43-0.33
Promotion4.955.05-0.11
Confidence7.427.21+0.21
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised26.8%21.1%
Maintained56.0%48.8%
Lowered8.9%11.6%
Withdrawn0.3%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims0.20×2.2%11.1%
The Question Left Hanging0.65×31.2%48.0%
201570.83%
201673.27%
201773.73%
201875.06%
201972.95%
202058.58%
202173.23%
202270.45%
202373.19%
202474.23%
202569.59%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-6.3%-7.2%
Interquartile range-24.2% to +12.2%
Share beating SPY40.5% (95% CI 40%–41%)39.5%
Observations19,75122,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
SBFGQ2 20252025-07-25A
DOCQ2 20252025-07-25C
USCBQ2 20252025-07-25B+
HCAQ2 20252025-07-25C
AONQ2 20252025-07-25C
NWGQ2 20252025-07-25B+
BFHQ2 20252025-07-25B
FFICQ2 20252025-07-25B+

4Discussion

A careful reader should conclude that calls the model flags as having underwritable guidance tend to sound more candid, specific, and confident, and are more often followed by raised guidance. These are descriptions of co-occurrence within one labeling system, not evidence that the label causes anything or that the label helps predict returns. The returns gaps (median -0.063 vs -0.072; beat rate 40.5% vs 39.5%) are modest and could reflect many confounds, including which stocks have liquid trading data at all. The 2020 dip to 58.58% likely reflects the pandemic environment rather than a change in disclosure quality. Nothing here should be read as a trading signal.

5Limitations

The language scores are produced by AI readers and are inherently noisy; a YES label may partly reflect fluent writing rather than genuinely underwritable guidance. The returns sample covers only 19,751 of the YES calls out of 22,449 base calls and is skewed toward liquid names, so return comparisons may not generalize. Our own forward tests falsified directional prediction from these labels, so no claim of predictive power is made. Finally, LLMs partially remember famous stocks' histories from training data, which can contaminate any backtest by leaking future information into the labels themselves. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “The Guidance Was Fine All Along: What Model-Endorsed Calls Look Like.” Artul.ai Earnings-Call Research Library, Study No. 30. https://artul.ai/research/guidance-worth-underwriting-earnings-calls

Related studies

Ammunition, Not Smoke: Inside the Calls Where CriticKnow What You Know: Calls Where Confidence Matched tThe Question Left Hanging, Fewer and Farther BetweenTalking the Skeptic Down: 66% of Calls Leave the DouThe Admission Is Not the Confession: Calls Flagged aThe Question Left Hanging: 79,206 Calls That End on
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.