Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 45Hypotheses TestedUpdated 2026-08-28

All Three Cylinders Firing: A Quarter of Calls Say Everything's Improving

By Artul.ai Research Group · n = 125 earnings calls · First published 2026-08-28
Abstract

This study examines earnings calls where management signaled simultaneous improvement across the front, middle, and back of the business. Of 495 calls from 2015 to 2024, 125 calls (25.3%, 95% CI 21.6% to 29.3%) met the criterion. Compared to the base, these calls showed higher confidence (7.89 vs 7.33), higher promotion (5.54 vs 5.11), and much lower stress (1.91 vs 2.39). They raised guidance 35.2% of the time versus 21.8% in the base, and lowered it only 2.4% versus 13.3%. The pattern over-indexed among skeptic-reassured calls (1.28x) and under-indexed where results were worse than direction (0.43x). Post-call returns were available for 58 calls, with a median of -7.1% versus -10.0% for the base sample.

Key findings
  • 125 of 495 calls (25.3%) answered yes to all-three-parts-improving, with a 95% CI of 21.6% to 29.3%.
  • These calls showed higher confidence (7.89 vs 7.33) and lower stress (1.91 vs 2.39) than the base.
  • Guidance was raised on 35.2% of these calls versus 21.8% of the base, and lowered on just 2.4% versus 13.3%.
  • The 58 calls with return data had a median post-call return of -7.1%, versus -10.0% for the 196-call base.
  • The pattern was 1.28x overrepresented among skeptic-reassured calls and 0.43x underrepresented where results were worse than direction.

1Introduction

Most earnings calls are a mixed report: one part of the business surging while another quietly stalls. Calls that claim everything is improving at once are rarer and worth studying, because they represent management's most confident framing of the enterprise. If that framing tracks with candor, guidance behavior, or analyst skepticism, it could help listeners calibrate how they read such calls. Using Artul.ai's research library of 495 calls from 2015 through 2024, this study identifies the 125 calls that answered yes to the hypothesis that front, middle, and back of the business were all improving in management's own words, and profiles their tone, guidance actions, thematic lifts, and post-call returns.

2Data & methodology

The corpus comprises 495 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "Front, middle, and back of the business all improving at once, in management's o" (n = 125; 25.3% of the reference set, 95% Wilson interval 21.6%–29.3%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The clearest differences are behavioral. These calls scored higher on confidence (7.89 vs 7.33), promotion (5.54 vs 5.11), and specificity (7.83 vs 7.64), and lower on stress (1.91 vs 2.39), evasion (2.63 vs 2.72), and candor (6.85 vs 6.91). Guidance tells a consistent story: 35.2% raised versus 21.8% in the base, and only 2.4% lowered versus 13.3%. Thematically, the pattern over-indexed on 'Skeptic Reassured' (1.28x lift, present in 71.3% of these calls) and under-indexed on 'Results Worse Than Direction' (0.43x) and 'When the CFO Dominates' (0.67x). The annual share fluctuated without a steady direction, from 0.09 in 2015 down to 0.0 in 2020, up to 0.15 in 2022, and 0.05 in 2024.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.856.91-0.06
Evasion2.632.72-0.08
Specificity7.837.64+0.19
Stress1.912.39-0.48
Promotion5.545.11+0.43
Confidence7.897.33+0.56
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised35.2%21.8%
Maintained52.8%52.9%
Lowered2.4%13.3%
Withdrawn0.8%1.0%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
A Tiny Fraction of the Market1.49×45.6%30.7%
Early Products Growing Fast1.36×52.8%38.8%
Founder-Led Companies1.36×28.8%21.2%
Skeptic Reassured1.28×91.2%71.3%
Results Worse Than Direction0.43×20.0%46.9%
The Hidden Segment0.47×11.2%23.6%
When the CFO Dominates0.67×9.6%14.4%
The Question Left Hanging0.69×29.6%42.8%
Underused Fixed Costs0.73×28.8%39.4%
20150.09%
20160.06%
20170.11%
20180.07%
20190.02%
20200.00%
20210.09%
20220.15%
20230.14%
20240.05%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-7.1%-10.0%
Interquartile range-26.7% to +6.7%
Share beating SPY39.7% (95% CI 28%–53%)38.3%
Observations58196
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
CRGOQ1 20242024-05-20C+
WRBYQ1 20242024-05-09A
ECPGQ1 20242024-05-08B
FWONKQ1 20242024-05-08C+
AZEKQ2 20242024-05-08B+
LINCQ1 20242024-05-06B+
AESQ1 20242024-05-03C+
CWQ1 20242024-05-02B+

4Discussion

A careful reader should conclude that calls making the everything-improving claim are a minority of the corpus, about a quarter, and that they come with more confident tone and friendlier guidance actions than the average call. That is a description of association, not a diagnosis of quality. These calls still had a median post-call return of -7.1%, and 39.7% beat the base median benchmark context, similar to the base's 38.3%. Nothing here shows the claim predicts outcomes or that management's framing causes results; confident framing can precede good quarters or bad ones.

5Limitations

The fields used here are AI-read from transcripts and are noisy by construction; tone scores and thematic labels can misclassify. The returns sample covers 22,449 calls and is skewed toward liquid names, and only 58 of these 125 calls have usable return data. Our own forward tests falsified directional prediction, so no trading implication should be drawn. Finally, LLMs partially remember famous stocks' histories, contaminating any backtest of labeled call characteristics. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “All Three Cylinders Firing: A Quarter of Calls Say Everything's Improving.” Artul.ai Earnings-Call Research Library, Study No. 45. https://artul.ai/research/hypothesis-front-middle-and-back-of-the-business-all-improving-at-once-in-mana

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecDon't Stick to the Script: Earnings Calls Where AnswHave Your Cake and Expand the Base Too: Repeat GrowtPressed Harder, Answered Straight: A Profile of 183
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.