Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 40Hypotheses TestedUpdated 2026-08-28

Pressed Harder, Answered Straight: A Profile of 183 Earnings Calls

By Artul.ai Research Group · n = 183 earnings calls · First published 2026-08-28
Abstract

We study earnings calls from 2015 through 2024 in which the research hypothesis "Earned edge, pressed harder" was answered YES. Of 485 calls in the corpus, 183 (37.73%, 95% CI 33.53% to 42.13%) qualify. Relative to the base population, these calls show higher candor (6.83 vs 6.92 on the base, a small gap), lower stress (2.19 vs 2.41), and higher confidence (7.58 vs 7.32) and promotion (5.47 vs 5.09). Management raised guidance on 28.42% of these calls versus 22.06% in the base, and lowered it on 8.20% versus 13.40%. The most overrepresented phrases include "Early Products Growing Fast" (1.45x). Among 84 calls with forward returns, the median was -15.42%.

Key findings
  • 183 of 485 calls (37.73%, 95% CI 33.53% to 42.13%) answered YES to the hypothesis "Earned edge, pressed harder".
  • These calls show lower stress (2.19 vs 2.41) and higher confidence (7.58 vs 7.32) than the base population.
  • Guidance was raised on 28.42% of qualifying calls versus 22.06% of base calls, and lowered on 8.20% versus 13.40%.
  • The phrase "Early Products Growing Fast" appears 1.45x more often than expected, with "A Tiny Fraction of the Market" at 1.44x.
  • Among the 84 qualifying calls with measured forward returns, the median return was -15.42% versus -8.56% for the base set.

1Introduction

Anyone who listens to earnings calls regularly knows the moment: an analyst pushes past the prepared remarks, and management either tightens up or opens up. Calls where the edge was earned and the pushback pressed harder are a distinctive slice of the record — the moments when questioning arguably extracted real substance. If that pattern means anything, it should show up in tone, guidance behavior, recurring language, and what happened to the stock afterward. This study examines the 183 calls from 2015 through 2024 that answered YES to the hypothesis "Earned edge, pressed harder," comparing their measured tone, guidance actions, characteristic phrases, and subsequent returns against the broader corpus of 485 calls.

2Data & methodology

The corpus comprises 485 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "Earned edge, pressed harder" (n = 183; 37.7% of the reference set, 95% Wilson interval 33.5%–42.1%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The tone deltas are modest but consistent with a less defensive posture: stress runs 2.19 versus 2.41 in the base, while confidence (7.58 vs 7.32) and promotion (5.47 vs 5.09) run higher. Candor is essentially flat at 6.83 versus 6.92, and specificity edges up at 7.71 versus 7.65. Guidance behavior leans positive: raises on 28.42% of calls versus 22.06% in the base, and cuts on only 8.20% versus 13.40%. The characteristic language clusters around growth framing — "Early Products Growing Fast" at 1.45x, "A Tiny Fraction of the Market" at 1.44x, and "Founder-Led Companies" at 1.35x, with no underrepresented phrases. The forward-return picture is sobering: a median of -15.42% versus -8.56% in the base, and a beat rate of 34.52% versus 38.95%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.836.92-0.09
Evasion2.752.71+0.04
Specificity7.717.65+0.07
Stress2.192.41-0.22
Promotion5.475.09+0.38
Confidence7.587.32+0.27
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised28.4%22.1%
Maintained53.6%52.6%
Lowered8.2%13.4%
Withdrawn1.1%1.0%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Early Products Growing Fast1.45×55.7%38.6%
A Tiny Fraction of the Market1.44×43.7%30.3%
Founder-Led Companies1.35×28.4%21.0%
20150.17%
20160.14%
20170.13%
20180.11%
20190.01%
20200.00%
20210.15%
20220.19%
20230.22%
20240.06%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-15.4%-8.6%
Interquartile range-34.4% to +4.9%
Share beating SPY34.5% (95% CI 25%–45%)38.9%
Observations84190
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
SFIXQ3 20242024-06-04C+
NOAHQ1 20242024-05-30D
CRGOQ1 20242024-05-20C+
AZEKQ2 20242024-05-08B+
LINCQ1 20242024-05-06B+
NCQ1 20242024-05-05C+
ROCKQ1 20242024-05-01B+
TNETQ1 20242024-04-26C

4Discussion

A careful reader should conclude that calls matching this pattern look measurably different in tone and guidance mix — more confidence, less stress, more raises and fewer cuts, and growth-oriented language. What should not be concluded is that pressing harder caused candor, that these traits predict returns, or that the negative median return (-15.42%) means the pattern identifies underperformers. The return sample covers only 84 calls, the confidence intervals are wide, and none of these associations establish any directional edge. The study describes how these calls read, not what they foreshadow.

5Limitations

The behavioral fields are AI-read from transcripts and are inherently noisy; small deltas like the candor gap of -0.09 should not be over-interpreted. The returns sample covers 84 calls drawn from a base of 22,449 calls skewed toward liquid names, so it is not representative of the broader market. Our own forward tests falsified directional prediction on this data. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest-style analysis of earnings-call language and outcomes. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “Pressed Harder, Answered Straight: A Profile of 183 Earnings Calls.” Artul.ai Earnings-Call Research Library, Study No. 40. https://artul.ai/research/hypothesis-earned-edge-pressed-harder

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecDon't Stick to the Script: Earnings Calls Where AnswHave Your Cake and Expand the Base Too: Repeat GrowtSecond Front, Funded and Fueled: Who Says Yes on the
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.