Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 63Hypotheses TestedUpdated 2026-08-28

It Takes Two to Tango: Calls With Two-Sided Intensity

By Artul.ai Research Group · n = 88 earnings calls · First published 2026-08-28
Abstract

This study examines 88 earnings calls (17.7% of a 496-call corpus, 2015-2024) flagged by Artul.ai's research hypothesis 'Two-sided intensity' -- calls where both upbeat and downbeat signals register strongly. These calls skew toward founder-led companies (1.73x over-representation) and early products growing fast (1.61x). Their language profile shows lower stress (1.94 vs 2.39) and higher confidence (7.81 vs 7.33) than the base. Guidance was raised on 30.7% of these calls vs 22.0% in the base. Where 37 of these calls had returns data, the median was -13.0% vs -10.5% in the base sample, and 40.5% beat. The pattern peaked in 2022 at 0.14 and faded to 0.02 by 2024.

Key findings
  • 88 of 496 calls (17.7%, 95% CI 14.6%-21.3%) answered YES to the 'Two-sided intensity' hypothesis.
  • Founder-led companies are over-represented at 1.73x, and 'Early Products Growing Fast' at 1.61x.
  • These calls show higher confidence (7.81 vs 7.33) and lower stress (1.94 vs 2.39) than the base corpus.
  • In the 37-call returns sample, the median return was -13.0% vs -10.5% in the base, with 40.5% beating vs 38.1% in the base.

1Introduction

Most earnings calls have a temperature: warm and promotional, or cool and defensive. A smaller group is genuinely two-sided -- management pushing growth narratives while fielding pointed pushback in the same conversation. That mix is interesting because it is where a company's story meets its skeptics in real time. Followers of earnings calls often treat such calls as tension points worth watching. This study examines 88 calls from the Artul.ai corpus (2015-2024) that answered YES to the 'Two-sided intensity' hypothesis, describing their language profile, guidance behavior, thematic over- and under-representations, prevalence over time, and later return outcomes.

2Data & methodology

The corpus comprises 496 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "Two-sided intensity" (n = 88; 17.7% of the reference set, 95% Wilson interval 14.6%–21.3%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The language profile of these calls is calmer than the base: stress 1.94 vs 2.39, confidence 7.81 vs 7.33, and promotion 5.73 vs 5.10, while candor is slightly lower (6.78 vs 6.92). Thematically, founder-led companies (1.73x), a tiny fraction of the market (1.69x), and early products growing fast (1.61x) are over-represented; results worse than direction (0.55x) and the CFO dominating (0.62x) are under-represented. Guidance was raised on 30.7% of calls vs 22.0% in the base, and lowered on 5.7% vs 13.3%. Prevalence rose to 0.14 in 2022, then fell to 0.02 in 2024. The 37-call returns sample shows a median of -13.0% vs -10.5% in the base.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.786.92-0.14
Evasion2.692.70-0.01
Specificity7.727.65+0.07
Stress1.942.39-0.45
Promotion5.735.10+0.63
Confidence7.817.33+0.48
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised30.7%22.0%
Maintained56.8%53.0%
Lowered5.7%13.3%
Withdrawn1.1%1.0%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Founder-Led Companies1.73×36.4%21.0%
A Tiny Fraction of the Market1.69×52.3%30.8%
Early Products Growing Fast1.61×62.5%38.7%
Results Worse Than Direction0.55×26.1%47.2%
When the CFO Dominates0.62×9.1%14.5%
The Question Left Hanging0.66×28.4%42.7%
20150.06%
20160.05%
20170.08%
20180.06%
20190.01%
20200.00%
20210.06%
20220.14%
20230.07%
20240.02%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-13.0%-10.5%
Interquartile range-47.3% to +12.0%
Share beating SPY40.5% (95% CI 26%–57%)38.1%
Observations37197
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
CRGOQ1 20242024-05-20C+
WRBYQ1 20242024-05-09A
LINCQ1 20242024-05-06B+
SNVQ1 20242024-04-18B
PBRQ4 20232024-03-08D
NICEQ4 20232024-02-22B+
SBSQ3 20232023-11-10C+
DASHQ3 20232023-11-01C+

4Discussion

A careful reader should treat this as a description of a call archetype, not a verdict. Two-sided-intensity calls coincide with founder-led companies and growth narratives, and with more frequent guidance raises, but none of these associations establish that one causes the other. The returns comparison (median -13.0% vs -10.5%) describes the past sample only; it is not a forecast or an edge. The 40.5% beat rate overlaps with the base's 38.1%. Read the deltas as characterization of a slice of the corpus, and hold any stronger conclusions until the pattern is tested out of sample.

5Limitations

Artul.ai's call fields are AI-read and noisy, so profiles and thematic lifts carry labeling error. The returns sample covers only 37 of the 88 calls, drawn from a base of 22,449 calls skewed toward liquid names, so results may not generalize. Our own forward tests falsified directional prediction, and any backtest here inherits that risk. LLMs partially remember famous stocks' histories, which can contaminate labels and any retrospective performance comparison. The 2015-2024 window also mixes very different market regimes. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “It Takes Two to Tango: Calls With Two-Sided Intensity.” Artul.ai Earnings-Call Research Library, Study No. 63. https://artul.ai/research/hypothesis-two-sided-intensity

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecDon't Stick to the Script: Earnings Calls Where AnswHave Your Cake and Expand the Base Too: Repeat GrowtPressed Harder, Answered Straight: A Profile of 183
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.