Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 43Hypotheses TestedUpdated 2026-08-28

Validators, Assemble: A Profile of Calls Where External Backing Converges

By Artul.ai Research Group · n = 109 earnings calls · First published 2026-08-28
Abstract

This study examines 109 earnings calls (out of 497 scored, a 21.9% share with a 95% CI of 18.5% to 25.8%) that answered YES to the hypothesis "External validators converging," drawn from calls between 2015 and 2024. These calls skew promotional (5.51 vs 5.10) and confident (7.49 vs 7.34), with near-baseline candor and specificity. Guidance was raised on 26.6% of them versus 21.9% in the base. Topic lifts are striking: "Volume About to Step Up" appears 1.59x over expectation and "A Tiny Fraction of the Market" 1.58x. Among 38 calls with forward returns, the median was -6.1% versus -10.0% for the base, though only 47.4% beat, statistically indistinguishable from the base's 38.8%.

Key findings
  • Calls flagged for converging external validators make up 21.9% of the corpus (109 of 497), with a 95% CI of 18.5% to 25.8%.
  • These calls score higher on promotion (5.51 vs 5.10) and confidence (7.49 vs 7.34) than the baseline, while candor (6.84 vs 6.91) and stress (2.31 vs 2.38) are slightly lower.
  • Guidance was raised on 26.6% of these calls versus 21.9% in the base, and lowered on only 10.1% versus 13.1%.
  • Among the 38 calls with forward returns, the median was -6.1% versus -10.0% for the base, and 47.4% beat versus 38.8%, with a CI of 32.5% to 62.7%.

1Introduction

Analysts often treat converging external validators—analyst upgrades, peer results, third-party data points lining up behind a company's story—as a signal that the narrative has independent support. But an earnings call is also a performance, and a management team that knows the outside world agrees with it may lean into promotion and confidence rather than plain facts. Understanding how these calls sound, what topics they emphasize, and what guidance they give helps anyone who reads transcripts separate corroborated stories from well-rehearsed ones. This study profiles the 109 calls that answered YES to the hypothesis "External validators converging" and compares their language, guidance behavior, topic emphasis, and forward-return context against the broader corpus.

2Data & methodology

The corpus comprises 497 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "External validators converging" (n = 109; 21.9% of the reference set, 95% Wilson interval 18.5%–25.8%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The behavioral profile is subtle: promotion rises to 5.51 from 5.10 and confidence to 7.49 from 7.34, while candor (6.84 vs 6.91), evasion (2.80 vs 2.71), and specificity (7.72 vs 7.65) sit near baseline. Guidance skews positive—26.6% raised versus 21.9% base, 10.1% lowered versus 13.1%. Topic lifts tell a sharper story: "Volume About to Step Up" (1.59x), "A Tiny Fraction of the Market" (1.58x), and "Founder-Led Companies" (1.52x) dominate, while "When the CFO Dominates" runs at 0.70x. The annual share peaked at 0.14 in 2023 after a 0.00 low in 2020. Forward returns on 38 calls show a median of -6.1% versus -10.0% base, but the beat rate of 47.4% (CI 32.5% to 62.7%) overlaps the base's 38.8%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.846.91-0.07
Evasion2.802.71+0.09
Specificity7.727.65+0.07
Stress2.312.38-0.07
Promotion5.515.10+0.41
Confidence7.497.34+0.15
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised26.6%21.9%
Maintained55.0%53.1%
Lowered10.1%13.1%
Withdrawn0.9%1.0%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Volume About to Step Up1.59×40.4%25.4%
A Tiny Fraction of the Market1.58×48.6%30.8%
Founder-Led Companies1.52×32.1%21.1%
Early Products Growing Fast1.37×53.2%38.8%
Underused Fixed Costs1.32×51.4%39.0%
When the CFO Dominates0.70×10.1%14.5%
20150.09%
20160.05%
20170.11%
20180.05%
20190.01%
20200.00%
20210.08%
20220.13%
20230.14%
20240.05%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-6.1%-10.0%
Interquartile range-34.2% to +16.1%
Share beating SPY47.4% (95% CI 32%–63%)38.8%
Observations38196
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
MMLPQ2 20242024-07-18B
CRGOQ1 20242024-05-20C+
EROQ1 20242024-05-10A
AZEKQ2 20242024-05-08B+
AESQ1 20242024-05-03C+
WECQ1 20242024-05-01A
DUOTQ4 20232024-04-01F
ASMQ4 20232024-03-21C+

4Discussion

A careful reader should conclude that calls with converging external validators sound measurably more promotional and confident, emphasize growth-framing topics, and more often raise guidance. That is a description of the sample, not a verdict on management quality or a signal about what comes next. The returns comparison—median -6.1% versus -10.0%, beat rate 47.4% versus 38.8%—has wide uncertainty with only 38 observations, and overlapping confidence intervals mean no reliable difference in outcomes can be claimed. These findings characterize how such calls read, not whether that tone is deserved or predictive.

5Limitations

All language and topic fields are AI-read and inherently noisy, so deltas of a few tenths should be treated as suggestive. The returns sample covers only 38 of these calls, and the underlying 22,449-call dataset is skewed toward liquid names. Our own forward tests falsified directional prediction from these signals. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest-style comparison and inflate apparent patterns. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “Validators, Assemble: A Profile of Calls Where External Backing Converges.” Artul.ai Earnings-Call Research Library, Study No. 43. https://artul.ai/research/hypothesis-external-validators-converging

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecDon't Stick to the Script: Earnings Calls Where AnswHave Your Cake and Expand the Base Too: Repeat GrowtPressed Harder, Answered Straight: A Profile of 183
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.