Findabl

by3cubed.ai

How Stable Are AI
Recommendations?

2,580 RUNS. 26 QUESTIONS. FOUR ENGINES.

Findabl ran 2,580 AI recommendation tests across ChatGPT, Gemini, Claude, and Perplexity to measure whether brands keep getting named.

The result: AI recommendation is not a rank you win. It is a probability that changes by engine, query, and retrieval.

AI visibility is not a ranking.

It's a probability that resets with every query, every engine, and every model update.

WHAT ACTUALLY MOVES THE ANSWER

AI recommendation is the result of four distinct forces. They compound, and they do not stay fixed.

1

MODEL FLOOR

Built-in variation persists even in a controlled setting.

78-87%
2

NORMAL SAMPLING

Language generation changes before retrieval begins.

67-78%
3

LIVE WEB

Different retrieved sources change the recommendation.

48-68%
4

ENGINE DIVERGENCE

Four engines agree on the same winner only 27% of the time.

27%

Track the answer, not a theoretical score. One-time audits cannot represent a system that changes from run to run.

2,580

REPEATED RUNS

Direct measurements, not modeled estimates.

26

BUYER QUESTIONS

Unaided category questions across three markets.

4

AI ENGINES

ChatGPT, Gemini, Claude, and Perplexity.

376

SOURCE DOMAINS

External domains cited across all engines.

1

RECOMMENDATION STABILITY ISN'T FIXED

With web search off, the same question returns the same top brand 78-87% of the time. Turn on live search, how people really use these tools, and consistency drops fast.

ChatGPT

78→48%

Gemini

78→54%

Claude

87→50%

Perplexity

68% web-only

WORDING ISN'T THE VARIABLE.

The web retrieval layer is. Different runs pull different sources, and different sources change the answer.

2

ENGINES DON'T AGREE WITH EACH OTHER

Only 27% of questions got the same top brand from all four engines simultaneously. Three out of four questions had at least one engine break away.

UNANIMOUS

ADHD telehealth: majority on all 4 engines

SPLIT DECISION

HNW RIA: 4 engines, 4 different brands

A COMPETITIVE MAP, NOT NOISE.

A brand can lead on ChatGPT and be invisible on Perplexity. The gaps are measurable and claimable.

3

A FEW PUBLICATIONS CONTROL THE ANSWER

Of 376 domains cited across all engines, the top 10 accounted for 28.6% of citations. Chambers alone was cited 250 times.

YOUR SITE CHAMBERS8.2%

The model doesn't reward your homepage. It rewards the publication it already trusts. Earned media is the new backlink.

4

NO BRAND PERMANENTLY OWNS AN ANSWER

The median brand appeared in just 10% of responses for its own category question. 85.7% of brand-question pairs fell below 50%.

NO PERMANENT WINNER.
NO LOCKED RANKING.
NO SET ANSWER.

Just a probability that resets with every query.

5

WHERE TO MOVE (2-MINUTE SELF-CHECK)

Market fragmentation predicts how big the opportunity is. Run the same consistency check on your own category.

0.68-1.00

Clear leader

Defend & monitor drift

0.40-0.70

Moderate

Find adjacent sub-questions

0.15-0.35

Fragmented

Move first on gatekeepers

See where your category sits. Measure appearance rate across all four engines.