We Ran the Same AI Citation Test Three Times. Only 11% of the Competitors Survived.

The short answer

We ran the identical citation test three times in a row on our own domain, same five questions, same three engines, nothing changed in between. The headline verdict of cited or not cited was identical in all 45 measurements. The list of competitors said to be taking our place kept 1 domain out of 9. One of those two numbers is safe to build a report on. The other is not, and almost nobody checks.

Every AI visibility tool, ours included, hands you a report built from a single run. It says whether an engine names you, and it lists the competitors it named instead. We wanted to know something uncomfortable about our own product: if we run that test again five minutes later, changing nothing at all, how much of that report survives?

The method

Three consecutive runs against crawlbit.app, our own site, on 5 August 2026. Five buyer-intent questions, no brand name in any of them. Three answer engines per question, ChatGPT, Perplexity and Claude, each queried through its API with live web search. That is 15 measurements per run, 45 in total.

One detail matters more than the rest: we pinned the questions. The same five strings went to all three runs. Without that, we would have been measuring two things at once and would not have been able to tell them apart.

Result 1: the verdict never moved

Across all 45 measurements, the answer was the same: not cited. Every question, every engine, all three runs.

EngineRun 1Run 2Run 3Verdict
ChatGPT0 / 50 / 50 / 5stable
Perplexity0 / 50 / 50 / 5stable
Claude0 / 50 / 50 / 5stable

Yes, that is our own tool scoring zero on its own category. We publish it because the alternative is to ask you to trust a number we would not show for ourselves. It also gives the stability result its meaning: the verdict is a solid thing to measure against.

One honest limit, and it matters. Zero is a floor. A site that is never cited cannot drift downward, so this run tells you the verdict is stable at the floor. A brand sitting at 2 or 3 out of 5, right at the threshold where an engine may or may not include it, could well move between runs. We have not measured that case yet, and we will not claim it.

Result 2: the competitor list is mostly noise

Now the part of the report that clients actually act on, the "who gets named instead of you" list. Each run returned five competitor domains. Across the three runs, nine distinct domains appeared. Exactly one appeared in all three.

Appeared inDomainsWhat it means
3 runs out of 31 of 9Structurally present for these questions
1 or 2 runs8 of 9Retrieval artefact, not a finding

That is an 89% churn rate on the single most actionable part of the deliverable. The cause is not a bug in anyone's tool: answer engines run a live search on every call, and the retrieved set varies between calls even for an identical question. A domain that shows up because it happened to be retrieved once may simply not be there next time.

The practical consequence is blunt. If someone hands you a competitive analysis of your AI visibility built on one run, most of that competitor list is not reproducible. Ask them how many runs it rests on. If the answer is one, you are looking at a snapshot of a sampling process, presented as a finding.

The bug we found in our own code before spending a cent

Chasing this, we read our own implementation and found something worse than the noise. Our citation test regenerated its five questions on every run. An AI wrote fresh buyer questions each time the test was launched.

For a first audit, that is fine. For a 30-day re-test, it is fatal. If your baseline asked one set of questions and your re-test asks another, the before and the after are two different measurements. A score moving from 1/5 to 2/5 could be the off-page work, or it could be that the second set of questions was simply easier. Nobody could tell, including us.

This is the question to put to any provider selling you a before and after: does your re-test replay the exact questions of the baseline? If it does not, the comparison is decoration.

What we changed

Both findings are now fixed in the product, and they shipped the same day we measured them.

The wider pattern: this is not a tooling problem

Reproducibility is the newest part of a familiar picture. Earlier this summer we ran live citation tests on ten SEO agencies and consultants, in their own market and language, asking the buyer question for their own service. Seven of the ten were not named at all. Those tests used a single engine, so treat the figure as indicative rather than definitive, and the domains stay anonymous because those businesses did not ask to be in a case study.

What was consistent across all ten was where the answers came from. The engines quoted comparison articles and directories, almost never the agencies' own websites. It matches what Muck Rack found across 25 million citations in ChatGPT, Claude and Gemini: about 84% come from third-party sources rather than brand-owned sites.

So the three findings stack. Your own pages are not what gets you named. The third-party pages that do get you named are measurable. And the measurement itself needs more than one run before its details mean anything.

How to read any AI visibility report from now on

Frequently asked questions

Is an AI citation test reproducible?

Partly. In three identical runs across ChatGPT, Perplexity and Claude, the cited or not cited verdict was identical in all 45 measurements. The competitor list was not: of 9 domains reported, 1 appeared in all three runs. The yes/no answer is reproducible, a single run's competitor list largely is not.

Why does the competitor list change between two reports?

Answer engines run a live web search for every question, and the retrieved set varies from call to call even when the question is identical. A domain that appears because it happened to be retrieved once will not necessarily be retrieved again. Only competitors that recur across runs are structurally present.

Can a 30-day re-test prove that AI visibility work paid off?

Only if it asks the exact same questions as the baseline. If the tool regenerates its questions on each run, before and after are two different measurements and any movement is unreadable. Ask your provider directly whether the re-test replays the original questions.

How many runs before trusting a competitor list?

More than one. In our measurement a single run reported five competitors, of which roughly one was durable. Treat recurrence as the signal: three runs out of three is a real competitor, one run out of three is a candidate.

See which competitors are actually durable

Run a free CrawlBit scan. It tests whether answer engines name you, shows the exact pages they quoted instead, and tracks how often each competitor recurs across your runs rather than reporting a single snapshot.

Scan my site free →