We Ran the Same AI Citation Test Three Times. Only 11% of the Competitors Survived.
We ran the identical citation test three times in a row on our own domain, same five questions, same three engines, nothing changed in between. The headline verdict of cited or not cited was identical in all 45 measurements. The list of competitors said to be taking our place kept 1 domain out of 9. One of those two numbers is safe to build a report on. The other is not, and almost nobody checks.
Every AI visibility tool, ours included, hands you a report built from a single run. It says whether an engine names you, and it lists the competitors it named instead. We wanted to know something uncomfortable about our own product: if we run that test again five minutes later, changing nothing at all, how much of that report survives?
The method
Three consecutive runs against crawlbit.app, our own site, on 5 August 2026. Five buyer-intent questions, no brand name in any of them. Three answer engines per question, ChatGPT, Perplexity and Claude, each queried through its API with live web search. That is 15 measurements per run, 45 in total.
One detail matters more than the rest: we pinned the questions. The same five strings went to all three runs. Without that, we would have been measuring two things at once and would not have been able to tell them apart.
Result 1: the verdict never moved
Across all 45 measurements, the answer was the same: not cited. Every question, every engine, all three runs.
| Engine | Run 1 | Run 2 | Run 3 | Verdict |
|---|---|---|---|---|
| ChatGPT | 0 / 5 | 0 / 5 | 0 / 5 | stable |
| Perplexity | 0 / 5 | 0 / 5 | 0 / 5 | stable |
| Claude | 0 / 5 | 0 / 5 | 0 / 5 | stable |
Yes, that is our own tool scoring zero on its own category. We publish it because the alternative is to ask you to trust a number we would not show for ourselves. It also gives the stability result its meaning: the verdict is a solid thing to measure against.
One honest limit, and it matters. Zero is a floor. A site that is never cited cannot drift downward, so this run tells you the verdict is stable at the floor. A brand sitting at 2 or 3 out of 5, right at the threshold where an engine may or may not include it, could well move between runs. We have not measured that case yet, and we will not claim it.
Result 2: the competitor list is mostly noise
Now the part of the report that clients actually act on, the "who gets named instead of you" list. Each run returned five competitor domains. Across the three runs, nine distinct domains appeared. Exactly one appeared in all three.
| Appeared in | Domains | What it means |
|---|---|---|
| 3 runs out of 3 | 1 of 9 | Structurally present for these questions |
| 1 or 2 runs | 8 of 9 | Retrieval artefact, not a finding |
That is an 89% churn rate on the single most actionable part of the deliverable. The cause is not a bug in anyone's tool: answer engines run a live search on every call, and the retrieved set varies between calls even for an identical question. A domain that shows up because it happened to be retrieved once may simply not be there next time.
The practical consequence is blunt. If someone hands you a competitive analysis of your AI visibility built on one run, most of that competitor list is not reproducible. Ask them how many runs it rests on. If the answer is one, you are looking at a snapshot of a sampling process, presented as a finding.
The bug we found in our own code before spending a cent
Chasing this, we read our own implementation and found something worse than the noise. Our citation test regenerated its five questions on every run. An AI wrote fresh buyer questions each time the test was launched.
For a first audit, that is fine. For a 30-day re-test, it is fatal. If your baseline asked one set of questions and your re-test asks another, the before and the after are two different measurements. A score moving from 1/5 to 2/5 could be the off-page work, or it could be that the second set of questions was simply easier. Nobody could tell, including us.
This is the question to put to any provider selling you a before and after: does your re-test replay the exact questions of the baseline? If it does not, the comparison is decoration.
What we changed
Both findings are now fixed in the product, and they shipped the same day we measured them.
- Re-tests replay the baseline questions. When you re-test a domain, CrawlBit asks the exact questions of your previous run, so a change in score means the web moved, not the wording. Starting a deliberately fresh baseline stays possible, but it is now an explicit choice rather than a silent default.
- Every competitor carries its recurrence. Instead of a flat list, each competitor is reported with how many of your stored runs it appeared in. A domain seen in three runs out of three is a real competitor. A domain seen once is a candidate. That distinction is now on the page rather than in our heads.
The wider pattern: this is not a tooling problem
Reproducibility is the newest part of a familiar picture. Earlier this summer we ran live citation tests on ten SEO agencies and consultants, in their own market and language, asking the buyer question for their own service. Seven of the ten were not named at all. Those tests used a single engine, so treat the figure as indicative rather than definitive, and the domains stay anonymous because those businesses did not ask to be in a case study.
What was consistent across all ten was where the answers came from. The engines quoted comparison articles and directories, almost never the agencies' own websites. It matches what Muck Rack found across 25 million citations in ChatGPT, Claude and Gemini: about 84% come from third-party sources rather than brand-owned sites.
So the three findings stack. Your own pages are not what gets you named. The third-party pages that do get you named are measurable. And the measurement itself needs more than one run before its details mean anything.
How to read any AI visibility report from now on
- Trust the verdict, interrogate the list. Cited or not cited held perfectly for us. The supporting detail did not.
- Ask how many runs. One run is a sample. Recurrence across runs is a finding.
- Ask whether the re-test replays the baseline questions. If not, the before and after cannot be compared.
- Check what the answer cited, not just who it named. The pages an engine quotes are your actual target list.
Frequently asked questions
Partly. In three identical runs across ChatGPT, Perplexity and Claude, the cited or not cited verdict was identical in all 45 measurements. The competitor list was not: of 9 domains reported, 1 appeared in all three runs. The yes/no answer is reproducible, a single run's competitor list largely is not.
Answer engines run a live web search for every question, and the retrieved set varies from call to call even when the question is identical. A domain that appears because it happened to be retrieved once will not necessarily be retrieved again. Only competitors that recur across runs are structurally present.
Only if it asks the exact same questions as the baseline. If the tool regenerates its questions on each run, before and after are two different measurements and any movement is unreadable. Ask your provider directly whether the re-test replays the original questions.
More than one. In our measurement a single run reported five competitors, of which roughly one was durable. Treat recurrence as the signal: three runs out of three is a real competitor, one run out of three is a candidate.