Two engines, one fact, two answers
Ask ChatGPT what a software company charges for its mid-tier plan. Ask Perplexity the same question. Ask Google. There is exactly one correct answer — the company publishes it on a page you can open in a browser — and yet the three responses can differ, in the price, in the plan name, in what the plan includes, or in whether the plan still exists.
This is a strange failure mode, and it is worth being careful about what is actually going wrong, because two very different problems get discussed under the same heading.
What has actually been measured
The best public evidence on cross-engine inconsistency comes from SparkToro. On 28 January 2026, Rand Fishkin published research conducted with Patrick O’Donnell of Gumshoe.ai: 600 volunteers ran 12 prompts through ChatGPT, Claude and Google’s AI search a combined 2,961 times during November and December 2025, at 60 to 100 runs per prompt per platform. The result was that two responses to the same prompt return the same list of brands less than one per cent of the time, and the same list in the same order less than one time in a thousand. Claude was marginally steadier than the others and still highly variable. The method is published in full.
That study is the reference point for this topic, and it is properly built: real replication, a stated sample, a public method. What it measured was which brands the engines named. It did not test whether the engines agreed about any fact concerning those brands, and it does not claim to have.
A statistic we could not trace
There is a second figure in circulation, quoted in GEO commentary, holding that AI engines named the same brand only around one time in five. We tried to find its origin on 2 September 2026 and could not. Repeated searches surfaced no primary study, no method, and no sample description — only restatements. The nearest verifiable figure we found in the same numeric neighbourhood measures something entirely different: the share of brand-name searches that return an AI Overview at all.
We are not asserting that the study does not exist. We are reporting that we looked, with the specific claim in hand, and could not locate it — and that everyone repeating the number is in the same position.
That is worth pausing on, because it is the exact failure this article is about, occurring in the literature about the failure. A number with no locatable source gets repeated until repetition substitutes for evidence. It is the same mechanism as a price that is wrong on your pricing page: nobody is lying, nobody checked, and the claim propagates on the strength of having been stated confidently. If a statistic cannot be traced to a method, it should be quoted with that caveat attached or not quoted at all.
Inconsistent about which brand, inconsistent about what is true
These are separate problems with separate consequences.
Recommendation inconsistency is the SparkToro finding: ask for the best tools in a category and you get a different shortlist each time. This is a marketing and measurement problem. It means visibility tracking needs replication to mean anything, and that a single-run rank check is close to worthless.
Factual inconsistency is different. Here the engine has already decided to talk about your company, and the question is whether what it says is true. When Cursor’s support assistant told users in April 2025 that a subscription was limited to one device, it was not failing to mention a brand. It was stating a policy that did not exist, about the company it worked for, to that company’s paying customers. The logouts were a session bug; the policy was invented. That incident is documented by The Register and catalogued as Incident 1039 in the AI Incident Database.
A dashboard measuring share of voice would have scored that answer as a success.
Why engines disagree about a fact with one correct answer
Four causes, and only one of them is the engine’s fault.
They are reading different corpora
Each engine crawls with its own agent under its own rules, and the agent that governs whether you appear in answers is often not the one people block or allow. Appearance in ChatGPT’s answers is governed by OAI-SearchBot, not GPTBot, which is training-only. Perplexity uses PerplexityBot for its index, while its user-initiated fetches do not respect robots.txt at all. Google’s AI surfaces are fed by Googlebot; Google-Extended controls Gemini training and does not govern AI Overviews. Three engines, three indexes, three different subsets of your site — and three different subsets of everything written about you elsewhere.
They are reading different points in time
A fact you changed last week exists in several states at once: the current version on your site, an older version in one engine’s index, an older version still in a third-party listing, and a cached copy somewhere behind all of them. The engines are not disagreeing about the present. They are each reporting a different past, accurately.
Your own surfaces disagree with each other
This is the uncomfortable one, and in our experience it is the most common.
The price is on the pricing page. It is also in a help-centre article, in the API documentation, in a comparison page written for a campaign two years ago, in a PDF a partner still hosts, and in a launch blog post nobody has opened since. When those disagree, an engine choosing between them is not hallucinating. It is faithfully reporting a contradiction that already exists in your published corpus, and the engine picked a different branch than you would have.
The uncomfortable implication is that cross-engine disagreement is partly a measurement of your own internal contradiction. Some of the variance sits in the engines. Some of it is yours, and it was there before any AI read it.
And non-determinism sits on top of all of it
Given the SparkToro result, some observed disagreement is simply sampling. Two engines producing different answers once tells you almost nothing. This is why any credible disagreement rate has to come from repeated runs, and why a screenshot of two chat windows side by side is an anecdote rather than a measurement.
How we are measuring it
Our cross-engine disagreement rate is M3 in the Stale Answer Benchmark, whose full protocol we published before collecting any data. In short: sixteen companies, four claims each — a price, a policy, a capability, a compliance status — with the company’s own current published page as ground truth, queried repeatedly against three engines in independent sessions.
M3 is defined as the share of claims where at least two engines gave materially different answers to the same question. Two design decisions matter for interpreting it:
Repeated runs, not single shots. A claim is only counted as a disagreement if the difference persists across runs. Otherwise we would be reporting non-determinism as if it were engine divergence.
Disagreement is scored separately from correctness. Three engines can agree and all be wrong — reproducing the same stale figure from the same outdated source. That case is a failure with a completely different fix, and collapsing it into a disagreement metric would hide it. Agreement is not accuracy.
Results are not in yet. When they are, they will be published in aggregate and anonymised, with the scoring rubric and the borderline calls included.
What you can do before anyone’s data lands
One part of this you can act on today without any tooling, because it is the part that belongs to you.
Pick your four claims. The price a buyer asks about, the policy that appears in contracts, the capability sales leans on, the compliance status procurement checks. Four is enough to start.
Find every place each one is stated. Site, docs, help centre, PDFs, partner listings, old campaign pages, your own blog archive. Search your domain for the figure rather than trusting your memory of where it lives.
Count the contradictions. Most teams doing this for the first time find at least one, usually in a surface nobody owns. Every contradiction you find is a branch an engine can legitimately pick.
Give each claim one home and a real date. Decide which page is canonical, make the others point at it, and ensure the visible date and the machine-readable date agree. If a page carries a verification date, that date should mean somebody actually re-checked the fact — otherwise it asserts nothing beyond the publication date.
None of this guarantees an engine will answer correctly. It removes the failure mode where the engine answers incorrectly and is, technically, quoting you.
The honest summary
What is established: AI engines are highly inconsistent about which brands they name, rigorously demonstrated at scale by SparkToro and Gumshoe.ai in January 2026.
What is not established: how often engines disagree about specific verifiable facts concerning a company, how often they are simply wrong, and how much of that variance originates in the company’s own contradictory publishing rather than in the engines. We could find no published study measuring any of it. That is the gap our benchmark is aimed at, and this page will be updated when we have something to report rather than something to assert.
