There is a scenario every GEO dashboard on the market will report as a success.
An engine recommends you first, in a positive tone, citing your own domain — and quotes a price you stopped charging fourteen months ago. Visibility: high. Sentiment: positive. Share of voice: winning. Citation source: your site. Every metric green.
Nothing in that stack was measuring whether the answer was correct. This is not a defect in the tools. It is a scope boundary, and it is worth naming precisely, because the boundary is invisible from inside the dashboard.
Three different measurements
The category talks about “AI answer quality” as one thing. It is at least three, and what these tools measure blurs them.
Visibility — does the brand appear, in what position, in how many of the sampled answers? A counting problem. Every tool on the shortlist does this.
Sentiment — is the surrounding language positive, neutral, or negative? A classification problem, and a solved-enough one. Most tools do this.
Factual accuracy — are the specific assertions in the answer — the price, the tier limits, the certification, the integration list, the policy — actually correct today? A verification problem. It requires a source of truth, and it is categorically harder than the other two.
The first two can be computed from the answer text alone. The third cannot be computed from the answer text at any level of sophistication, because the answer text does not contain the fact needed to check it.
What the nine actually measure
Two of the nine are no longer independent, which changes who owns the roadmap — Semrush to Adobe, Scrunch to Sitecore. The dates:
Deal | Milestone | Date |
|---|---|---|
Adobe / Semrush | Announced | 19 November 2025 |
Adobe / Semrush | Completed | 28 April 2026 |
Adobe / Semrush | AI Optimization merged into Adobe Brand Visibility | 17 June 2026 |
Sitecore / Scrunch | Announced | 3 June 2026 |
Tool | Visibility | Sentiment | Claim-level accuracy |
|---|---|---|---|
Profound | Yes | Yes | Partial — FactCheck |
Semrush (an Adobe company) | Yes | Yes | None documented |
Ahrefs Brand Radar | Yes | Yes | None documented |
Peec AI | Yes | Yes | None documented |
Otterly.AI | Yes | Yes | None documented |
Athena | Yes | Yes | Partial — Oracle, Enterprise only |
Evertune | Yes | Yes | Narrow — prices only |
Scrunch AI (a Sitecore company) | Yes | Yes | None documented |
Brandlight | Yes | Yes | None documented |
“None documented” means the vendor publishes no such capability on any page we could read. It is a statement about the documentation, not proof that nothing exists behind an enterprise login.
Six of the nine document no claim-level checking of any kind. A seventh, Evertune, has exactly one narrow case. That leaves two real exceptions — and all three are worth reading closely, because none of them is quite what the category summaries say it is.
The exceptions, read closely
Profound FactCheck
FactCheck launched on 14 July 2026. It analyses AI responses at sentence level, isolates verifiable claims with a proprietary claim-detection model, classifies each as accurate, inaccurate or not relevant, and routes remediation by who owns the source. It scores those claims against a Knowledge Base the customer builds from crawls, direct uploads, and Google Drive and Notion syncs. It is the most advanced claim-checking in the category by a distance.
Who can buy it is less settled than the coverage suggests, and the ambiguity is Profound’s own. The press release says FactCheck “comes included with existing Profound configurations for brands” — read literally, that covers every brand tier, including the $99/month Starter plan. The launch blog says it is available “for all Profound customers on enterprise plans at no additional cost.” The two statements are mutually exclusive, and FactCheck has no row in the pricing comparison table on either billing tab, so the page cannot break the tie.
What is clearly gated is programmatic access: the POST /v2/reports/factcheck endpoint requires the REST API, which is Enterprise-only and separately available only on request. If the tier gate matters to your budget, it is a question for the vendor, not for a comparison table.
Three documented constraints are worth knowing before you commit: you can connect only one Knowledge Base and cannot change it after setup, results take seven days to appear, and Profound’s own guidance is to build the prompt set at 100–200 fact-based prompts.
Athena’s Oracle
Athena’s Enterprise tier lists “Oracle discrepancy detection” alongside a Knowledge Base and claim review. That is the entirety of the public record: a single line item on a pricing table. Athena publishes no technical documentation for Oracle — what it compares answers against, who maintains that source, and how often it refreshes are all unstated. The capability may be excellent. There is no way to tell from outside.
It is also the tier you cannot try. Oracle sits behind a custom Enterprise contract, above the Starter plan, which lists at $295/month on Athena’s default monthly billing, or $245/month billed annually — a string that appears only in the site’s JavaScript bundle, not in the served HTML. Athena publishes no annual total, so none is quoted here. That places Oracle outside the tools you can evaluate without a sales call; the wider tier structure is in the pricing breakdown.
Evertune’s one narrow case
“None” is too strong for Evertune. Inside Shopping Intelligence, Evertune describes “fact-checking the prices AI is quoting against your actual prices” — real verification, scoped to exactly one class of fact, with no stated tier gate. It is not general claim checking and Evertune does not present it as such. It is instructive precisely because it is so narrow: even the simplest working case is a diff against a list the customer supplies.
Two things the coverage gets wrong
The hallucination detection that is not there
Scrunch is widely credited with hallucination detection. It does no claim-level accuracy checking. Its pricing page contains no mention of hallucination, accuracy, misinformation or fact-checking in any tier, and its published metric set is brand presence matched as a binary, plus three-bucket sentiment — “positive, mixed, or negative”.
The “91% fewer hallucinations” figure that circulates alongside Scrunch’s name is a misread of direction. It is a claim about the Agent Experience Platform reducing hallucinations by serving structured content to agents — the opposite operation from detecting a false claim in an answer. The third-party line calling Scrunch “the only AI visibility tool with dedicated hallucination detection” has no support on Scrunch’s own pages. That error belongs to the review layer, not to Scrunch’s marketing.
Neither exception was first to market
Bluefish shipped “AI Accuracy” with “Brand Vault” on 5 May 2026, roughly ten weeks before FactCheck, under a headline about bringing brand verification to AI channels for the first time. Profound’s launch copy calls FactCheck “the first way for brands to analyze AI…” — the headline is truncated in our source, so treat the exact wording as unconfirmed. Bluefish is not on the shortlist, but it shipped first. Treat both firstness claims as marketing.
The assumption the working ones share
Of the accuracy capabilities documented well enough to inspect, every one outsources the ground truth to the customer. Profound checks extracted claims against a Knowledge Base the customer builds and maintains. Evertune checks quoted prices against the customer’s actual prices. Athena publishes nothing about Oracle’s source of truth — which is not evidence that Athena is different, only that you cannot check.
No tool in this category maintains an independent authoritative corpus of brand facts. “Verified accuracy” means “consistent with a document set you uploaded.”
Profound is the only vendor in the set that documents the resulting failure mode in writing, and it deserves credit for that. Its help centre states that “the accuracy of your results depends on keeping this source of truth up to date,” and then names the specific way it breaks: “a claim about an old pricing tier might be marked accurate if the Knowledge Base doesn’t include the updated pricing page.”
Read the direction of that failure carefully. It is a false positive, not a miss. A missed error leaves you where you started. A false positive puts a green tick beside a wrong answer and tells you the work is done. A brand that updates its prices but not its Knowledge Base gets a clean score while every engine quotes last year’s number, and the score is technically correct. The engine and your Notion agree. They are agreeing about something that stopped being true last year.
The uncomfortable version: most companies have never reconciled their own claims against each other. The pricing page, the sales deck, the onboarding email, the support macro and the FAQ were written at different times by different people, and nobody has ever diffed them. Point a checker at that corpus and it will faithfully report agreement with whichever version it happened to crawl. Buying one of these is buying a content-operations obligation, not a dashboard.
The number underneath the number
Before you can ask whether an answer is true, you have to ask how many answers you looked at. Eight of the nine either take one draw per prompt per cycle or decline to publish a number at all; only Evertune claims multi-sampling, and its own homepage gives that figure two ways — “100 times” per model in one place, “up to 100x” in another. The IAB’s August 2026 position is blunt about what that means: “single-response measurement is not measurement,” and any metric derived from one response per query “reflects a sample of one.”
Nor is the mention count clean. OpenAI began showing labelled sponsored placements in ChatGPT to logged-in US Free and Go users on 9 February 2026 and opened a self-serve ads manager on 5 May 2026. No vendor in the set documents whether it separates sponsored from organic placements, or whether it samples logged-in or logged-out — which now decides whether ads are visible to the scraper at all.
Why sentiment is not a substitute
Teams reach for sentiment as a proxy for accuracy. It does not work, and the failure is directional.
A factually wrong answer that is flattering scores as positive. A factually correct answer describing a genuine limitation scores as negative. Optimising sentiment therefore pushes, at the margin, away from accuracy — you are rewarded for engines saying nice things and penalised for engines saying true ones.
The stale-price case is the clean illustration. An engine quoting your discontinued lower price is describing you generously. Sentiment reads positive. Your sales team spends the quarter explaining the difference on calls.
Six questions to ask a GEO vendor
Any vendor on this list will answer these. The answers separate them faster than a feature matrix.
What is your ground truth, and who maintains it? If the answer is “we crawl your site,” the tool measures consistency with your site, not correctness.
Does a stale source of truth show up as an error, or as a pass? Profound documents that it shows up as a pass. Ask everyone else.
What happens when two of my own sources disagree? Does it flag the internal contradiction, or silently pick one?
Do you check claims, or mentions? A claim is “the Pro plan includes SSO.” A mention is your brand name in a sentence.
How many times do you sample the same prompt, and what triggers a re-check — a schedule, or a change to the fact? Calendars, mostly. That determines your worst-case staleness window.
Can it see surfaces no crawler reaches? The sales deck, the support macro, the partner PDF. The claims a crawler can reach are a subset of the claims your buyers see.
What this is not
This is not a claim that the nine tools are defective. They do what they say: they measure representation in AI answers, and several do it very well. Nor is “none documented” the same as “does not exist” — this sweep covered public pages, and an enterprise-only capability with no public page would not surface. Athena’s Oracle is a single line item on a pricing table, which is fair warning that other vendors may have similarly under-documented features.
It is a claim about what remains unmeasured after you have bought one. Visibility tooling tells you the engine is talking about you. It does not tell you the engine is right, and no amount of prompt tracking will produce that answer, because the required input — a maintained, reconciled record of what is currently true about your company — does not exist inside the tool.
That record is the missing layer. Every category of company fact has a system of record: master data has MDM, product attributes have PIM, regulated pharma claims have Veeva, security posture has Vanta. The claims written in ordinary sentences — pricing, policy, capabilities, certifications, metrics — have none.
Buy the visibility tool. Then ask what it is checking against.
How this was verified
Feature coverage in the table was re-read at source on 1–2 September 2026. This update corrects and re-sources the 1 September version in the places below.
What changed. Evertune moved from “No” to “Narrow”: it documents price verification inside Shopping Intelligence, so a blanket “none” was wrong for Evertune specifically. Profound’s FactCheck tier gate moved from settled to reported-as-conflicting — the press release and the launch blog disagree and the pricing page has no FactCheck row, so this article reports the conflict rather than picking a side. Athena’s Oracle is now marked undocumented; the earlier draft implied a knowable architecture behind it. Scrunch’s absence of claim checking is now stated affirmatively from its own pricing page, read at source, rather than inferred from Sitecore’s acquisition release — the same pass established that the circulating “91% fewer hallucinations” figure describes hallucination reduction by the Agent Experience Platform, not detection. Brandlight was read at source: the homepage names six engines and brandlight.ai/pricing returns HTTP 404, so Brandlight publishes no pricing and none is quoted here.
Prices and provenance. Athena’s $295/month Starter and $245/month annual figure come from Athena’s own plans page, the annual figure from the site’s shipped JavaScript rather than the served HTML. Profound’s $99/month Starter is from Profound’s pricing page in its default state; “Billed yearly” is a per-card toggle that ships off, and toggling it shows $82.50/month. Neither vendor publishes an annual total, and none is calculated here. No price in this article is filled in from a third-party aggregator.
Ownership. Adobe completed its acquisition of Semrush on 28 April 2026, per Adobe’s newsroom. Sitecore announced its acquisition of Scrunch on 3 June 2026, and that acquisition is company-confirmed. The widely reported ~$225M price is not: Sitecore disclosed no financial terms and both companies declined to comment on the valuation.
Still unverified. Athena’s Oracle ground truth, owner and refresh model. Whether FactCheck is available below Enterprise. Sampling depth: Otterly, Scrunch, Semrush, Brandlight and the Ahrefs Index publish no figure, and Profound’s one response per prompt per engine per day is our arithmetic on its published response caps, not a vendor statement. The ChatGPT advertising material and Scrunch’s current tier contents come from summarised reads and should be re-checked at source before anyone relies on them. “Partial” and “Narrow” in the table mean the vendor ships a named capability, not that it solves the problem.
No first-party testing was conducted for this article. Where we make a comparative claim, it is from documentation, not from a benchmark — see the testing protocol for how to run your own. Knowledge Company published this comparison from vendor documentation alone and says so, rather than implying a benchmark it did not run.
Sources: Profound pricing · About FactCheck · Working with FactCheck · Athena plans · Scrunch pricing · Sitecore/Scrunch · Evertune · Bluefish AI Accuracy · Brandlight · Adobe completes Semrush acquisition · Testing ads in ChatGPT