ESSAY

ESSAY

·

·

GEO Tooling

GEO Tooling

9 GEO Tools, 4 Measurement Methods

The nine platforms on every 2026 GEO shortlist don't measure the same thing. Sorted by method rather than by feature list, the shortlist gets a lot shorter.

The nine platforms on every 2026 GEO shortlist don't measure the same thing. Sorted by method rather than by feature list, the shortlist gets a lot shorter.

14 min read

In short

In short

Knowledge Company mapped the nine tools on every 2026 GEO shortlist to four different measurement methods: live prompt sampling, model-level statistical sampling, index-scale prompt corpora, and server-side agent analytics. Two tools using different methods will report different visibility for the same brand, and both can be correct. And only one of the nine samples a prompt more than once per cycle — everything else you are shown is a single draw.

Knowledge Company mapped the nine tools on every 2026 GEO shortlist to four different measurement methods: live prompt sampling, model-level statistical sampling, index-scale prompt corpora, and server-side agent analytics. Two tools using different methods will report different visibility for the same brand, and both can be correct. And only one of the nine samples a prompt more than once per cycle — everything else you are shown is a single draw.

Every GEO shortlist in 2026 contains roughly the same nine names, and every comparison sorts them by feature checklist: engines covered, prompts tracked, seats included, does it have an API.

That sorting hides the distinction that decides what the numbers mean. The nine use four different methods to produce a visibility score, and two tools using different methods will disagree about your brand while both are correct.

It hides something simpler too. Eight of the nine either ask each prompt once per cycle or decline to publish the number. One claims to sample repeatedly, and states its own figure two ways on the same page.

Two of the nine are also no longer independent. Semrush is an Adobe property, announced 19 November 2025 and completed 28 April 2026. Scrunch is a Sitecore company, announced 3 June 2026 with terms disclosed by neither party. A comparison presenting these nine as standalone startups is describing 2025.

“The nine” here are the names that recur on 2026 shortlists; the IAB counts more than 20 providers in the category.

The nine, sorted by method and by depth

Tool

Method

Draws per prompt, per cycle

Profound

Live sampling + agent analytics

1 per prompt per engine per day [our arithmetic]

Semrush (an Adobe company)

Clickstream-derived corpus + live prompt tracking

not published

Ahrefs Brand Radar

Index-scale prompt corpus

Index not published; Custom Prompts metered as checks

Peec AI

Live prompt sampling

1, stated outright

Otterly.AI

Live sampling + agent analytics

not published; a binary 1/0 per day, prompt and engine implies one

Athena (AthenaHQ)

Live sampling on a user-set schedule

1 credit = 1 AI response

Evertune

Model-level statistical sampling

100× per model — also written “up to 100x” on the same page

Scrunch AI (a Sitecore company)

Live sampling + delivery to agents

not published

Brandlight

Live prompt sampling

not published

And the entry prices, each as the vendor prints it:

Tool

Entry price

Prov.

Profound

$99/month Starter, month-to-month; $82.50/month billed yearly

[vendor page]

Semrush

Semrush One Starter $199, $165.17 annual-equivalent; standalone AI toolkit $99/mo per domain

[vendor page] + [vendor code]

Ahrefs Brand Radar

Index $199/mo select platforms, $699/mo all platforms; Custom Prompts Basic $50 /mo

[vendor page]

Peec AI

€85 /mo in the EU, $95 /mo elsewhere, month-to-month; €70 / $80 annual

[vendor page] + [vendor code]

Otterly.AI

$ 29 /month Lite; annual tab shows $25 /month

[vendor page] + [archived page]

Athena (AthenaHQ)

Essential Free, $25 free credit / Includes 300 credits; Starter $295/month or $245/month, billed annually

[vendor page] + [vendor code]

Evertune

$800 / month Pro, no billing term stated, demo-gated

[vendor page]

Scrunch AI

Starter $250 per month (billed annually) or $300 month-to-month

[vendor page]

Brandlight

Not published — brandlight.ai/pricing returns HTTP 404

[vendor page]

Prices were read at source on 1–2 September 2026 and quoted as the vendor prints them; the marks are defined in How this was verified. No annual totals appear, because Profound, Athena and Otterly publish annual pricing as a per-month figure and no yearly total anywhere. Every yearly figure quoted for them is somebody’s multiplication.

The column nobody prints

Ask a model the same question twice and you can get two different answers. That is not a temperature setting anyone can switch off. The mechanism is batch non-invariance: your request is batched with other users’ requests, batch composition changes the reduction order inside normalisation and attention kernels, and the logits move. At temperature 0, a thousand completions of one prompt on Qwen3-235B produced 80 unique outputs — identical for 102 tokens, then diverging.

Fixing that requires control of the inference stack, which no tool here has: all nine query hosted endpoints or scrape consumer interfaces. The variance is irreducible for every product in the set — an architectural fact, not a vendor shortcoming.

Which makes draws per prompt the most consequential specification in the category, and the one nobody puts on a pricing page. Peec is the exception, and its worked example is the clearest disclosure in the set: 25 prompts across 3 models for 30 days is 2,250 answers.

The IAB’s August 2026 report on measuring visibility in the AI era states the consequence: “Single-response measurement is not measurement.” Visibility on a query is a distribution, not a value, and a metric built from one response per query is a sample of one.

Nobody knows the right depth yet:

Source

What it says about depth

Schulte, Bleeker and Kaufmann (arXiv:2604.07585)

At least seven runs per prompt per day

Zatuchin (arXiv:2607.13304)

Returns collapse after the fifth

The IAB

No threshold at all

The arithmetic underneath is checkable, though. Take a visibility rate near 25%. Measured over 72 answers, it carries a 95% confidence interval of about ±10 points. Measured over 300 answers, about ±5. That is our binomial arithmetic rather than a vendor claim, and it is why a fifty-prompt single-run reading cannot carry the conclusions drawn from it.

Method 1 — Live prompt sampling

Used by: Peec, Otterly, Semrush, Brandlight, Athena, Profound, Scrunch.

You supply prompts. The platform sends them to ChatGPT, Perplexity, Gemini, Copilot and Google’s AI surfaces on a schedule, and records whether your brand appears, in what position, with what sentiment, and which URLs were cited.

It is a sample, not a census, and the prompt list is the product. Two agencies running the same tool on the same brand with different lists will report different visibility, and neither is lying.

Then check what coverage means on the plan you buy: the marketed engine count and the tracked engine count are frequently different numbers.

  • Semrush. Five engines self-serve — ChatGPT, AI Overviews, AI Mode, Perplexity, Gemini — with Claude, Copilot, Grok and DeepSeek on Enterprise AIO. But the metered daily Prompt Tracking feature supports three: ChatGPT Search, AI Mode and Gemini, desktop only. Perplexity sits in the weekly Brand Performance database instead, and Semrush’s own knowledge base does not agree with itself on whether Prompt Tracking reaches AI Overviews.

  • Otterly. Six engines read from public web interfaces; Claude tracked via API on Sonnet with web search, which makes its Claude series methodologically different from the rest of its own dataset.

  • Peec. UI scraping for six consumer surfaces, with the five Enterprise models tagged API on its own pricing table.

  • Athena. Credits are the unit and are not uniform: six engines cost one credit per response, five cost five. Athena states both sides — its plans table says a credit is one AI response, its calculator says the heavier models cost five each.

  • Profound, Scrunch, Brandlight. Profound captures the consumer experience rather than API output, across three pipelines. Scrunch mixes browser automation with official platform APIs per platform. Brandlight publishes no technical documentation, only a marketing description.

Five of the nine market consumer-interface capture as what makes their data real. Whether that survives platform terms of use we cannot say here: OpenAI’s terms page returned HTTP 403 to us this pass.

Method 2 — Model-level statistical sampling

Used by: Evertune.

Same idea, different statistical commitment: each prompt sampled many times per model rather than once, with prompt selection grounded in a large consumer panel. Both halves need qualifying, and both qualifications come from Evertune’s own pages.

The depth is stated four ways:

Where it appears

The figure given

Homepage card

100 times across every AI model

Homepage product copy

“up to 100x per model”

FAQ

100 times per model across 11+ models

Docs

100 times per model

“Up to” is a materially different commitment, and this is the company’s central differentiator.

The model count is a count of surfaces. The eleven separate ChatGPT from ChatGPT Search and Gemini from Gemini Search, and include AI Mode and AI Overviews, which Evertune’s own docs describe as living inside Google Search rather than standing alone as LLMs. Roughly six model families are involved, and the FAQ’s “11+” enumerates exactly eleven.

The panel shifts unit across pages:

Page

What it counts

Homepage

over 150 million user prompts

Methodology page

over 150 million real conversations

Company overview

150 million prompts

How it is sourced, recruited, consented or compensated is not published.

The base-model differentiator carries a caveat too: the foundational-knowledge measurement for ChatGPT runs on ChatGPT-5.4-mini, a mini-tier model per Evertune’s own docs table, not the flagship.

Even so, this is the only tool in the set that treats a visibility figure as a distribution.

Method 3 — Index-scale prompt corpora

Used by: Ahrefs Brand Radar; partially by Evertune’s panel.

This method starts from a large corpus of prompts people actually send. You lose control of the prompt list and gain the ability to discover demand you would never have thought to track.

Two things about it are widely misreported. The first is that the corpus has no single size. On the same day, Ahrefs publishes:

Where

Corpus size

Brand Radar

467M+

AI Visibility Index

468M+ — its own per-platform table there sums to 468,564,875

Pricing

475M+

Methodology post

a table summing to roughly 376M

The second is that the pricing is two products. The Index is $199/mo for selected platforms and $699/mo for all of them. The much-quoted bundle belongs instead to the $50/mo Custom Prompts Basic package:

Custom Prompts Basic

Allowance

Prompts

83 a day

Checks

2,500 monthly

Overage

$0.020 per check

On the rendered page that list sits in the adjacent column, and the $199 card carries no feature list at all.

Coverage has two holes worth knowing before you count engines. Claude is not in the Index — custom prompts only, tracked via API, consuming eight checks per update. Grok is listed as supported, but collection is suspended following platform policy changes.

Method 4 — Server-side agent analytics

Used by: Profound and Otterly. Scrunch sits adjacent rather than inside this method — it formats and delivers content to agents rather than observing them.

The other three methods observe the engine from outside. This one observes it from your own server: which AI crawlers fetched which URLs, how often, and what they did next.

Otterly gates agent analytics above the entry tier. Profound sells agent credits alongside prompt tracking, with a choice between overage billing and pausing usage at the limit. Scrunch extends the idea into delivery: its Agent Experience Platform, which Sitecore describes as delivering content formatted for LLMs “in a way that AI agents can read and use.”

It is the only one of the four methods based on observed fact rather than sampling. It cannot tell you what an engine said about you. It can tell you, without statistical caveat, exactly how many times OAI-SearchBot fetched your pricing page last week.

What counts as a mention

The plainest problem in the category is that a “mention” has no shared definition. A linked citation, an unlinked brand name and a paraphrase of your positioning are three distinct events, and no vendor in the set documents how it counts them.

It got harder in 2026. OpenAI began showing labelled sponsored placements to logged-in US free and Go users on 9 February 2026 and opened a self-serve ads manager on 5 May 2026. If a sponsored block renders inside a surface these tools capture, a mention count may be counting a paid placement. We found no disclosure from any of the nine on whether it separates sponsored from organic, or whether it captures logged-in or logged-out sessions — which now decides whether the ads are visible to the scraper at all.

Why two tools disagree about your brand

  • Different prompt populations. Yours, versus a panel’s, versus an index’s.

  • Different sample depth. One draw per prompt versus a hundred, and in four cases a depth the vendor does not publish.

  • Different engine mixes. Athena’s $295 Starter reaches ten models, metered at one or five credits each. Profound’s two tiers differ on engines, prompts and responses at once.

  • Different definitions of a mention.

Profound’s two tiers, as published:

Tier

Engines

Prompts

Responses a month

$99 Starter

ChatGPT only

50 unique

1,500

$399/month Growth

Three

100

9,000

The Growth figure is the month-to-month price; billed yearly it is $332.50/month.

One further trap is arithmetic rather than method. Dividing entry price by prompt allowance produces a per-prompt figure that looks comparable and is not: Profound’s fifty Starter prompts run on one engine and Peec’s fifty run on three, Ahrefs meters checks, Athena meters responses with a five-times engine multiplier, and Scrunch meters custom prompts, industry prompts, personas and page audits separately. The tidy “$1 to $2 per tracked prompt” this category gets summarised with describes nothing; the pricing structures repay reading directly.

The discipline that follows: never compare a number from one tool to a number from another. Compare each tool’s series to its own history, and settle a testing protocol before you trust either.

Which method answers which question

Your question

The method that answers it

Are we visible for the 40 prompts our sales team actually hears?

Live prompt sampling (Method 1)

Is our position stable, or did we get lucky on Tuesday?

Model-level sampling (Method 2)

What are buyers asking that we’ve never tracked?

Index-scale corpora (Method 3)

Are AI crawlers reaching our real pages?

Agent analytics (Method 4)

Most teams need Method 1 plus Method 4, and buy neither 2 nor 3 until a board deck forces the question. The free and no-signup layer is where the cheapest real evaluation lives.

What none of the four methods do

All four answer versions of are we present, and where. None of the four methods answers is what the engine said about us true. Two vendors sell a separate feature that tries — Profound’s FactCheck and Athena’s Oracle — and both rest on ground truth the customer maintains.

A brand can score in the ninetieth percentile for visibility while every engine quotes a price it stopped charging fourteen months ago. Sentiment analysis will read that answer as positive. Share of voice will count it as a win. That gap is the subject of what GEO tools do not check.

How this was verified

Provenance marks: [vendor page] the vendor’s own page, read at source · [vendor code] its shipped price module or component props · [vendor docs] its help centre or docs · [press release] a vendor or newswire release · [secondary] reputable secondary · [archived page] an Internet Archive snapshot · [not published] not published · [our arithmetic] our arithmetic on vendor numbers, not a vendor statement.

What changed since the 1 September version:

Profound’s billing basis was inverted. $99/month is the month-to-month price; “Billed yearly” is a per-card toggle that ships off and shows $82.50/month when switched on. Both terms are published. No annual total is.

Semrush moved from unconfirmed to primary. Its monthly prices and their annual-equivalents are static literals in its own price-table module, and also appear in the served HTML:

Monthly

Annual-equivalent

$199

$165.17

$299

$248.17

$549

$455.67

Peec resolved. European visitors are shown euro prices, month-to-month against annual:

Month-to-month

Annual

€85

€70

€205

€180

€425

€360

Elsewhere the page shows dollars:

Month-to-month

Annual

$95

$80

$245

$205

$495

$420

Pages describing €85 as the annual price have it the wrong way round.

The Ahrefs feature bundle was reassigned from the $199 Index tier to the $50 Custom Prompts package, settled by rendered column geometry rather than extracted text.

Brandlight was downgraded to “not published”. Its pricing page returns HTTP 404 and its homepage carries no figures; the ~$199/mo price in circulation traces to aggregators and to a vendor-owned but programmatically generated subdomain. Its homepage names six engines and publishes no total.

Scrunch and Evertune both now publish an entry price on the page. Evertune’s replaced an earlier “starts at $3,000 per month” listing; Scrunch’s page was restructured. Scrunch’s own FAQ still names the retired Core tier, so the pricing page is the citation. A September read lists Starter as:

Scrunch Starter

Allowance

User licences

Three

Custom prompts

350

Industry prompts

1,000

Personas

Three

Page audits

Five

Platforms

Six

That read came through a summarising fetch — treat the tier contents as indicative, the prices as read.

Deliberately not printed:

  • any annual total;

  • the claim that three engines agree on a brand only 21% of the time, which comes from one vendor’s experiment at 150 answers whose own page concedes a single run per prompt with no temperature control;

  • and “GEO lifts visibility up to 40%”, whose underlying paper is peer-reviewed but measures re-ranking among five already-retrieved sources fed to a 2023-era model.

Still open:

  • whether Athena’s 300 free credits recur;

  • whether Evertune’s “100,000 prompts tracked” counts distinct questions or total responses;

  • Peec’s agency price list, not captured at source.

Unverified numbers propagate here. Most GEO listicles reproduce each other’s tables, and a price that changed in March is still quoted in September because nobody re-read the page. If you will make a purchasing decision from a table, you should know which cells were read and which were inherited. Every price and capability above was read at source by Knowledge Company on the dates tagged, not carried across from another comparison table.

Sources: Profound pricing · Semrush pricing · Semrush KB 1503 · Adobe completes the Semrush acquisition · Ahrefs AI Visibility Index · Ahrefs Custom Prompts · Ahrefs Brand Radar help · Peec pricing · Otterly pricing · Otterly engine coverage · Athena plans · Athena credit calculator · Evertune FAQ · Evertune methodology · Evertune models available · Scrunch pricing · Sitecore on the Scrunch acquisition · Brandlight · IAB, Measuring Visibility in the AI Era, August 2026 · Defeating nondeterminism in LLM inference · OpenAI, testing ads in ChatGPT

Summarize and read with