- Ask "a percentage of what?" before comparing any AI visibility score. The same brand can post wildly different share numbers depending on the denominator.
- Record the surface, not just the model: product, harness and interface change retrieval behaviour on the same underlying model.
- Keep six states separate: searched, fetched, cited, mentioned, recommended, converted. Collapsing them hides the action.
- Vendor scores are fine as inputs when their formula travels with them. Composite scores without formulas are weak evidence.
- Keep first-party evidence (Search Console, GA4, server logs, your own answer ledger) as the base layer, with vendor tools on top.
AI search metrics become useful when every number carries its denominator: the prompt set, the answer surface, the market, the date range and the source state behind it. Track presence, citation, mention, recommendation, referral and conversion as separate events, because each one implies a different fix. A visibility percentage with no denominator cannot tell your team what to build next, which is why so many AI dashboards produce charts and no decisions.
What AI search measurement is actually for
Measurement has one job: to tell you what to change next. Every metric below earns its place by pointing at a specific fix, and any metric that cannot do that belongs in a reference file rather than an operating dashboard.
AI search makes this harder than classic search for a structural reason. A ranking is a stable position on a list. An AI answer is generated, varies between runs, differs across surfaces, and frequently produces no click at all. So AI measurement is sampling, and sampling requires discipline that ranking reports never needed.
Rule one: every number carries its denominator
The most common failure in this category is a percentage with no stated base. "Share of ChatGPT" can mean the share of tracked responses that mention you, the share of all brand mentions across those responses, the share of citations, or the share of a category subset. The same brand can post four very different numbers, each technically correct inside its own calculation, and none of them comparable with another vendor's chart.
So for every reported metric, hold six fields beside it:
- the exact prompt set, and how it was chosen
- the answer surface, including product and interface
- the market and language
- the date range and sampling frequency
- the denominator and formula
- the source state being counted
If a tool cannot supply those, its number is directional at best. That is not a reason to avoid tools; it is a reason to write the fields into your own reporting template.
Rule two: the surface is more than the model name
A row labelled only "Claude" or "ChatGPT" loses the context that explains the result. The same underlying model behaves differently depending on the product wrapped around it: whether it searches the web, which pages it reads, how many sources it uses, and how it structures the answer.
Profound's July 2026 research illustrates the size of the effect. Across 24,135 sampled responses, it reports Claude searching the web in 93% of responses while Claude Code searched in 13%, with the two surfaces mentioning only around 20% of the same brands on average for the same prompt. Treat those figures as first-party vendor research with a stated sample rather than settled fact, and take the structural lesson: record model, product, harness, interface, task type, prompt, market and date. Our breakdown of that study is in Claude SEO.
Rule three: keep the six states separate
These describe different events with different fixes, and blending them into one "visibility" number destroys the diagnosis.
| State | What it means | What a gap here implies |
|---|---|---|
| Searched | The system ran a retrieval step | Your category may not trigger retrieval at all |
| Fetched | Your page was accessed | Crawlability and access are working |
| Cited | The answer displayed your source | Fetched but uncited means content or authority |
| Mentioned | Your brand appeared in the text | Mentioned without citation means third-party sources carry you |
| Recommended | The answer favoured you for the buyer | Mention without recommendation means proof and comparison gaps |
| Converted | A referral produced a business outcome | Needs joined first-party attribution |
A page can be fetched without being cited. A brand can be mentioned without being recommended. A source can be cited for a fact while a competitor gets the shortlist. Each of those is a different week of work.
The metrics worth reporting
Visibility and presence
Presence rate is the cleanest: the share of tracked responses containing your brand at least once. It stays close to the observed event and is easy to explain. Composite visibility scores blend presence, position, prominence and sentiment, so ask for the formula before treating movement as meaningful.
Citation and source share
Which domains the answers cite in your category, and how often yours appears among them. This is the single most actionable metric in AI search, because a cited-domain list is a to-do list for source development.
Position, prominence and sentiment
Where you appear inside the answer, how much space you get, and how you are characterised. Useful context, rarely decisive alone, and worth recording per surface.
Referral and business outcome
AI-assisted sessions and what they do next. Google Analytics now classifies AI Assistants as a channel while treating Google's own AI Overviews and AI Mode under organic search, so referral data undercounts influence by design. Treat AI referrals as a floor on impact, never a ceiling.
Accuracy
The underrated one. Whether answers describe your prices, services and limits correctly. An inaccurate mention can cost more than an absent one, and the fix is usually a clear page fact plus corroborating sources.
Build the answer ledger
Whatever tools you use, keep a first-party record you control. One row per observation:
- Prompt ID and exact prompt text
- Intent class: informational, commercial or comparative
- Surface and run date
- Raw answer text and format
- Searched state and retrieved URLs
- Cited URLs and sources block
- Brand mention, position and recommendation state
- The page or claim that appears to support the answer
- Referral or enquiry where attribution exists
- Confidence and reviewer note
Keep the raw answer. A score without the answer prevents anyone checking whether the metric describes the buyer's real decision, and raw answers are the only record you cannot reconstruct later.
Common mistakes
- Comparing scores across tools with different prompt sets and parsing rules.
- Reacting to single-day movement in a non-deterministic system.
- Letting the prompt set go stale so it measures an old guess about buyer language.
- Reporting AI visibility with no linked action, which turns the dashboard into decoration.
- Treating vendor panel statistics as independent benchmarks. They are first-party claims from that vendor's sample.
How we run this
Inside the Managed Search Loop, AI-answer observation sits beside Search Console, analytics and live SERPs as one evidence stream, and every observation ends in one of two places: a ranked queue item (a page to fix, a source to earn) or an explicit decision to keep watching. Fixes ship through the Agentic Website with approval and live verification, then the same prompt set gets re-observed. The tooling question that sits underneath this is covered in what an AI visibility tracker can and cannot do, and the budget split in depth vs delivery.
FAQs
What are the most important AI search metrics?
Presence rate, citation and source share, recommendation state, accuracy, and AI-assisted referrals with their business outcome. Each needs its denominator and surface recorded to be comparable over time.
How is AI visibility measured?
By sampling: a fixed prompt set run repeatedly across surfaces, with responses parsed for mentions, citations, position and sentiment, then aggregated. Every resulting score is a percentage of tracked responses.
What is a good AI visibility score?
There is no absolute benchmark, because the score depends entirely on your prompt set and competitor set. Read it relatively: your share against named rivals on the same prompts, trending over weeks.
Can I measure AI search in Google Analytics?
Partly. GA4 classifies AI Assistants as a channel while Google's AI Overviews and AI Mode fall under organic search, so referral data understates influence. Use it as a floor and keep a separate answer ledger.
How often should I sample?
Daily if your tool automates it, reviewed weekly, with monthly matched-window comparisons. Answers vary run to run, so trends over weeks are the only reliable signal.
Should I trust vendor AI visibility scores?
As inputs, when the formula, prompt set and sample travel with them. As independent proof of your performance, no; they are measurements from one vendor's panel.
How many prompts do I need to track?
Enough to cover your commercial question space across brand, category, comparison and buying-intent phrasings. Many businesses do real work with 15 to 50, and the set should be refreshed against real query language.
What is the difference between citations and mentions?
A citation is a displayed source reference; a mention is your brand appearing in the answer text. You can have either without the other, and the fixes differ: citations point at your pages, mentions at your third-party sources.
What should I do first?
Write down 15 real buyer questions, run them across two surfaces, and record mentions and cited sources with dates. That manual ledger teaches you more in a fortnight than any dashboard bought before you know what you are looking at.
Turn the numbers into a decision
Every metric here should end in a page action, a source action, or an explicit decision to keep measuring. If a number on your dashboard has never caused any of those three, delete it. Show us the numbers you need to trust, or start with an AI Visibility Audit.