No standard SEO tool currently measures whether AI engines recommend your brand. Ahrefs, SEMrush and Google Search Console were built to track rankings and clicks, and AI recommendation produces neither.
Search Console reports queries that produced impressions and clicks in Google Search. It does not report whether Chat GPT named you when someone asked what tool to buy. Ahrefs and SEMrush track positions in ranked results, which is a different output from a synthesized answer.
Several vendors now offer AI visibility tracking, and some are useful. But they are early, their methodologies vary, and their results frequently diverge from what a user actually sees partly because these systems produce variable output and partly because API access differs from the consumer interface.
The reliable approach right now is manual, and it takes about thirty minutes a month.
One argument for starting now rather than waiting: When proper tooling matures, having six months of your own baseline will be worth considerably more than beginning from zero on launch day. Trend data cannot be backfilled.
What should you measure?
Two metrics carry the weight: appearance rate and share of voice
Appearance rate :
Definition: The percentage of tested prompts in which your brand is named.
If you run 20 prompts across two engines 40 data points and you are named in 12, your appearance rate is 30%.
This is your baseline and the headline number. It is simple, it is intuitive, and it moves in a direction anyone can understand.
Its weakness: It moves for reasons unrelated to your work. Engines change. Categories gain coverage. Everyone’s numbers rise together and you conclude your effort is succeeding.
Share of voice
Definition: Your brand’s mentions as a proportion of all vendor mentions across the same prompt set.
If 40 prompts produce 160 total vendor mentions and 18 of them are yours, your share of voice is 11.25%.
This is the more honest number because it controls for category-level movement. It answers the question a founder actually has: are we gaining relative to competitors, or is the whole category getting more visible?
Why it matters: Appearance rate rising from 8 to 12 out of 40 looks like progress. If share of voice fell from 14% to 11% over the same period, you gained absolutely and lost competitively. Only the second number tells you that.
Report share of voice, Track both.
Three secondary metrics worth recording :
Position in the answer being named first carries more weight than being fifth. Record the ordinal.
Accuracy rate: Of the times you are mentioned, how often is the description correct? Being described wrongly is a different problem from being absent, and often more urgent.
Source concentration: Which sources the engines cite when answering your category’s questions. This is the most actionable data point you will collect, and almost nobody records it.
How do you build the tracking system?
Build a spreadsheet, define a fixed prompt set, and run it monthly under consistent conditions.
Step 1: Define your prompt set:
Fifteen to twenty prompts, covering four types:
Type Count Example
Core buying question 4–5 “Best [category] tool for [specific customer situation]”
Comparison 4–5 “Comp are top [category] platforms for [segment]”
Alternatives 3–4 “[Competitor] alternatives”
Direct / brand 2–3 “Is [your product] good for [use case]?”
Write them in buyer language, not marketer language. Buyers describe situations; marketers describe categories.
Then freeze them Changing phrasing between months destroys the comparison. Write them once, in a document, and use them verbatim.
Step 2: Set the conditions:
- Fresh session every time Prior conversation history contaminates output.
- Same engines every month ChatGPT, Perplexity and Google AI Overviews at minimum
- Three runs per prompt for your most important queries. These systems produce variable output; a single result is noise.
- Same week of the month , so you are not comparing across an engine update.
- Logged out or in a consistent account state. Personalization affects results.
Step 3: Build the sheet:
- Column Purpose
- Date Month of test
- Prompt ID Fixed reference number
- Prompt text Verbatim
- Engine ChatGPT / Perplexity / AI Overviews
- Run # 1, 2, 3
- Appeared (Y/N) | Binary
- Position Ordinal in the answer
- All vendors named In order this feeds share of voice
- Sources cited From Perplexity
- Accuracy notes Any incorrect claims about you
- Screenshot link Evidenc
The all vendor’s named column is what makes share of voice calculable. The sources cited column is what makes the whole exercise actionable rather than merely diagnostic.
Step 4: Calculate:
- Appearance rate = (prompts where you appeared) ÷ (total prompt-engine combinations)
- Share of voice = (your mentions) ÷ (total vendor mentions across all prompts)
- Accuracy rate = (mentions with correct description) ÷ (total mentions)
- Chart all three monthly. Three months of data is where the trend becomes readable one month is a snapshot.
What do the numbers mean?
Benchmarks vary by category, but rough interpretation:
Share of voice Read
Under 5% Effectively invisible. Competitors are being recommended instead.
5–15% Present but not prominent. Usually appearing on specific queries, absent from broad ones.
15–30% Strong. Consistently in the consideration set.
Over 30% Category leader position in AI recommendation.
Context matters more than the absolute figure . In a category with four vendors, 15% is weak. In a category with forty, it is strong. Calculate what an equal share would be 100 divided by the number of credible vendors and compare against that.
How do you track AI referral traffic?
AI referral traffic can be partially tracked in GA4, but it undercounts significantly.
What to set up: In GA4, examine the session source/medium report for referrals from chat gpt. com, perplexity.ai, gemini.google.com and copilot.microsoft.com. Create a segment or exploration grouping these.
Why it undercounts?
Three reasons:
1. Many AI-influenced sessions never show as AI referrals. A buyer reads a recommendation, then searches your brand name directly. That records as branded organic or direct.
2. Not all engines pass referrer data consistently.
3. Google AI Overviews clicks are often recorded as regular Google organic.
The implication: Treat AI referral traffic as a floor, not a measurement. The real influence is larger and mostly invisible.
A better proxy: Track branded search volume against non-branded. Rising branded search with flat non-branded frequently indicates AI-mediated discovery people are learning your name somewhere else and then searching for it.
How should you report this to a board or founder?
Report share of voice, direction of travel, and the competitive gap. Not appearance rate alone, and never rankings.
A workable one-page structure:
1. Share of voice, this month vs three months ago. One number, one trend line. This is the headline.
2. The competitive picture. Your share against your three closest competitors. This is what a founder actually wants to know, and it frames everything else.
3. Where you gained and where you did not? Which prompt types moved? Specific queries usually move before broad ones say so, so that flat broad-query numbers do not read as failure.
4. Accuracy issues. Any incorrect claims found, and what is being done about them.
5. What we did, and what we are doing next. Tied to the numbers above?
What to leave out: Rankings, impressions, keyword counts, and anything requiring the reader to understand retrieval mechanics. If the report needs a glossary, it will not be read.
On honesty in reporting. Include a section on what did not move. Every founder knows not everything works, and a report with no misses makes the hits less believable. This is the single easiest way to build trust in a channel where the mechanics are opaque and the client cannot independently verify most of what you say.
What are the limits of this measurement
Four limitations worth stating rather than hiding.
1. Output is variable : The same prompt produces different answers. Multiple runs reduce noise but do not eliminate it. Treat everything as directional.
2. Personalization affects results: What you see may differ from what a buyer in another location, with a different history, sees.
3. Prompt phrasing changes everything: Your set is a sample of buyer language, not a census of it. Real buyers phrase things in ways you have not tested.
4. Attribution to your work is imperfect: If share of voice rises, you cannot cleanly separate your content work from a competitor’s decline, an engine update, or a third party publishing a roundup that happened to include you.
None of this makes the measurement worthless. It makes it directional which is still considerably better than the alternative, which is no visibility at all into a channel that is influencing your pipeline.
Anyone presenting AI visibility numbers with the precision of a rank tracker is overstating what the method can deliver.
Frequently Asked Questions:
Are there tools that track AI search visibility?
Several exist and some are useful, but they are early, methodologies differ, and results frequently diverge from what users see in the consumer interface. Manual tracking remains more reliable and takes about thirty minutes monthly.
How many prompts do I need?
Fifteen to twenty gives a workable signal. Fewer produces noise; more becomes a burden that stops getting done.
Can I see AI visibility in Google Search Console?
No. Search Console reports Google Search impressions and clicks. AI Overview appearances are not broken out separately.
How often should I measure?
Monthly. More frequent testing produces noise from output variability rather than signal.
Why do results differ between runs of the same prompt?
These models generate variable output by design. Run important prompts three times and record the pattern.
What’s a good share of voice?
Depends on category density. Calculate equal share 100 divided by the number of credible vendors and compare. Being at or above equal share is a strong position.