You cannot improve a number you have never measured, and there is no report that hands you this one. Tracking brand mentions in AI search means building the measurement yourself: a fixed sheet of questions, run against each engine on a schedule, with the answers recorded rather than eyeballed.
This is the practical half of AI search visibility. That guide defines what is being measured and why no dashboard reports it. This one is the procedure.
What you are actually tracking
Three separate things, and conflating them is the most common mistake in this work.
Named. The answer says your brand name.
Cited. Your name carries a link back to your site, or your domain appears in the source list.
Quoted. Your specific wording, a number you published, or a framing that is recognisably yours turns up in the answer, whether or not you were credited.
Quoted without cited is the interesting case and it is more common than people expect. It means your page was good enough to use and not distinct enough to attribute. That is a fixable problem, and you would never see it if you only counted links.
Step 1: build the prompt sheet
Twelve to twenty questions, in the words a buyer would use, not the words you use internally. Cover four kinds:
Category questions. “Best tool for bulk checking whether pages are indexed.” This is where you find out whether you exist in the model’s picture of your market at all.
Problem questions. “Why are my pages crawled but not indexed.” The largest group, and usually where the first mentions arrive, because these have a factual answer somebody has to have written.
Comparison questions. “X or Y for a small site.” Comparison content is disproportionately quoted by answer engines because it is already structured as a decision.
Identity questions. “Who is [your name], what does [your brand] do.” A blunt test of whether the engines know you exist. Zero here with mentions elsewhere means your content is being used but your identity is not attached to it.
Write them in a spreadsheet with one row per prompt. That sheet is your instrument, and changing its wording invalidates every comparison you have made so far. If you must add prompts, add them as a new block with its own start date.
Step 2: run it cold
Fresh session, signed out where you can be, no history, no follow-up questions, no location personalisation you can avoid. Personalised context is the single fastest way to measure yourself as more visible than you are, because the systems that remember you have already been told you matter.
Run the identical sheet against each engine you care about. For most sites that is ChatGPT, Google AI Overviews, Google AI Mode, Gemini and Perplexity. Add Claude if your buyers are technical.
Do not ask follow-ups. A follow-up changes the question, and in AI Mode it counts as a new query entirely, which is one of the reasons the Search Console reporting is so hard to read.
Step 3: record five fields per answer
One row per prompt per engine per run:
- Date and engine.
- Named yes or no.
- Cited yes or no, with the URL if there is one.
- Quoted yes or no.
- Who was cited instead, the two or three sources that got the position you wanted.
Field five is the one people skip and the one that pays for the whole exercise. Those sources are the answer to what the engine considered good enough. Open them. In almost every case the difference is not authority, it is that they stated the specific thing in a liftable sentence and you buried it.
Add a sixth column for a one-line note when something surprises you. Six months later, that column is the most useful part of the sheet.
Step 4: score it, simply
Two numbers per engine per run: mention rate, and citation rate. Mentions divided by prompts, citations divided by prompts. That is enough. A weighted composite score feels more serious and tells you less, because when it moves you cannot see which half moved.
Report them separately per engine. Never average across engines. A site cited constantly by Perplexity and never by AI Overviews has a specific, diagnosable situation, and a single blended score erases it.
Step 5: repeat monthly
Same sheet, same wording, same cold conditions, same day of the month if you can. Monthly is the right cadence for almost everyone. Weekly mostly measures the non-determinism described in the FAQ below, and a run that takes an hour by hand will quietly stop happening if you schedule it weekly.
Two runs is not a trend. Give it three or four before you conclude anything about direction.
Tracking mentions in ChatGPT specifically
ChatGPT is the engine most people mean when they ask about this, so it is worth its own notes.
It answers from two different places, and they behave differently. Some answers come from what the model already holds, with no retrieval at all. Others trigger a live search, and those return sources you can click. The same prompt can do either depending on wording, so record which one happened. If your prompt sheet only ever produces the non-retrieval kind, you are measuring the model’s memory of the past, not your current site.
The live half depends on ChatGPT’s crawlers being able to fetch you. This is worth checking before you conclude anything about your content, because it produces exactly the same symptom as being unquotable: whether robots.txt blocks ChatGPT covers the user agents involved and which one does what. More than one site I have looked at was blocking the retrieval bot while wondering why it was never a source.
The tools sold for this are, in the main, this same procedure run at scale against a prompt list you supply. That is a genuine convenience once your sheet is long. It is not a different measurement, and no tool sees inside the model.
Tracking mentions in Gemini and AI Overviews
These two sit on the same index, which makes them the cheapest pair to track and the most connected to ordinary SEO work.
Because retrieval here runs through Google’s index, ranking for the query is close to a prerequisite. If you are not in the index, or not on the first page for the phrasing you are testing, a zero tells you about your rankings rather than about your AI visibility. Check the ranking first. If the page is not indexed at all, that is the actual bottleneck and nothing on this page will move until it is fixed.
There is now a generative AI performance report in Search Console, which counts impressions where links to your site were shown in a generative feature. It is worth watching, and it does not replace this procedure: it has no clicks, no query dimension worth the name, and no way to tell you which prompt produced the impression or who was shown alongside you. What the new report can and cannot settle goes through it in detail.
If you would rather not be summarised at all, that is a legitimate position with its own controls, covered in keeping content out of AI Overviews.
What to do with the sheet once you have it
The point of tracking is the diagnosis, not the dashboard. After two runs, sort your rows by the fifth column and read the sources that beat you. You will find one of four situations, and each has a different fix.
You are not fetchable. Zero across the retrieval engines, fine in Google. Access problem. Check the crawler rules per vendor: Claude and Perplexity each run more than one agent, and Google-Extended does something different again.
You are fetchable and unquotable. You rank, you are crawled, you are never used. The answer is somewhere in paragraph nine of your page and the sources beating you put it in the first sentence under a heading that matches the question.
You are quoted and not credited. The wording is yours and the link is not. Usually an attribution problem: no named author, no evidence trail, nothing that makes you safe to name.
You are absent from the category entirely. Nobody in your market is being asked about you. That is a positioning problem before it is an SEO one.
What to do about each is how to improve brand visibility in AI search engines. Then re-measure next month and see which of those moved. That loop, run for six months, is worth more than any single optimization you could make today.
If you want the sheet built and run against your site rather than the instructions for building it, that is AI search optimization.
Quick answers
What counts as a brand mention in AI search?
Three things worth counting separately. A named mention is the model saying your brand name in its answer. A cited mention is your name attached to a link back to your site. A quoted mention is your actual wording reproduced, whether or not you were named. They have different value and they move independently, so a tracker that collapses them into one score hides the thing you want to see. Count all three per answer and you will notice, for example, that your sentences are being used while somebody else's site is getting the link.
How many prompts do I need before the numbers mean anything?
Ten to twenty is enough to start, and the size matters less than keeping the sheet fixed. These systems are non-deterministic, so a single prompt run once tells you almost nothing and the same prompt run monthly for six months tells you a great deal. Adding prompts later is fine as long as you track the new ones as a separate group rather than folding them into an existing average, which would make the trend line lie.
Do I need a paid tool to track this?
Not to start, and starting by hand is better anyway because you read the answers rather than a chart of them. A tool earns its price at the point where the sheet is long, the cadence is weekly, and you want five engines covered without spending an afternoon on it. What no tool does is tell you why you were left out, which is the part that actually changes your plan.
Why do the same prompts give different answers on different days?
Because sampling is built into how these models generate text, and because the retrieval layer underneath several of them re-runs a live search each time. Two runs a day apart can pull different sources with nothing about your site having changed. This is why the method here fixes the wording, uses cold sessions, and compares months rather than days. Treat any single-run movement as noise until a second run agrees with it.
Can I track mentions in ChatGPT and Gemini the same way?
The recording is identical, the running is not. Gemini and AI Overviews both sit on Google's index, so the same query can be checked in both and the results usually rhyme. ChatGPT and Perplexity retrieve through their own crawlers, so their answers can diverge sharply from anything you see in Google. Keep one column per engine rather than one score across all of them, or a strong result in one place will paper over a zero somewhere else.
