"AI visibility" gets used for at least three different things: how much traffic assistants send you, how often your name appears anywhere near an AI product, and whether a model recommends you when someone asks. Only the third is measurable with any rigour, and it is the one worth the name.
This post is the definition we work from, the four outcomes it splits into, and the rules that stop the number being flattering nonsense.
AI visibility, defined
AI visibility is the rate at which an assistant names your brand in its answer to a question someone in your category would actually ask.
Three words in that sentence are doing work.
Rate, not score. A percentage with a denominator attached, not an index number scaled to 100 by a vendor who will not show the arithmetic.
Names, not mentions. The brand appearing in the sentence a person reads, as distinct from the URL list underneath it. Those are different events and we report them separately.
An assistant, singular. ChatGPT is not Gemini. Measuring both and averaging them produces a figure that describes neither.
The reason any of this became a business problem rather than a curiosity is visible in the click data. Pew Research Center followed real browsing behaviour in March 2025 and found that users who met an AI summary clicked a traditional search result in 8% of visits, against 15% for users who did not, roughly half as often. They clicked a link inside the summary itself in 1% of visits. Ahrefs, looking at 300,000 keywords, measured a 34.5% lower average click-through rate for the top-ranking page when an AI Overview was present.
So the answer is increasingly the destination. If you are not in it, the ranking you worked for is being read to someone who never scrolls far enough to see it.
Three things AI visibility is not
It is not rank tracking. A Google result is retrieved from an index and is stable enough that checking once tells you something true. An AI answer is generated, and generated fresh. Two runs of the same question an hour apart can name different brands with nothing having changed in the world. Everything downstream of that difference, including how often you have to check, follows from it.
It is not brand monitoring. Social listening tools watch what humans publish. AI visibility watches what a model says when asked. A brand can be written about heavily and still be absent from the answers, usually because the pages describing it are not the pages being retrieved. The citation data is where that gap shows up.
It is not AI referral traffic. Analytics can show you sessions with a ChatGPT or Perplexity referrer, and that number is worth watching, but it measures only the people who clicked. The Pew figures above are the reason it undercounts badly: most people who read an answer about your category never click anything. Referral traffic is the visible tip of an outcome that mostly happens off your property.
The outcome splits four ways, not two
Most tools report one thing: were you mentioned. That collapses two independent questions into one, and the collapse hides the only part that tells you what to do next.
Every answer puts you in one of four states:
Named and cited. You are in the sentence and your page is in the sources. This is the outcome you want and the one to protect.
Named, not cited. The model recommends you from what it already knows, without reading you today. Pleasant, and fragile: nothing you publish is holding that position up.
Cited, not named. Your page is in the source list and a competitor is in the sentence. The model read you and recommended somebody else. This is the most useful finding in the category and the one people most often misdiagnose, because it looks like an absence and is actually a framing problem. More content will not fix it. Different framing on the page being read might.
Neither. You are not in the answer or the sources. That is a retrieval problem, and the first thing to check is whether the assistant can reach your site at all. Google publishes the list of its crawlers including Google-Extended, and the equivalent for other assistants is a robots.txt question most teams have never audited.
Keeping these four apart is the whole reason our visibility tracking reports them as separate states rather than one percentage.
One check tells you nothing
This is the rule that most vendor screenshots quietly break.
Because answers are generated rather than retrieved, a single run is a coin flip. Ask "best project management software" three times and you can get three different orderings. A screenshot of the one run where you appeared is not evidence, and neither is the one where you did not.
What produces a defensible number is repetition with the wording held constant. Our rule, stated in full on the methodology page: every rate carries the count behind it, and below five checks we show the fraction rather than a percentage. "Named in 2 of 4" is honest. "50%" from four checks implies a precision four checks cannot support.
Thirty checks on a fixed question, per engine, is roughly where a real change starts to separate itself from ordinary variance. That is also why changing the wording of a tracked question restarts the series: a trend measured across two different prompts is not a trend.
Engines disagree, and the disagreement is the signal
The assistants retrieve differently and they answer differently, which is why we report all six separately and never blend them.
ChatGPT runs its own search product. Anthropic added web search to Claude in 2025, and Claude cites fewer sources per answer while being noticeably more willing to say it does not know. Google runs two surfaces with separate retrieval: AI Overviews, and the AI Mode tab launched in 2025. We have measured those two naming different brands for the same question on the same day.
A single blended "AI visibility score" averages those disagreements away. If you are strong on Perplexity and invisible on Gemini, the blended number is a middling figure that tells you to do nothing in particular.
What actually moves it
Less exotic than the category implies. Google's own guidance for site owners says structured data is not required for generative AI search and there is no special schema.org markup to add, and that its generative features sit inside the core ranking systems rather than beside them.
What is left is ordinary, in this order:
- Be retrievable. Indexed, crawlable, not blocked by a directive nobody audited. This is binary and it disqualifies more sites than people expect.
- Answer the question in liftable prose. Assistants assemble from passages. A page that answers "how much does X cost" in one clear paragraph gets lifted; the same answer spread across a pricing table and three footnotes does not.
- Fix how third parties describe you. The pages cited in your category are mostly not yours. Comparison sites, roundups and forums brief the model before your homepage does.
Semrush's study of AI search traffic patterns is a reasonable read on how that traffic behaves once it arrives, though treat every volume figure in this category, ours included, as a proxy.
Where to start
Pick ten to thirty questions a buyer types before they know your name. Run them on each assistant on a fixed schedule, unchanged. Store every answer in full, with its citations, so any number you later quote can be opened and checked. Report per engine, with denominators.
If you want the mechanics rather than the definition, the how to track Google AI Overviews guide covers one surface end to end, and how to rank in AI Overviews covers what moves the result once you can see it. Our plans start at two prompts, which is a floor rather than a starting point, and every plan runs all six engines.
The measurement is not complicated. It is just repetitive, and that is the part people skip.

