Someone asked ChatGPT for the best tool in your category this morning. It gave them three names and linked four sources.
You have no idea whether you were one of the names. You have no idea which pages it read to decide. And if it read your page and recommended a competitor anyway, that is the most useful thing you could possibly learn about your marketing this quarter, and there is no report in your analytics that will ever show it to you.
That gap is what we built AIMentionTracker to close.
AIMentionTracker is an AI visibility tool that tracks whether AI assistants name your brand when someone asks about your category, and which pages they cite when they answer. Six assistants, on a schedule, with every answer kept in full.
It is live today. Here is what is in it, what it refuses to do, and what it costs.
Why this stopped being optional
For twenty years the job was a position in a list. Rank first, earn the click.
That arrangement is coming apart, and there is now hard data on it instead of speculation. Pew Research Center tracked real browsing behavior in March 2025 and found that users who encountered an AI summary clicked a traditional search result in 8% of visits, compared with 15% for users who did not. Roughly half as often. They clicked a link inside the summary itself in 1% of visits.
Ahrefs measured the same effect from the other direction. Across 300,000 keywords, they found a 34.5% lower average click-through rate for the top-ranking page when an AI Overview was present.
Read those together and the conclusion is uncomfortable. The answer is becoming the destination. Your best page can rank first, be read by a model, be summarized for someone who never scrolls, and leave no measurable trace anywhere you currently look.
Whether you are named inside that answer is now a separate question from where you rank. Almost nobody measures it honestly.
What actually happens when you run it
Not a feature list yet. The mechanics, because the mechanics are the product.
You pick prompts, not keywords. Whole questions, phrased the way a person types them into an assistant. Paste your domain and we read your site, then propose a category, up to six competitors you genuinely lose deals to, and eight to twelve prompts a buyer types before they know your name. You edit all of it. Each suggestion arrives with real Google search volume beside it, because nobody can measure true AI prompt volume and we would rather show you a proxy and name it than invent a number.
Every prompt runs against every engine on a fixed schedule. Same wording, same market, every time, so a change in the answer means something changed in the world rather than in your phrasing. Five engines run six days a week and Claude runs twice, because Claude costs seven to ten times more per answer. We would rather tell you that than quietly bill for it.
Every answer is stored whole. Not a score derived from it. The full text, which brands were named and in what order, and every URL cited. You can open any number on any chart and read the answers underneath it. That sounds obvious. Most tools in this category discard the text and ask you to trust the percentage.
Our failures leave the denominator. If collection breaks on our side, that run is excluded rather than counted as you being absent. A tool that logs its own outages as your invisibility is measuring itself.
The seven things it measures
Visibility tracking. How often each engine names you, as a rate with its sample size attached. Every answer lands in one of four states, kept apart deliberately: named and cited, named but not cited, cited but not named, or neither. Collapsing those into one mention percentage is how most tools lose the only signal that tells you what to do next. That is AI visibility tracking.
Citation tracking. Which pages the assistants actually read to build the answer, with the domains that keep appearing ranked. This is where you find out that the pages briefing your category are mostly not yours, and it flags every answer that cites your page while naming somebody else.
Sentiment analysis. Being named is not being recommended. Three ratings per mention, each shown with the sentence that produced it, so you can check the rating rather than take it on faith. You will occasionally find an assistant naming you accurately and describing you in a way that loses the deal.
Share of voice. Your share of the answers against whoever is named beside you, reported three ways: per category, per engine, per prompt. The per-prompt number is the one you can act on, because one prompt where a rival owns the answer is a page you can go write. We record whoever turns up, not only the competitors you told us about.
Prompt tracking. The tracked set itself, with search volume beside each one, and the discipline that a rewritten prompt starts a new series instead of continuing an old one. A trend measured across two different prompts is not a trend.
Brand monitoring. Six kinds of change watched for and emailed once each: a visibility drop against its recent baseline, a rise worth knowing about, an engine blackout where one assistant stops naming you while the others carry on, a competitor overtaking you on a prompt you were holding, an answer citing your page without naming you, and a sentiment shift. Findings inside their own margin of error are suppressed, so a quiet week sends nothing. An alert you learn to ignore is worse than no alert at all.
Answer archive. Every answer, searchable and filterable by engine, by prompt, and by whether you were named. Finding every answer where one particular competitor appeared beside you is the kind of query that turns a dashboard into an argument you can take into a meeting.
All seven ship on every plan. None of them is held back for a higher tier.
Three rules we will not bend
These are why the product sometimes looks less impressive than its competitors in a screenshot. A number with its work shown is a smaller number.
Every rate carries its sample size. Assistants generate a fresh answer every single time, so one check is a coin flip. Below five checks you see "2 of 4", never "50%", because four runs cannot support that precision. Around thirty checks per prompt per engine is where a real change starts separating itself from ordinary variance.
Engines are never blended. They retrieve differently and they disagree. We have measured Google AI Overviews and Google AI Mode naming different brands for the same prompt on the same day, and those two are both Google. A blended score averages that away and leaves you knowing only that you have a problem somewhere. So we report every engine on its own.
Cited and named stay separate. Your page in the sources while a competitor sits in the sentence is not absence. It means retrieval already works and your framing is losing, which is a different repair job and a much cheaper one. Citation tracking reports it as its own outcome, and it is the finding that surprises new users more than any other.
The full rules are on the methodology page, in public, before you pay us anything.
What it will not do
No single blended visibility score. No estimate of how many humans saw an AI answer, because nobody outside the assistant vendors knows and a confident figure would be fabricated. No claim that schema markup moves AI Overviews, because Google states plainly that there is no special schema.org markup you need to add for generative AI search.
Then a list of things that simply are not built yet. No REST API. No MCP server. No white labeling or client-facing branded reports. No historical backfill, so we cannot tell you what ChatGPT said about you in June. Measurement starts the day you do.
We also do not track Grok, Llama or DeepSeek. They answer from training data without live retrieval, which makes a rate on them a different measurement wearing the same label. Saying so costs us a row on somebody's comparison table. We would rather be right.
Three problems, and telling them apart
Nearly every visibility failure is one of three things, and they need opposite responses.
The assistant cannot reach you. A crawling problem, usually one robots.txt line nobody audited. Each assistant uses named crawlers: Google publishes its full list including Google-Extended, OpenAI documents its crawlers separately by purpose so blocking the training bot and the search bot are different decisions, and Perplexity documents PerplexityBot. Google also warns that nosnippet, max-snippet and noindex restrict how your content appears in its AI experiences. This tier is binary, cheap to fix, and disqualifies more sites than anyone expects.
It reaches you and does not cite you. A depth problem. The page touches the topic without answering the question the way a model needs it answered.
It cites you and names somebody else. A framing problem, and the one people most often misdiagnose as needing more content. More pages give the model more to retrieve from a source it already picked. Rewriting the one paragraph being cited beats publishing ten new posts.
A mention percentage tells you none of this. The citation list does: never in the sources on any answer, in them occasionally, or in them constantly while somebody else gets named. That is three different readings of the same stored data, and it is why we keep the sources rather than only counting names.
What it costs
Four plans: $9, $19, $29 and $49 a month. Yearly is ten months' price.
One thing changes between them: how many prompts you track. Two, five, nine or fifteen. Every plan includes all six engines, full answer text with citations, five brands, twenty-five competitors, twenty-five teammates, email alerts, and any of 212 markets in 128 languages.
Compare that to how this category usually prices. Otterly's $29 plan carries four engines and sells Gemini and Google AI Mode at $9 a month each plus Claude at $29, so the full set runs $76. Peec caps you at three models per project below its enterprise tier. Profound publishes no paid price at all. The full comparison of nine tools is on the site, including where we come off worse.
Seven days free. A card is required and nothing is charged until the trial ends, which we print on the button instead of in the small print.
If you would rather see how we stack up before reading any of that, the nine-tool comparison names its criterion up front.
Who should not buy this
Worth being direct, because a refund on day two helps nobody.
If you need one number for a board deck, you will find this frustrating. The product is built to refuse that.
If you want agent traffic analytics, shopping visibility or content generation, none of it is here.
If you need history from before you signed up, we cannot manufacture it.
And if two prompts sounds too small to learn anything from, you are right in one sense. Two prompts is 256 answers a month, which is plenty of signal on those two, but it is a corner of your category rather than the category itself. Start higher if you want the shape of the whole thing.
The honest state of things
It is new. The company was incorporated this month. I built it because I wanted it, could not find one that showed its work, and got tired of dashboards printing a confident percentage over four runs.
The site you are reading is the proof. We sell AI visibility, so if assistants cannot read, parse and cite this site then the product's own claim is refuted by its homepage. Every page is real HTML at build time, robots.txt welcomes every crawler by name, and there is an llms.txt stating what the product does in terms a model can quote directly.
If you want the concept rather than the announcement, start with what AI visibility is. To see how the six assistants differ in practice, read how AI assistants decide who to name. And the plans are here when you want to find out whether the assistants are recommending you or somebody else.
No funding announcement. No waitlist. Nothing revolutionized. Just a tool that shows you what the assistants said, and the arithmetic behind every number it prints.

