Methodology
How we measure, and what we refuse to claim
This page exists so you can check us. Every rule below constrains what we are able to report, including several that would make the product look better if we dropped them.
14-day trial, no card required
The rules
These are enforced in the product, not aspirations written on a marketing page.
- 1
Counts always carry their denominator
Every rate is shown as a fraction of the answers behind it. A percentage with no sample size is a claim you cannot check, and at low volume it is usually wrong. Below five answers we show the fraction only, because one run is a draw, not a rate.
- 2
Engines are never averaged together
Each assistant retrieves differently. A blended score hides which one you are losing, which is the only part that tells you what to do. Every figure is per engine, labelled with how it was collected.
- 3
Mentions and citations are different things
Being named in the answer text and being cited as a source are separate states, and a brand can have one without the other. We report them separately because the fixes are unrelated.
- 4
Our failures are excluded, not counted against you
If a probe fails to collect, it leaves the denominator entirely rather than being recorded as an answer that did not name you. An outage on our side must never read as your brand disappearing.
- 5
The same question, every time
Wording changes the answer, so the prompt is held constant. When a question changes we treat it as a new question rather than continuing the old series, because a trend across two different prompts is not a trend.
- 6
Answers are stored verbatim
Every response is kept in full with its citations. Any number in the product can be traced back to the sentence it came from. A metric you cannot audit is a metric you should not trust.
What we do not measure
Stated plainly, because the gaps in a measurement matter as much as its coverage.
We do not measure models that answer without live retrieval, such as Grok, Llama or DeepSeek in their default configurations. What they say reflects training data you cannot influence this quarter, and putting them beside retrieval-based engines would imply a comparability that does not exist.
We also do not estimate how many people saw a given answer. Nobody outside the assistant vendors knows that, and a confident number would be invented. See the engines we do track and how a run works.
Questions about our measurement
Why do you refuse to show a single AI visibility score?
What is the minimum sample before you show a percentage?
How do you decide a brand was mentioned?
Do answers vary if you run the same question twice?
Measure your own brand
The trial runs on the same rules described above. No card required.