Methodology

How we measure, and what we refuse to claim

This page exists so you can check us. Every rule below constrains what we are able to report, including several that would make the product look better if we dropped them.

14-day trial, no card required

The rules

These are enforced in the product, not aspirations written on a marketing page.

  1. 1

    Counts always carry their denominator

    Every rate is shown as a fraction of the answers behind it. A percentage with no sample size is a claim you cannot check, and at low volume it is usually wrong. Below five answers we show the fraction only, because one run is a draw, not a rate.

  2. 2

    Engines are never averaged together

    Each assistant retrieves differently. A blended score hides which one you are losing, which is the only part that tells you what to do. Every figure is per engine, labelled with how it was collected.

  3. 3

    Mentions and citations are different things

    Being named in the answer text and being cited as a source are separate states, and a brand can have one without the other. We report them separately because the fixes are unrelated.

  4. 4

    Our failures are excluded, not counted against you

    If a probe fails to collect, it leaves the denominator entirely rather than being recorded as an answer that did not name you. An outage on our side must never read as your brand disappearing.

  5. 5

    The same question, every time

    Wording changes the answer, so the prompt is held constant. When a question changes we treat it as a new question rather than continuing the old series, because a trend across two different prompts is not a trend.

  6. 6

    Answers are stored verbatim

    Every response is kept in full with its citations. Any number in the product can be traced back to the sentence it came from. A metric you cannot audit is a metric you should not trust.

What we do not measure

Stated plainly, because the gaps in a measurement matter as much as its coverage.

We do not measure models that answer without live retrieval, such as Grok, Llama or DeepSeek in their default configurations. What they say reflects training data you cannot influence this quarter, and putting them beside retrieval-based engines would imply a comparability that does not exist.

We also do not estimate how many people saw a given answer. Nobody outside the assistant vendors knows that, and a confident number would be invented. See the engines we do track and how a run works.

Questions about our measurement

Why do you refuse to show a single AI visibility score?
Because it would be the most saleable and least honest thing on the screen. The engines disagree with each other, often sharply, and averaging that disagreement produces a number that describes no real surface. If one number is genuinely what you need, the per-engine figures are all there to combine yourself, with the sample sizes attached so you know what you are combining.
What is the minimum sample before you show a percentage?
Five answers. Below that we show the fraction, such as 1 of 3, rather than 33 percent. The rounding is not the problem; the implied precision is.
How do you decide a brand was mentioned?
We match the brand name and its known variants against the answer text, then record the position of the first mention. Ambiguous cases, where a brand name is also a common word, are resolved with the surrounding context rather than a bare string match, and the full answer is stored so you can check any call we made.
Do answers vary if you run the same question twice?
Yes, substantially. That variance is the reason this product runs on a schedule rather than on demand. A single check of a generated answer is close to meaningless, and any tool reporting one as a status is reporting a coin flip.

Measure your own brand

The trial runs on the same rules described above. No card required.