Method

Measuring AI visibility without fooling yourself

The failure mode is not dishonesty. It is running the question once, seeing your name, and believing it.

Language models are non-deterministic: the same question asked three times produces three different answers, with different companies named. A single run tells you what happened once. Any AI visibility figure built on one run per question is a screenshot presented as a metric, and it falls apart the first time a client re-runs it themselves.

This is the part of AEO most likely to embarrass whoever sells it. Rank tracking taught a generation of marketers that a number checked today means the same thing tomorrow. Citation does not work that way, and pretending otherwise has a short half-life.

What varies between runs

Holding the question identical, the answer still moves with sampling randomness in the model itself, which pages the retrieval step happens to fetch that second, the model or interface version, the account state, and the market the query resolves to. Most of these are invisible in the output.

Account state is the one that catches people. Asking from your own logged-in account, with memory on and months of history about your company, is not a measurement of anything except how well the assistant remembers you. It is the single most common way a founder concludes they are visible when they are not.

The protocol

These are the parameters we publish and hold ourselves to. They are not exotic; they are the minimum for a number that survives being checked.

ParameterSettingWhy
Question set20, fixed and version-stamped before measuringA set edited between runs can be made to show anything
Runs3 per question, per engine, in separate fresh sessionsOne run cannot be told apart from randomness
OrderRandomised per runRemoves position effects within a session
Session stateLogged out, memory off, no historyOtherwise you measure your own footprint
MarketPinned to one country per tracked setAnswers differ by market; mixing them averages two different realities
Recorded per rowEngine, model or interface version, timestampMakes a disputed figure reproducible
CadenceMonthly, inside the same 72-hour windowRemoves drift from comparing different points in a release cycle

Report frequency, not existence

The output of this is not “you are visible in ChatGPT”. It is a frequency: cited in 7 of 20 questions, 14 of 60 runs. That phrasing does two things. It states the sample, so anyone can judge how much weight it carries. And it makes improvement legible — 14 of 60 becoming 23 of 60 is a real movement, where “yes” becoming “yes” is not.

Count citations and mentions separately

A citation is an attributed source — linked, footnoted, or named as where the information came from. A mention is your brand appearing in the answer text with no attribution. Merging them inflates the headline number, and it destroys the most useful diagnosis available.

Rising mention share with flat citation share is specific and fixable: the models know you exist but have nothing of yours worth quoting. That is an extraction and entity problem, not an awareness problem, and it points at different work than the same numbers moving together.

The five things to refuse

  • A screenshot as evidence. One answer, once, proves nothing in either direction.
  • A prompt set that changes between measurements. If the questions moved, the trend is an artefact.
  • “We track all the major engines” without a per-report statement of which were actually reached. Coverage varies with tooling and access; a blanket claim is usually false.
  • A single blended visibility score with no definition. If you cannot reconstruct it from the underlying counts, it is decoration.
  • Any guarantee of AI rankings. Non-deterministic systems cannot be guaranteed. Anyone offering it is either confused or counting on you not checking.

What this protocol still cannot do. Twenty questions run three times is a representative sample, not a census — citation share is an estimate with a margin, not a meter reading. It cannot attribute revenue to a citation. And it measures the engines it reached, on the dates it ran. We say all of this in the reports too. The full protocol, including its limits →

Sources

Every figure above is linked to the page that published it. Where a number is self-reported by the company that benefited from it, this post says so.

FAQ

Related questions

Language models are non-deterministic. Sampling randomness, which pages the retrieval step fetches at that moment, the model or interface version, account state and market all vary between runs. This is why a single run cannot support a visibility claim.

At minimum three runs per question per engine, in separate fresh sessions with randomised question order. tapFunnel runs twenty version-stamped questions three times each and reports frequency — for example 'cited in 7 of 20 questions, 14 of 60 runs' — rather than a yes or no.

Because a logged-in session with memory on and history about your own company measures your footprint with that account, not the model's default behaviour toward your category. It is the most common reason a founder concludes they are visible when they are not.

A citation is an attributed source — linked, footnoted or named as where the information came from. A mention is the brand appearing in the answer text with no attribution. Counting them separately gives a specific diagnosis: rising mentions with flat citations means models know you exist but have nothing of yours worth quoting.

See what the engines say about you.

Send your domain. The check comes back free, with the findings named and the fixes attached.