Methodology

How we measure AI visibility.

Nine steps and a fixed protocol, run the same way every time. Published in full so you can check our work — or run it yourself and hold us to it.

The problem with most AI visibility numbers

“How do you know 30% isn’t just randomness?”

It is the right question, and most AEO reporting cannot answer it. Language models are non-deterministic: ask the same question twice and you can get two different answers, with two different sources, an hour apart. A screenshot of ChatGPT naming your competitor proves almost nothing on its own.

So the measurement has to be built to survive that. Fixed questions, multiple runs, clean sessions, recorded engine versions, and frequency reported rather than a single result. Everything below is the protocol we run, at the values we actually run it at — not the values that would sound most impressive.

The method

Nine steps, every engagement.

Steps 1 to 6 are the free audit. Steps 7 to 9 are the work you would pay for.

1

Define the question set

We write 20 questions a real buyer would actually type — pulled from your sales calls, your category terms, and the comparison queries people run before they shortlist. Not keywords. Questions. The set is written down and version-stamped before anything is measured.

2

Establish the baseline

We run the full set before changing a single thing on your site, and record the result. Without a baseline there is no result later, only a claim. The baseline is part of your report whether it flatters you or not.

3

Test across engines

Every question goes to each answer engine we can reach under the same conditions your buyer meets — logged out, no memory, no personalisation, market pinned. Each report names the engines it covers and the dates it covers them. We do not report an engine we could not reach.

4

Run each question three times

Language models are non-deterministic: the same question can return a different answer an hour later. One run is an anecdote. We run each question three times per engine, in separate fresh sessions, and report how often you were cited — not whether you were cited once.

5

Record the citations

For every run we record three things: were you named, were you linked as a source, and where in the answer it happened. A brand named in passing and a brand cited as the source are different outcomes and we never merge them.

6

Identify who is cited instead

For the questions you lose, we record which competitor domains and which third-party sources — review sites, Reddit threads, directories, trade press — the models quote in your place. This is usually the most useful page of the report: it tells you exactly where the citation is going and what you would have to displace.

7

Diagnose the gaps

Three layers, in this order. Technical: crawler access, rendering, schema, whether the page can be read at all. Content: whether there is a quotable, self-contained answer on the page a model can lift. Entity: whether anything outside your own website corroborates that you exist and are credible.

8

Implement, in dependency order

Technical blockers first — nothing else matters while GPTBot is disallowed. Then the pages that answer the questions you are losing. Then the entity signals, which are the slowest to move and the hardest for a competitor to copy.

9

Re-test and report the delta

Same questions, same engines, same conditions, run at a fixed monthly interval so the comparison is like-for-like. The report leads with the change in Citation Share against the baseline, with the run counts visible, alongside AI referral traffic and the leads attributed to it.

The measurement protocol

The parameters, stated.

Every one of these changes the number you get. A report that does not state them is not a measurement.

Parameter Our setting Why it matters
Questions per set 20, fixed for the engagement Written before the baseline and version-stamped. Changing the set mid-engagement would make every comparison meaningless, so it changes only at a quarterly review — and the change is recorded in the report.
Runs per question, per engine 3, in separate fresh sessions Three runs per question per engine gives 60 data points per engine on a 20-question set. We report frequency — “cited in 7 of 20 questions, 14 of 60 runs” — never a single screenshot.
Question order Randomised on every run Asking the same questions in the same sequence lets earlier answers influence later ones. The order is shuffled per run so it cannot.
Session state Logged out, memory off, no history Personalisation is the fastest way to fool yourself into thinking you are cited. A logged-in account that has visited your site for months will show you to yourself. We test the way a stranger meets you.
Location Pinned to one market per set Answers differ by country. One tracked set is one market — US or UK. A second market is a second set with its own baseline, never averaged into the first.
Engines and versions Recorded per run Every row carries the engine, the model or interface version where the engine exposes it, and the date and time of the run. When an engine is unreachable for a reporting period, the report says so rather than quietly dropping it.
Citation vs mention Counted separately, always A citation is your domain given as an attributed source — linked, footnoted, or named as where the answer came from. A mention is your brand appearing in the prose with no attribution. Mentions are worth having; they are not citations and we never report them as such.
Repeat citations in one answer Counted once per run Being cited three times in one answer is one cited run. Position and prominence are recorded separately, in Citation Quality, so that a strong placement is not hidden and a weak one is not inflated.
Cadence Baseline, then monthly Re-tests run inside the same 72-hour window each month. Engines change behaviour over weeks; measuring on wildly different dates measures the engine, not your work.
The metrics

Five numbers, defined.

“AI visibility” is not a measurement. These are.

Citation Share The headline number
The share of your tracked questions in which your domain is given as an attributed source, across all runs. Cited runs ÷ total runs. One number, always reported with its run count, always against the baseline.
Mention Share Brand presence without attribution
The share of runs where your brand is named in the answer but not cited as a source. Rising Mention Share with flat Citation Share is a specific, fixable diagnosis: the models know you exist but have nothing of yours worth quoting.
Competitor Share Who holds the answer
The same measurement run for each named competitor on the same question set. It converts “we are not visible” into “these three companies hold 60% of our category questions, here is which one to displace first.”
Citation Quality Where in the answer you land
Not all citations are equal. We score whether you appear inside the direct answer or in a trailing source list, your position among the sources, and whether the model recommends you or merely lists you.
Source Authority The third parties that outrank your own site
The ranked list of external domains cited on your questions — review platforms, communities, directories, trade press. It is the shortest route to citation for most companies, because the model already trusts those sources and you can be present on them in weeks rather than quarters.
Limits

What this method cannot tell you.

It is a sample, not a census. Twenty questions run three times is a representative sample of how the engines answer your category. It is not every question every buyer will ever ask, and Citation Share is an estimate with a margin, not a meter reading.

No engine publishes its citation logs. Nobody — no tool, no agency, no platform — has access to what a model actually returned to a specific buyer. Everyone in this field is sampling from the outside. Anyone telling you otherwise is selling you something.

AI referral traffic is undercounted. Several assistants strip the referrer, so a share of AI-driven visits land in your analytics as direct traffic. We report what is measurable and flag the gap rather than filling it with an estimate.

Correlation, not proof of cause. Citation Share moving after we change your pages is strong evidence the changes worked. It is not a controlled experiment, and we will not describe it as one.

If a competing agency shows you a cleaner number than this, ask them these ten questions: how many runs, fixed or ad-hoc prompts, randomised order, which engines, which model versions, logged in or out, which country, how do you define a citation, how do you count repeats, and how often do you re-run. The answers are the whole difference between a measurement and a screenshot.

Find out if AI recommends you.

Send us your domain. You get the full audit back — free, no call required, yours to keep either way.