Nine steps and a fixed protocol, run the same way every time. Published in full so you can check our work — or run it yourself and hold us to it.
It is the right question, and most AEO reporting cannot answer it. Language models are non-deterministic: ask the same question twice and you can get two different answers, with two different sources, an hour apart. A screenshot of ChatGPT naming your competitor proves almost nothing on its own.
So the measurement has to be built to survive that. Fixed questions, multiple runs, clean sessions, recorded engine versions, and frequency reported rather than a single result. Everything below is the protocol we run, at the values we actually run it at — not the values that would sound most impressive.
Steps 1 to 6 are the free audit. Steps 7 to 9 are the work you would pay for.
We write 20 questions a real buyer would actually type — pulled from your sales calls, your category terms, and the comparison queries people run before they shortlist. Not keywords. Questions. The set is written down and version-stamped before anything is measured.
We run the full set before changing a single thing on your site, and record the result. Without a baseline there is no result later, only a claim. The baseline is part of your report whether it flatters you or not.
Every question goes to each answer engine we can reach under the same conditions your buyer meets — logged out, no memory, no personalisation, market pinned. Each report names the engines it covers and the dates it covers them. We do not report an engine we could not reach.
Language models are non-deterministic: the same question can return a different answer an hour later. One run is an anecdote. We run each question three times per engine, in separate fresh sessions, and report how often you were cited — not whether you were cited once.
For every run we record three things: were you named, were you linked as a source, and where in the answer it happened. A brand named in passing and a brand cited as the source are different outcomes and we never merge them.
For the questions you lose, we record which competitor domains and which third-party sources — review sites, Reddit threads, directories, trade press — the models quote in your place. This is usually the most useful page of the report: it tells you exactly where the citation is going and what you would have to displace.
Three layers, in this order. Technical: crawler access, rendering, schema, whether the page can be read at all. Content: whether there is a quotable, self-contained answer on the page a model can lift. Entity: whether anything outside your own website corroborates that you exist and are credible.
Technical blockers first — nothing else matters while GPTBot is disallowed. Then the pages that answer the questions you are losing. Then the entity signals, which are the slowest to move and the hardest for a competitor to copy.
Same questions, same engines, same conditions, run at a fixed monthly interval so the comparison is like-for-like. The report leads with the change in Citation Share against the baseline, with the run counts visible, alongside AI referral traffic and the leads attributed to it.
Every one of these changes the number you get. A report that does not state them is not a measurement.
| Parameter | Our setting | Why it matters |
|---|---|---|
| Questions per set | 20, fixed for the engagement | Written before the baseline and version-stamped. Changing the set mid-engagement would make every comparison meaningless, so it changes only at a quarterly review — and the change is recorded in the report. |
| Runs per question, per engine | 3, in separate fresh sessions | Three runs per question per engine gives 60 data points per engine on a 20-question set. We report frequency — “cited in 7 of 20 questions, 14 of 60 runs” — never a single screenshot. |
| Question order | Randomised on every run | Asking the same questions in the same sequence lets earlier answers influence later ones. The order is shuffled per run so it cannot. |
| Session state | Logged out, memory off, no history | Personalisation is the fastest way to fool yourself into thinking you are cited. A logged-in account that has visited your site for months will show you to yourself. We test the way a stranger meets you. |
| Location | Pinned to one market per set | Answers differ by country. One tracked set is one market — US or UK. A second market is a second set with its own baseline, never averaged into the first. |
| Engines and versions | Recorded per run | Every row carries the engine, the model or interface version where the engine exposes it, and the date and time of the run. When an engine is unreachable for a reporting period, the report says so rather than quietly dropping it. |
| Citation vs mention | Counted separately, always | A citation is your domain given as an attributed source — linked, footnoted, or named as where the answer came from. A mention is your brand appearing in the prose with no attribution. Mentions are worth having; they are not citations and we never report them as such. |
| Repeat citations in one answer | Counted once per run | Being cited three times in one answer is one cited run. Position and prominence are recorded separately, in Citation Quality, so that a strong placement is not hidden and a weak one is not inflated. |
| Cadence | Baseline, then monthly | Re-tests run inside the same 72-hour window each month. Engines change behaviour over weeks; measuring on wildly different dates measures the engine, not your work. |
“AI visibility” is not a measurement. These are.
It is a sample, not a census. Twenty questions run three times is a representative sample of how the engines answer your category. It is not every question every buyer will ever ask, and Citation Share is an estimate with a margin, not a meter reading.
No engine publishes its citation logs. Nobody — no tool, no agency, no platform — has access to what a model actually returned to a specific buyer. Everyone in this field is sampling from the outside. Anyone telling you otherwise is selling you something.
AI referral traffic is undercounted. Several assistants strip the referrer, so a share of AI-driven visits land in your analytics as direct traffic. We report what is measurable and flag the gap rather than filling it with an estimate.
Correlation, not proof of cause. Citation Share moving after we change your pages is strong evidence the changes worked. It is not a controlled experiment, and we will not describe it as one.
If a competing agency shows you a cleaner number than this, ask them these ten questions: how many runs, fixed or ad-hoc prompts, randomised order, which engines, which model versions, logged in or out, which country, how do you define a citation, how do you count repeats, and how often do you re-run. The answers are the whole difference between a measurement and a screenshot.