Mechanics

How an AI assistant decides who to cite

Three gates stand between your page and a citation. Most sites fail one of them without ever finding out which.

A model must clear three separate gates before your name appears in an answer: it has to retrieve your page during the live query, extract a self-contained statement from it, and attribute that statement to you as an identifiable entity. Failing any one costs the citation, and they fail for different reasons and have different fixes.

The reason this matters more than it sounds: the three gates are invisible from the outside. Nothing in your analytics distinguishes “the model never fetched us” from “it fetched us and found nothing quotable”. Both look like silence.

Gate one: retrieval

When an assistant answers a live question it sends a named agent to fetch candidate pages. Those agents — OAI-SearchBot, ChatGPT-User, PerplexityBot and their equivalents — are permissioned separately from Googlebot in your robots.txt. Allowing Google says nothing about allowing them.

This gate also closes for duller reasons: a page that renders its content in client-side JavaScript, a slow response, a soft 404, a login wall, a CDN challenge served to anything that does not look like a browser. A model does not retry. It moves to the next candidate, which is your competitor.

The trap here is that blocking is not binary. Training crawlers and retrieval agents are different classes with very different costs, and most audits report them as one alarm. The distinction, and what we found across 137 UK B2B sites →

Gate two: extraction

Having fetched the page, the model needs a statement it can lift — short, self-contained, and true on its own without the paragraph around it. This is where well-optimised marketing sites fail hardest, because the things that make a page persuasive to a human make it useless to an extractor.

Prose that builds to a point gives a model nothing to take from the middle. A benefit-led headline — “Ship faster, worry less” — answers no question anybody typed. A pricing page that says “contact us” cannot be quoted in an answer about pricing. A comparison written as a narrative loses to a competitor's table.

What survives extraction is dull to write and effective to publish: the literal question as a heading, the answer in one or two sentences directly beneath it, then the nuance for the human who kept reading. Definitions, tables, specifications, numbers with their units, and explicit statements of scope all extract cleanly.

Gate three: attribution

The third gate is the one nobody expects, and the data on it is stark. In Semrush's AI visibility study, only 6–27% of the most-mentioned brands in a category also ranked as top cited sources, depending on the industry. Being talked about and being credited are close to independent variables.

A model will happily absorb a fact from your page and state it without naming you. To attach your name it needs to resolve you as a distinct entity — consistent naming, structured data with stable identifiers, corroboration elsewhere that you are who the page says you are. Where that resolution fails, the safe move for the model is to state the fact unattributed, or to credit a source it is more confident about.

Why third-party sources dominate

This is also why the most-cited domains in AI answers are almost never vendor sites. Peec AI's analysis of 30 million sources, published March 2026, found Reddit the most-cited domain across ChatGPT, Google AI Mode, Gemini, Perplexity and AI Overviews, with YouTube and LinkedIn next, and Wikipedia and Forbes also in the top five. Platform preferences differed — ChatGPT leaned to Wikipedia, Reddit and editorial sites; Perplexity to Reddit, LinkedIn and G2 for B2B queries — but the pattern held.

Those domains clear all three gates by construction. They are open to retrieval, structurally quotable, and unambiguously attributable. A vendor page claiming its own product is best clears gate three but is discounted at gate two on grounds of self-interest; the same claim on G2 or in a Reddit thread carries weight the vendor site cannot generate about itself.

The practical reading is not “go spam Reddit”. It is that citation-worthiness is partly outside your domain, and a plan that only touches your own site has a ceiling built into it.

Diagnosing which gate you are failing

SymptomLikely gateWhere to look
Never appears, on any phrasing, in any engineRetrievalrobots.txt by named agent, render mode, status codes, CDN rules
Competitors cited on questions you answer betterExtractionWhether your pages contain liftable answers under question-shaped headings
Your facts appear in answers without your nameAttributionEntity schema, naming consistency, third-party corroboration
Cited on brand questions, absent on category onesAttribution & corroborationPresence in the sources models trust for your category

The last row is the most common pattern in B2B and the most misread. A company that appears when you ask about it by name concludes it is visible. Ask the category question instead — the one a buyer types before they know you exist — and the picture usually changes.

Sources

Every figure above is linked to the page that published it. Where a number is self-reported by the company that benefited from it, this post says so.

FAQ

Related questions

No. Semrush's AI visibility study found that only 6–27% of the most-mentioned brands in a category also ranked as top cited sources, depending on industry. Mention and citation are close to independent: a model can state a fact from your page without naming you.

Peec AI's March 2026 analysis of 30 million sources found Reddit the most-cited domain across ChatGPT, Google AI Mode, Gemini, Perplexity and AI Overviews, followed by YouTube and LinkedIn, with Wikipedia and Forbes also in the top five.

Usually extraction. Answering well for a human and being quotable by a model are different properties. A page that builds to its point across several paragraphs gives an extractor nothing to lift; a competitor with the question as a heading and the answer beneath it wins the citation with worse content.

Only partly. Retrieval and extraction are yours to fix. Attribution depends on corroboration in sources the model already trusts for your category, which sits outside your domain — so a plan confined to your own site has a ceiling.

See what the engines say about you.

Send your domain. The check comes back free, with the findings named and the fixes attached.