Access first, structure second, freshness third — and the pages it will not cite whatever you do.
Perplexity cites pages it can fetch, parse and attribute. Allow PerplexityBot and Perplexity-User in robots.txt and at the CDN, put a self-contained answer directly under a question-shaped heading, and keep the page dated and current.
Perplexity is the most legible of the assistants to optimise for, because it shows its sources inline. You can see exactly who it chose and read what made them choosable.
Perplexity operates more than one named agent, and they do different jobs:
PerplexityBot — indexes pages so they can be surfaced as sources. Block this and you are not in the pool at all.Perplexity-User — fetches a page in response to a specific user action. Blocking it means a user who asks about your page gets nothing back.Allow both in robots.txt, then verify what each actually receives — managed bot rules at the CDN override the file and are the most common cause of a silent block. How to test that properly →
Perplexity composes an answer from short passages and attributes each one. That rewards a specific page shape:
| Do | Instead of |
|---|---|
| A heading that is the literal question | “Our approach to X” |
| A complete answer in the first forty words under it | Three paragraphs of build-up |
| A table for anything comparative | Prose describing a comparison |
| A figure with its source and date inline | “Studies show…” |
| One question per section | Several questions answered in one block |
The test to apply to your own page: take any single sentence out of it and read it alone. If it still makes a complete, checkable claim, it is quotable. If it needs the paragraph around it, it is not.
Perplexity leans toward current material, particularly for questions where the answer changes. Put a visible date on the page, keep dateModified accurate in your schema, and update the substance rather than the timestamp. A page dated this month that says the same thing it said last year is worse than an honest old date.
Check your work the obvious way. Ask Perplexity your buyer’s category question and read the sources it names. Those are your real competitors for the citation, and the page it chose over yours will show you why. Run it several times — the source list varies between runs, which is exactly why a single check is not a measurement.
Freeze a set of buyer questions, run each one more than once, and record how often you appear rather than whether you appeared. Citation is probabilistic: appearing in one run out of three is a real result and a different result from three out of three. Our protocol →
Perplexity's indexing crawler. It fetches pages so they can be surfaced as sources in answers. It is separate from Perplexity-User, which fetches a page in response to a specific user action, and both are permissioned independently of Googlebot.
Only if you do not want to be cited by Perplexity. Unlike a training crawler, blocking it has a direct and immediate cost: you are removed from the pool of pages it can name as a source.
Usually one of three reasons: it cannot fetch your page, your page has no self-contained statement worth lifting, or it cannot confidently identify you as an entity. Read the sources it named for your category question — the page it chose will usually show which of the three applies.
Not by itself. Perplexity favours current material for questions whose answers change, but frequency without substance does not help. Updating the timestamp on an unchanged page is worse than an honest old date.
Freeze a set of buyer questions, run each several times, and record how often you appear rather than whether you appeared at all. The source list varies between runs, so a single check tells you very little.