Perplexity is an answer engine that shows its work — every answer carries numbered, inline citations linking straight to the pages it drew from. That makes citation the one AI-visibility outcome you can literally count: open the Sources panel and see whether you made the list. Getting on it comes down to three controllable things — letting PerplexityBot reach your site, publishing a page that answers the specific buying question in plain HTML, and giving Perplexity a claim worth citing over a competitor's. Here's what actually decides it, including the parts of Perplexity's crawling that aren't fully in your control.
Perplexity attaches a numbered citation to nearly every claim in its answers, with a Sources panel listing exactly which pages it used. That's a different shape of visibility than getting recommended by ChatGPT, where the reasoning behind a shortlist stays mostly hidden. With Perplexity you can ask your own buyer questions and read the citation list like a scoreboard — your site is either on it or it isn't, for that exact question.
Do this: ask Perplexity three real buying questions about your category and check the Sources panel on each answer — that's your current citation rate, not a guess.
Perplexity runs two distinct crawlers, and only one of them is bound by robots.txt. PerplexityBot is the indexing crawler that builds Perplexity's own knowledge base ahead of time, and Perplexity's documentation states it respects robots.txt. Perplexity-User is a different, real-time agent that fetches a page at the moment someone pastes a URL or asks about it directly in a chat — Perplexity has described this one as acting on a user's explicit request rather than as a bot, so it isn't held to the same robots.txt rule.
Proof: Perplexity — crawler documentation (PerplexityBot & Perplexity-User).
Check it now: confirm PerplexityBot isn't disallowed in robots.txt and isn't caught by a CDN's default AI-bot block — the same check that matters for GPTBot and ClaudeBot.
Cloudflare reported in August 2025 that Perplexity was using additional, undeclared crawlers that impersonated a regular Chrome browser to keep reading sites after their declared bots were blocked. Cloudflare said it verified the behavior with test domains that disallowed all automated access, then delisted Perplexity as a Cloudflare Verified Bot and added new rules to block the undeclared traffic. Perplexity publicly disputed Cloudflare's characterization of the findings.
Proof: Search Engine Journal — Cloudflare delists and blocks Perplexity from crawling websites.
Bottom line: allowing PerplexityBot is still the right, controllable move — it's the difference between being indexed on purpose and being read anyway. It just isn't a guarantee that it's the only way you're being read; server-log monitoring closes that gap.
Perplexity pulls the passage that answers a question most directly, not the page that's broadest or best-known. A dedicated page that states the answer to "best X for [use case]" or "X vs Y" in its first paragraph is far more quotable than the same claim buried three paragraphs into a general product page. This is the same principle covered in full on the answer-first writing guide.
Check it now: pick your top buying question and read your own page's first 100 words as if you were the model — does it answer the question, or set up the answer?
Because Perplexity leans on live web search rather than a fixed knowledge cutoff, a narrowly-scoped page updated recently can out-cite a bigger, older, more "authoritative" one. A page that names numbers, dates and specifics reads as fresher and more citable than one that speaks in generalities, even when the general page sits on a larger domain.
Rule of thumb: when you update a claim or a number on a key page, update the visible date on it too — Perplexity has no way to know a page changed if nothing on the page says so.
Perplexity's citations often include third-party coverage, reviews and comparison sites alongside a company's own pages. If independent sources never mention a business by name in the context of the buying question, Perplexity has less to cross-check a self-authored claim against — and a competitor's third-party mentions win the citation instead.
Do this: track which domains Perplexity already cites in your category and prioritize getting a real mention on the ones that come up again and again.
The fastest read on Perplexity citations is to ask it the exact question your buyers ask, then look at who's in the Sources panel and who isn't. Reflexa's Sources tool automates this across your real buying questions and shows which domains get cited, how often, and whether your own site is among them — the same view covered in the tools overview.
The free check reads your crawler access and shows real citation data — 3 minutes, evidence included.