Open robots.txt and search for GPTBot, OAI-SearchBot, ClaudeBot, Google-Extended and PerplexityBot, then run a live fetch too, since a CDN or bot wall can block a crawler that robots.txt says is allowed. 7 of 40 major sites checked (18%) blocked at least one major AI crawler outright (Reflexa AI Crawler Study, August 2026).
This is the first and most decisive layer in the chain the what is GEO pillar lays out: access, then identity, then content, then trust. A blocked crawler cancels every other GEO effort before it starts, so it is worth checking on purpose rather than assuming a CDN default got it right.
Eight user-agents cover the four engines that matter: OpenAI, Anthropic, Google and Perplexity. Each engine splits its crawler into a training role and a retrieval role, and the two behave differently when blocked.
| Bot | Sent by | Role |
|---|---|---|
| GPTBot | OpenAI | Trains future models |
| OAI-SearchBot / ChatGPT-User | OpenAI | Live web search & user-triggered fetch for ChatGPT |
| ClaudeBot | Anthropic | Trains future models |
| Claude-User / Claude-SearchBot | Anthropic | User-triggered fetch & search-quality crawling for Claude |
| Google-Extended | Trains Gemini & feeds AI Overviews | |
| PerplexityBot / Perplexity-User | Perplexity | Indexing & on-demand fetch for Perplexity answers |
Proof: OpenAI — GPTBot & OAI-SearchBot documentation; Perplexity — crawler documentation; full role-by-role detail in the Reflexa AI crawler guide.
Do this: check the retrieval bots first. They decide whether an engine can answer with you named today, not just whether a future model learns about you.
Open yoursite.com/robots.txt and look for a Disallow: / rule under that bot's own user-agent group, or under the catch-all User-agent: *. Engines apply the most specific matching group first; if a crawler has no group of its own, it falls back to the wildcard group, so a broad Disallow: / meant for something else can silently shut out every AI bot at once.
Proof: Search Engine Land — Anthropic clarifies how Claude bots crawl sites and how to block them.
Note: Anthropic's three Claude bots honor one shared rule, so there is no equivalent split for Claude the way there is for GPTBot and OAI-SearchBot; see should I block AI crawlers for the full engine-by-engine trade-off.
Because robots.txt is a request the crawler chooses to honor, not an access control, and a CDN or bot-management layer can override it silently. Cloudflare, Akamai and similar bot-wall products can challenge or block a request carrying an AI crawler's user-agent with a 403, a 503 or a "just a moment" page, while robots.txt still says the bot is allowed and a normal browser sees the site fine. That gap is exactly the case a file-only check misses.
Do this: pair the robots.txt read with a live request using each bot's real user-agent string, since that is the only way to see a CDN-level block.
Send your homepage one request per bot, using that bot's genuine user-agent, and compare the response to a normal browser request. Doing this by hand means finding eight exact user-agent strings and reading HTTP status codes for each one; the free AI Crawler Access Check does it in one pass, reporting a robots.txt verdict and a live-fetch verdict for every bot side by side, plus whether llms.txt and sitemap.xml exist. One honest caveat the tool states plainly: since its request does not come from a vendor's verified IP range, a CDN that allow-lists only confirmed bots will challenge it too and still let the real crawler through, so that result is reported as "check your bot rules" rather than a hard block.
Shortcut: run the check, then open your CDN or WAF's bot-management settings directly to confirm verified AI bots are on the allow list.
18% of major sites block at least one, and the block rate is close across engines rather than concentrated in one. Reflexa checked the robots.txt of 40 well-known companies in August 2026:
| AI crawler | Sent by | Sites blocking it |
|---|---|---|
| GPTBot | OpenAI | 7 of 40 · 18% |
| ClaudeBot | Anthropic | 7 of 40 · 18% |
| ChatGPT-User | OpenAI | 6 of 40 · 15% |
| Google-Extended | 6 of 40 · 15% | |
| OAI-SearchBot | OpenAI | 5 of 40 · 12% |
| PerplexityBot | Perplexity | 5 of 40 · 12% |
Blocking is concentrated by sector, not evenly spread: media and news sites block at 57% (4 of 7), against retail at 14% (1 of 7), SaaS at 7% (1 of 15), and finance and travel at 0%. The New York Times, BBC, Bloomberg and Reuters each block all six major AI crawlers outright, a deliberate stance to protect paid journalism.
Proof: Reflexa AI Crawler Access Study (robots.txt of 40 well-known companies checked August 22, 2026).
The nuance: for a publisher, that block can be a considered strategy. For a business that wants ChatGPT, Claude or Perplexity to recommend it, the same rule is usually a CDN default nobody chose on purpose.
None of these checks require buying anything, and each takes a few minutes:
Open yoursite.com/robots.txt and search for OAI-SearchBot, ChatGPT-User and GPTBot. If none of those carry a Disallow: / rule, and there's no broad Disallow: / under User-agent: *, robots.txt allows ChatGPT in; then run the free AI Crawler Access Check at /tools/ai-crawler-check to catch a CDN-level block robots.txt won't show.
No. A CDN, WAF or bot-management rule (Cloudflare, Akamai, Vercel) can challenge or block a crawler's real request while robots.txt says everything is allowed, since robots.txt is a polite request the crawler chooses to honor, not an access control.
The retrieval bots: OAI-SearchBot and ChatGPT-User for ChatGPT, Claude-User and Claude-SearchBot for Claude, and PerplexityBot for Perplexity. These fetch your page while answering a live question, so blocking one removes you from that engine's answers today, not just from future training.
7 of 40 major sites checked (18%) blocked at least one major AI crawler in robots.txt, concentrated in media (57%) versus SaaS (7%) and finance (0%) (Reflexa AI Crawler Study, August 2026). Some of that is a deliberate publisher stance; some is a CDN default nobody chose.
Yes. The AI Crawler Access Check at /tools/ai-crawler-check reads your robots.txt rule for eight AI user-agents and fetches your homepage with each bot's real user-agent, so a robots.txt pass and a CDN-level challenge both show up, in about 3 minutes with no signup.
The free check tests every major AI crawler against your site, robots and CDN-level, in 3 minutes, with the evidence.