Buyer's question

How do I check which AI bots can access my website?

Open robots.txt and search for GPTBot, OAI-SearchBot, ClaudeBot, Google-Extended and PerplexityBot, then run a live fetch too, since a CDN or bot wall can block a crawler that robots.txt says is allowed. 7 of 40 major sites checked (18%) blocked at least one major AI crawler outright (Reflexa AI Crawler Study, August 2026).

By Reflexa Technologies — the team building Reflexa, the AI Visibility Platform · September 28, 2026

In brief

This is the first and most decisive layer in the chain the what is GEO pillar lays out: access, then identity, then content, then trust. A blocked crawler cancels every other GEO effort before it starts, so it is worth checking on purpose rather than assuming a CDN default got it right.

Which AI bots do I actually need to check for?

Eight user-agents cover the four engines that matter: OpenAI, Anthropic, Google and Perplexity. Each engine splits its crawler into a training role and a retrieval role, and the two behave differently when blocked.

BotSent byRole
GPTBotOpenAITrains future models
OAI-SearchBot / ChatGPT-UserOpenAILive web search & user-triggered fetch for ChatGPT
ClaudeBotAnthropicTrains future models
Claude-User / Claude-SearchBotAnthropicUser-triggered fetch & search-quality crawling for Claude
Google-ExtendedGoogleTrains Gemini & feeds AI Overviews
PerplexityBot / Perplexity-UserPerplexityIndexing & on-demand fetch for Perplexity answers

Proof: OpenAI — GPTBot & OAI-SearchBot documentation; Perplexity — crawler documentation; full role-by-role detail in the Reflexa AI crawler guide.

Do this: check the retrieval bots first. They decide whether an engine can answer with you named today, not just whether a future model learns about you.

How do I read robots.txt to see if a bot is blocked?

Open yoursite.com/robots.txt and look for a Disallow: / rule under that bot's own user-agent group, or under the catch-all User-agent: *. Engines apply the most specific matching group first; if a crawler has no group of its own, it falls back to the wildcard group, so a broad Disallow: / meant for something else can silently shut out every AI bot at once.

Proof: Search Engine Land — Anthropic clarifies how Claude bots crawl sites and how to block them.

Note: Anthropic's three Claude bots honor one shared rule, so there is no equivalent split for Claude the way there is for GPTBot and OAI-SearchBot; see should I block AI crawlers for the full engine-by-engine trade-off.

Why isn't robots.txt enough on its own?

Because robots.txt is a request the crawler chooses to honor, not an access control, and a CDN or bot-management layer can override it silently. Cloudflare, Akamai and similar bot-wall products can challenge or block a request carrying an AI crawler's user-agent with a 403, a 503 or a "just a moment" page, while robots.txt still says the bot is allowed and a normal browser sees the site fine. That gap is exactly the case a file-only check misses.

Do this: pair the robots.txt read with a live request using each bot's real user-agent string, since that is the only way to see a CDN-level block.

How do I run a live fetch to test crawler access?

Send your homepage one request per bot, using that bot's genuine user-agent, and compare the response to a normal browser request. Doing this by hand means finding eight exact user-agent strings and reading HTTP status codes for each one; the free AI Crawler Access Check does it in one pass, reporting a robots.txt verdict and a live-fetch verdict for every bot side by side, plus whether llms.txt and sitemap.xml exist. One honest caveat the tool states plainly: since its request does not come from a vendor's verified IP range, a CDN that allow-lists only confirmed bots will challenge it too and still let the real crawler through, so that result is reported as "check your bot rules" rather than a hard block.

Shortcut: run the check, then open your CDN or WAF's bot-management settings directly to confirm verified AI bots are on the allow list.

How many sites actually block AI crawlers, and who?

18% of major sites block at least one, and the block rate is close across engines rather than concentrated in one. Reflexa checked the robots.txt of 40 well-known companies in August 2026:

AI crawlerSent bySites blocking it
GPTBotOpenAI7 of 40 · 18%
ClaudeBotAnthropic7 of 40 · 18%
ChatGPT-UserOpenAI6 of 40 · 15%
Google-ExtendedGoogle6 of 40 · 15%
OAI-SearchBotOpenAI5 of 40 · 12%
PerplexityBotPerplexity5 of 40 · 12%

Blocking is concentrated by sector, not evenly spread: media and news sites block at 57% (4 of 7), against retail at 14% (1 of 7), SaaS at 7% (1 of 15), and finance and travel at 0%. The New York Times, BBC, Bloomberg and Reuters each block all six major AI crawlers outright, a deliberate stance to protect paid journalism.

Proof: Reflexa AI Crawler Access Study (robots.txt of 40 well-known companies checked August 22, 2026).

The nuance: for a publisher, that block can be a considered strategy. For a business that wants ChatGPT, Claude or Perplexity to recommend it, the same rule is usually a CDN default nobody chose on purpose.

What to do this week

None of these checks require buying anything, and each takes a few minutes:

Keep reading

Frequently asked questions

What's the fastest way to check if ChatGPT can crawl my site?

Open yoursite.com/robots.txt and search for OAI-SearchBot, ChatGPT-User and GPTBot. If none of those carry a Disallow: / rule, and there's no broad Disallow: / under User-agent: *, robots.txt allows ChatGPT in; then run the free AI Crawler Access Check at /tools/ai-crawler-check to catch a CDN-level block robots.txt won't show.

Does a clean robots.txt guarantee AI bots can reach my site?

No. A CDN, WAF or bot-management rule (Cloudflare, Akamai, Vercel) can challenge or block a crawler's real request while robots.txt says everything is allowed, since robots.txt is a polite request the crawler chooses to honor, not an access control.

Which AI bots actually matter for getting recommended, not just trained on?

The retrieval bots: OAI-SearchBot and ChatGPT-User for ChatGPT, Claude-User and Claude-SearchBot for Claude, and PerplexityBot for Perplexity. These fetch your page while answering a live question, so blocking one removes you from that engine's answers today, not just from future training.

How common is it for a site to accidentally block AI crawlers?

7 of 40 major sites checked (18%) blocked at least one major AI crawler in robots.txt, concentrated in media (57%) versus SaaS (7%) and finance (0%) (Reflexa AI Crawler Study, August 2026). Some of that is a deliberate publisher stance; some is a CDN default nobody chose.

Is there a free tool that checks this for me?

Yes. The AI Crawler Access Check at /tools/ai-crawler-check reads your robots.txt rule for eight AI user-agents and fetches your homepage with each bot's real user-agent, so a robots.txt pass and a CDN-level challenge both show up, in about 3 minutes with no signup.

Stop guessing what's blocked. See who can reach you.

The free check tests every major AI crawler against your site, robots and CDN-level, in 3 minutes, with the evidence.

Run the free check →