There is no single bug. Usually it is one of five specific gaps: blocked crawler access, JavaScript-only text, entity ambiguity, no third-party evidence, or thin comparison content. Start with access: Reflexa's own study found 18% of major sites block at least one AI crawler outright in robots.txt, before content or authority get a chance to matter.
These five sit inside the same chain the what is GEO pillar walks through: access, then identity, then content, then trust. Here is where each link in that chain shows up specifically when the answer names a competitor instead of you.
Sometimes, yes, and it is the easiest of the five to fix once you can see it. OpenAI sends two separate bots: GPTBot, which crawls pages to help train future models, and OAI-SearchBot plus ChatGPT-User, which fetch pages live to answer a specific question inside ChatGPT. Blocking the wrong one, or blocking broadly with a single disallow rule, can remove you from the live answers even if your content would otherwise qualify.
| Category | Sites blocking at least one AI crawler |
|---|---|
| Media / news | 57% (4 of 7) |
| Retail | 14% (1 of 7) |
| SaaS | 7% (1 of 15) |
| Finance | 0% (0 of 5) |
| Travel | 0% (0 of 4) |
Proof: Reflexa AI Crawler Access Study (robots.txt of 40 well-known companies checked August 22, 2026); crawler roles from the Reflexa AI crawler guide.
Do this: run the free AI Crawler Access Check. It reads your robots.txt rules for every major bot and fetches your homepage with each bot's real user-agent, since a CDN or bot-wall block often will not show up in robots.txt at all.
Most AI crawlers do not execute JavaScript, so content that only appears after the page renders in a browser is invisible to them. A homepage, About page or pricing table can look complete to a person while its raw HTML response, the only thing most AI crawlers ever read, is close to empty. This is a common, entirely fixable gap, not a content-quality problem: the words are already written, they just are not present at the point AI actually fetches the page.
Rule of thumb: view your page's source, not the rendered page, and confirm your name, what you sell and your pricing are sitting in the raw HTML.
Do this: run the free What AI Actually Sees check. It shows your page's raw HTML next to the rendered version, side by side, so a missing block is obvious in seconds.
Naming a business is a claim about a specific real-world entity, so when an engine cannot confirm that "this domain," "this company name" and "this LinkedIn page" are the same thing, the safer move is to name a competitor whose identity resolves cleanly. A company that is "Acme Inc." on its homepage, "Acme" on LinkedIn and "Acme Software" in its own footer is handing the engine three candidates to reconcile instead of one confirmed match, exactly the drift covered in entity clarity.
Do this: run the free Schema & Entity check. It reads your homepage's Organization JSON-LD, your sameAs profiles and whether a Wikipedia or Wikidata anchor exists, the same signals engines use to settle who you are.
An engine weighs what your own site claims about you against what everyone else says, and with no third-party corroboration, a competitor with reviews, press or a Wikipedia page is the lower-risk pick. This is also part of why the same buyer question can get a different answer on a different day: independent research has found AI recommendations of brands and products highly inconsistent run to run.
Proof: SparkToro — New research: AIs are highly inconsistent when recommending brands or products.
Do this: run the free check at dashboard.reflexatechnologies.com/audit and read the sources it cites when it names a competitor instead of you. Those domains, review sites, comparison posts, press, are the third-party evidence the engine trusted over your own claims.
When a buyer asks "X vs Y," AI needs a page that actually compares the two, and if you do not have one, the engine cites whoever does, sometimes a competitor, sometimes a third-party review site neither of you controls. A features page that only lists your own product, with no pricing side by side and no honest trade-off, rarely gets pulled into an answer that is explicitly a comparison question. Structuring that page so the comparison is answered in the first sentence, per answer-first writing, is what makes it quotable.
Do this: search your own site for "[your product] vs [competitor]." If nothing comes up, run the free check at dashboard.reflexatechnologies.com/audit with that exact buyer question and see whose page the engine quoted instead.
None of these five checks require buying anything, and each takes a few minutes:
Bottom line: fix them in order. A perfectly written comparison page still loses if the crawler that would read it was blocked at the door.
Blocked or restricted crawler access is the most common front-door problem: in Reflexa's study of 40 well-known companies, 18% blocked at least one major AI crawler outright in robots.txt, which rules a company out before content or authority ever get weighed.
Yes. Blocks often come from a CDN or bot-protection rule rather than robots.txt, so a clean robots.txt file does not guarantee GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot can actually fetch your pages. The free AI Crawler Access Check tests both.
No. Structured data helps with entity ambiguity, one of five reasons, but a page engines cannot crawl or that has no third-party evidence will still lose even with perfect schema.
It varies by fix and by engine, from days for a crawler-access change to weeks for new third-party coverage to get indexed and trusted. Re-check weekly rather than expecting an overnight change.
No. Each reason has a free, no-signup check: AI Crawler Access Check, What AI Actually Sees, the Schema & Entity check, and the free audit at dashboard.reflexatechnologies.com/audit for third-party evidence and comparison questions.
The free check runs real buyer questions across ChatGPT, Claude and Perplexity — 3 minutes, evidence included.