AI crawlers can read your site only if four gates let them through. Your robots.txt and your CDN have to allow them, the page has to carry its content in the HTML they receive, they have to be able to reach the page, and they have to understand what they find. Most of these checks take minutes, and a closed gate sends you no error message.
- Amazon blocks seven AI crawler names in its robots.txt on purpose. The same lines added to a small store's file by an app are a mistake to find and fix.
- Google renders JavaScript. OpenAI, Anthropic and Perplexity do not document whether their crawlers do, so put the facts that matter in the initial HTML.
- RFC 9309 says robots.txt rules are not a form of access authorization, so a crawler that ignores them still reads your pages.
- Each vendor documents different rules for its crawlers, so check what blocking each name does before you keep or remove it.
Most site owners never check whether an AI crawler can get into their site at all. In our June 2026 study of 216 Shopify storefronts, two thirds scored below 50 out of 100 for AI legibility (the study). The figures below show how often the front door itself was the problem.
The four gates a crawler has to pass
A crawler meets these gates in the order a visit happens. A closed first gate ends the visit, and the later gates decide how much of your site the crawler reaches and understands.
| Gate | The question | What fails it | How to check |
|---|---|---|---|
| 1. Access | Are you letting the crawler in? | A Disallow line in robots.txt, often added by an app or a theme, or a CDN rule that blocks the bot | Open yourdomain.com/robots.txt, then your CDN's bot settings |
| 2. Reach | Can it find every page? | Orphan pages, dead URLs, an outdated sitemap, broken canonicals | Open yourdomain.com/sitemap.xml and check a few URLs |
| 3. Render | Can it read the page once it arrives? | Content that appears only after JavaScript runs, for a crawler that may not run it | View the page source and search it for a sentence you can see on screen |
| 4. Parse | Can it understand what it found? | Facts that live only in images, no structured data, ambiguous copy | Check images and structured data on your key pages |
Each gate has a place to check it today.
| What you are checking | Gate | Where to check it now |
|---|---|---|
| robots.txt rules, crawler by crawler | 1 | robots.txt in the AI era, and the crawler table below |
| A CDN or firewall that blocks what robots.txt allows | 1 | Is Cloudflare blocking AI from your site?, and your own CDN's bot settings |
| Page directives such as noindex and nosnippet | 1 and 4 | Google's AI features and your website for Google, and the Copilot answer below for Bing |
| Initial HTML against the rendered page | 3 | JavaScript rendering and AI crawlers |
Gate one in practice: robots.txt is a policy, not a lock
Your robots.txt file tells crawlers what you want. It does not stop anyone. RFC 9309, the standard behind the file, says its rules are not a form of access authorization (RFC 9309, September 2022). I hold CISA and CISSP certifications, and in an audit I would file robots.txt under policy, not under controls. A crawler that follows the standard obeys it, and anything else reads your pages anyway.
Amazon shows what a deliberate policy looks like. On September 26, 2026, amazon.com/robots.txt disallowed GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, PerplexityBot, CCBot and Google-Extended from the whole site. That is a business decision, and I respect it. The same lines copied into a small store's file by an app or a theme are a finding. Check what each crawler does before you keep or remove its line.
| Name in robots.txt | Vendor | What it does | What blocking it means |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Surfaces websites in ChatGPT search | Your pages are not shown in ChatGPT search answers, though they can still appear as navigational links |
| GPTBot | OpenAI | Collects content for training OpenAI's foundation models | A signal that your content should not be used for training |
| ChatGPT-User | OpenAI | Visits a page when a user asks ChatGPT to | OpenAI says robots.txt rules may not apply to these user actions |
| ClaudeBot | Anthropic | Collects content that could contribute to model training | A signal that your future materials should be left out of training |
| Claude-SearchBot | Anthropic | Indexes content to improve search results | Anthropic says it may reduce your visibility in user search results |
| Claude-User | Anthropic | Retrieves pages when a Claude user asks | Anthropic says it stops retrieval of your content in response to a user query |
| PerplexityBot | Perplexity | Surfaces and links websites in Perplexity search, not used for foundation models | Your pages leave Perplexity's search results |
| Perplexity-User | Perplexity | Fetches pages a user asked for | Perplexity says this fetcher generally ignores robots.txt rules |
| Google-Extended | A control token with no crawler of its own, used for AI training and grounding in some Google systems | Google says it does not affect inclusion or ranking in Google Search | |
| Applebot-Extended | Apple | An opt out from training Apple's generative foundation models | Apple says the pages can still be included in search results |
| CCBot | Common Crawl | Builds Common Crawl's open repository of web crawl data | Common Crawl stops adding your pages to that repository |
OpenAI says a robots.txt change takes about 24 hours to reach its search systems, and Perplexity says up to 24 hours. Check again the next day before you conclude anything.
Gates two to four: reach, read and understand
A crawler reaches pages by following links and your sitemap. A sitemap can also go stale inside Google with no error to warn you. Google keeps reading an old copy of the index, and it misses every page published after that copy. A fresh submission in Search Console fixes it. Redirect chains and contradictory canonicals make it harder to tell which URL is the real one.
Google renders JavaScript with an evergreen version of Chromium (JavaScript SEO basics). OpenAI, Anthropic and Perplexity say nothing in their crawler documentation about whether their crawlers do. So I treat anything that appears only after a script runs as content those crawlers may never see. If your price, product grid or description loads by script, check the initial HTML, and read JavaScript rendering and AI crawlers.
Gate four is about understanding. A fact that lives only in an image with no alt text is not in the text a crawler reads. You may also come across llms.txt. Google says Google Search does not use it and that keeping one for other services is fine (AI optimization guide).
A scan is a snapshot, and legibility drifts
A site that passes all four gates today can fail one next month. A theme update can overwrite a snippet, and an app update can add a Disallow line nobody typed. The machine readable layer is the one people never look at, so it regresses without anyone noticing while the site still looks fine. Scan again after every theme or app change. We run the same scan on our own site every month, and you can see our result on Our Own Score (why we scan ourselves).
What our scan checks, and what it does not
Since September 26, 2026, our scan has checked eight crawler names one by one against the robots.txt rule for the page it reads: GPTBot, OAI-SearchBot and ChatGPT-User from OpenAI, ClaudeBot, Claude-SearchBot and Claude-User from Anthropic, and PerplexityBot and Perplexity-User from Perplexity. It reports Google-Extended on its own as a control for Gemini training and grounding and never counts it as a block, since Google documents that it has no crawler of its own. The scan does not test Bingbot or Applebot today. A firewall can also answer our request differently from the way it answers the real crawler, so check your CDN logs as well.
Questions people actually ask
How do I know if AI crawlers can access my store?
Start with your robots.txt file at yourdomain.com/robots.txt and look for rules that name AI crawlers. Then check that your key content appears in the page source, not only after JavaScript runs. Check your CDN or firewall settings for bot rules too. A free scan checks whether eight crawlers from OpenAI, Anthropic and Perplexity are allowed in and reads each page as it is served. It does not test Bingbot or Applebot, so check those by hand.
Do AI crawlers read JavaScript?
Google does. Its documentation says Google Search runs JavaScript with an evergreen version of Chromium. OpenAI, Anthropic and Perplexity do not document whether their crawlers run JavaScript. So put the facts that matter, such as prices, product names and descriptions, in the initial HTML the server sends. Compare the page source with the rendered page to find what exists only after a script runs, and treat that content as invisible to crawlers that have not said otherwise.
Is being blocked the same as ranking poorly?
It is not, and it is worse. Ranking poorly means your page was considered and placed low. Being blocked or unreadable means the crawler never reached or parsed the page, so it was not considered at all. That makes access the first thing to check. It takes minutes, and every other improvement depends on it.
How do I stop Microsoft Copilot from using my content without leaving Bing search?
Microsoft documented two robots meta values for this on September 22, 2023, in a post that refers to Bing Chat rather than Copilot. NOARCHIVE keeps a page out of the chat answers and out of training for Microsoft's generative AI foundation models. NOCACHE lets answers use only the URL, title and snippet. A page with both is treated as NOCACHE, and Microsoft says either way the page stays in Bing search results. Its current help page did not load for us on September 26, 2026, so check it yourself before relying on a 2023 post.
Can a bot fake the GPTBot or Googlebot user agent?
Yes, a user agent string is text any client can send, and Google and Common Crawl both warn about spoofed crawler names. So verify the source instead of trusting the name. Google and Apple support a reverse DNS lookup, Microsoft offers a Verify Bingbot tool, and OpenAI, Anthropic, Perplexity, Google, Apple and Common Crawl publish their IP ranges as files. Treat a request that uses a crawler's name from an IP outside those ranges as an unverified bot.
Can I block AI crawlers by IP address instead of robots.txt?
You can, but the vendors document robots.txt as the opt out. Anthropic warns that blocking its IP addresses may not work correctly or keep an opt out in place, since its bots then cannot read your robots.txt. IP rules make sense as enforcement against bots that ignore robots.txt, and they must use the ranges each vendor publishes, which change over time. For crawlers you keep out for business reasons, use a robots.txt rule plus a CDN rule for unverified traffic.
Will blocking AI crawlers stop my content from being used to train AI models?
It works only for crawlers that honor the rule, and only for content collected after the rule is in place. OpenAI, Anthropic, Google, Apple and Common Crawl each document a robots.txt name for this: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot. None of those pages says the rule removes content collected earlier, and copies of your text on other websites follow those websites' rules. RFC 9309 says robots.txt is not access authorization, so private material needs a login.
Can crawlers read content behind a login or a consent wall?
They cannot read content that requires signing in. Google's JavaScript guidance says to return a 401 status for pages behind a login, and Anthropic says its bots will not try to bypass CAPTCHAs. A consent banner that only overlays the page leaves the HTML readable. Google advises against redirecting visitors to a separate page to collect consent, and a wall that withholds content until someone clicks withholds it from crawlers too. Keep anything you want found on public pages.
Sources
- OpenAI, Overview of OpenAI Crawlers. No update date is shown; read September 26, 2026.
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?, Claude Help Center; read September 26, 2026.
- Perplexity, Perplexity Crawlers. No update date is shown; read September 26, 2026.
- Google, Google common crawlers (last updated July 14, 2026), Verify requests from Google crawlers and fetchers (last updated March 20, 2026) and Google user triggered fetchers (last updated August 19, 2026); read September 26, 2026.
- Google, AI features and your website (last updated December 10, 2025), Understand the JavaScript SEO basics (last updated March 4, 2026), Avoid intrusive interstitials and dialogs (last updated December 10, 2025) and the AI optimization guide (last updated July 10, 2026); read September 26, 2026.
- Apple, About Applebot. No update date is shown; read September 26, 2026.
- Common Crawl, CCBot. No update date is shown; read September 26, 2026.
- IETF, RFC 9309, Robots Exclusion Protocol, September 2022; read September 26, 2026.
- Microsoft, Announcing new options for webmasters to control usage of their content in Bing Chat, Bing Webmaster Blog, September 22, 2023, and the Verify Bingbot tool; read September 26, 2026.
- Amazon, robots.txt, read with curl on September 26, 2026 at 13:58 UTC.
- Visibility Mesh, study of 216 Shopify storefronts scanned June 28, 2026, and study of 303 non ecommerce websites, July 2026.
See what a machine sees
You cannot tell from your browser whether AI crawlers can read your site. A free scan shows you in a few minutes. It checks whether the crawlers from OpenAI, Anthropic and Perplexity can fetch your pages, and it reads five of them the way a crawler receives them.
Your buyers are already asking AI. This is how you make your website readable to the assistants they ask.
Everything in this article is measurable on a live storefront. The Full Assessment and Roadmap reads your website the way AI crawlers receive it and hands you every fix in plain English, in the order we would make them.