Open /robots.txt on every host that serves your pages and find the group of rules each crawler follows. Apply the longest matching rule in that group to each URL that matters. A crawler with its own User-agent group ignores the rules written for everyone else. The file does not enforce anything. Check the firewall in front of the site as well.
One line in this file can stop every compliant crawler from reading the whole site. An auditor would call that a single point of failure. The line is Disallow: / under User-agent: *, and it takes eleven characters.
Read robots.txt the way you would read a firewall policy. For each request, find the rule that applies and write the answer down. Our scan does this for eight AI crawlers, and this guide is the manual version for any crawler you care about. You finish with a table that whoever edits the file can work from.
Before you start
- List every host that serves pages. That includes the bare domain, www and any shop or blog subdomain. Include http as well as https if both answer. Google's specification says a robots.txt file applies only to the host, protocol and port where it is served.
- Open Search Console on a Domain property or on a URL prefix property without a path, such as https://example.com/. The robots.txt report exists only for those.
- Keep the crawler list open. It names each AI crawler and says what it is for.
- Pick five to ten URLs that matter to the business. Include the home page, a product or service page and the sitemap.
Steps
Step 1. Open the file on every host and write down the response
Type the address into a browser, for example https://www.example.com/robots.txt, and repeat it for every other host on your list. A browser hides the status code. With curl, a developer can see it in one command per host.
curl -sI https://www.example.com/robots.txt
The first line of the response is the status code. The Content-Type line shows the format. If the file redirects, the Location line shows where to. Google follows at least five redirects for this file. If it still has not reached the file, it treats the file as a 404. Write down the host, the status and the final address, and date the note.
Step 2. Confirm the status, the format and the size
A healthy file answers 200 with plain text in UTF-8. It also stays under 500 KiB, because Google ignores anything past that point. Any other response changes what Google does.
| What /robots.txt returns | What Google does |
|---|---|
| 200 with plain text | Reads the rules. |
| A 4xx status other than 429, such as 404 | Crawls as if no robots.txt file existed, so nothing is blocked. |
| A 5xx status | Stops crawling the site for 12 hours while it keeps asking for the file. For the next 30 days it uses the last good copy, and with no copy it assumes no restrictions. |
| 200 with an HTML error page | Finds no valid rules, since Google ignores lines it cannot parse. |
Look hardest at the last row. An HTML error page at /robots.txt looks normal in a browser, and crawlers find no rules in it at all.
Step 3. Find the group each crawler follows
A group starts with one or more User-agent lines. The Allow and Disallow rules under them belong to that group. Google follows the one group whose user agent matches the crawler most specifically. It ignores every other group, and the order of groups does not matter. RFC 9309 works the same way. A crawler falls back to the * group only when no group names it.
User-agent: *
Disallow: /checkout
Disallow: /account
User-agent: GPTBot
Disallow: /blog/
This example hides a common mistake. In this file GPTBot follows only its own group. It may crawl /checkout and /account, which the owner probably did not intend. It may not crawl the blog. When you give a crawler its own group, copy in every rule from the * group that should still apply to it.
Step 4. Apply the longest matching rule to each URL
Take one crawler and one URL at a time. In that crawler's group, find every rule whose path matches the start of the URL path. The rule with the longest path wins. When an Allow and a Disallow rule are equally long, Google uses the less restrictive one. That is the Allow.
Paths are case sensitive. A rule for /Blog/ does not block /blog/. The * wildcard matches zero or more characters, and $ marks the end of the URL.
Write each answer in a table with one row per crawler and URL. Record the group, the rule that decides and whether the URL is allowed. Date the table. Keep it next to a saved copy of the file, so the next change has something to compare with.
For a second opinion, a developer can run the file through Google's open source robots.txt parser. Its README says it is the parser Googlebot uses in production, with small changes. Bing Webmaster Tools has a robots.txt tester of its own, which Microsoft announced in September 2020.
Step 5. Decide what each AI crawler may read
Each AI company runs more than one crawler, and each crawler does a different job. Blocking one of them does not block the others.
| Operator | Name in robots.txt | What the operator says it does | Does robots.txt control it? |
|---|---|---|---|
| OpenAI | OAI-SearchBot | Surfaces sites in ChatGPT's search features | Yes |
| OpenAI | GPTBot | Crawls content that may be used to train OpenAI's models | Yes |
| OpenAI | ChatGPT-User | Visits a page for certain actions a ChatGPT user starts | OpenAI says robots.txt rules may not apply |
| Anthropic | ClaudeBot | Collects content that could contribute to model training | Yes |
| Anthropic | Claude-SearchBot | Crawls to improve the quality of Claude's search results | Yes |
| Anthropic | Claude-User | Fetches a page when a Claude user asks a question | Yes, Anthropic says site owners can control it |
| Perplexity | PerplexityBot | Surfaces and links sites in Perplexity search results | Yes |
| Perplexity | Perplexity-User | Visits a page when a Perplexity user asks it to | Perplexity says it generally ignores robots.txt |
| Google-Extended | A token, not a crawler, for Gemini training and grounding | Yes, as a token that Google's crawlers read |
If you want your pages in AI answers, leave the search crawlers unblocked. Those are OAI-SearchBot, Claude-SearchBot and PerplexityBot. Whether to allow the training crawlers is a separate business decision. The operators run them apart from the search crawlers so that you can make it. OpenAI says a change can take about 24 hours to reach its search systems. Perplexity says up to 24 hours. OpenAI's page now also lists OAI-AdsBot, which visits only landing pages that someone submitted as a ChatGPT ad.
Step 6. Read the Search Console robots.txt report
In Search Console, open Settings and then the robots.txt report. It lists the robots.txt files Google found for the top 20 hosts in the property. Each row shows when Google last crawled the file and any warnings or errors. The two can differ, so compare that version with the file you saved in step 1. After an urgent fix, you can ask Google from the same report to crawl the file again.
Step 7. Check the Sitemap line
A Sitemap: line tells crawlers where your sitemap lives. Google wants a full address there. The line can point to another host. Open that address and confirm that it answers 200 and lists your current pages. A line that still names the old domain after a move sends crawlers to the wrong list. Once the line is right, check the sitemap next.
Step 8. Check the layer in front of the site
A firewall, CDN or bot manager answers before robots.txt is ever read. It can refuse a crawler that your file allows, and robots.txt will not show it. Open the bot settings in your CDN or firewall. Look for rules that name AI crawlers.
Cloudflare is one example. Its documentation, last updated July 1, 2026, announced new defaults for new domains starting September 15, 2026. Training and agent crawlers would be blocked on pages that show ads, and search crawlers would stay allowed. When we read the page on September 30, 2026, it still described the change in the future tense. So open Configure AI bot policies in your own Cloudflare dashboard and read what it says.
Anthropic warns that blocking its IP addresses may not work as an opt out, because a crawler that cannot reach your site cannot read your robots.txt either. Your server or CDN logs show whether each crawler's requests received a 200 or a 403. The guide to server logs shows you how to see which crawlers visit.
Check that it worked
Wait a full day after any change. Google may cache the file for up to 24 hours, and OpenAI and Perplexity give the same window. After that day, repeat steps 3 and 4 on the live file and read the Search Console report again.
The check passes when every important URL is allowed for Googlebot, Bingbot and each AI search crawler you want. The report should also show the new file with no errors. A CDN can still hold the old file. The guide on how to confirm the change is live covers the caches in between.
If it did not work
- If a crawler is still blocked, check the other hosts first. Each host serves its own file. The rules on www do not cover the bare domain or a subdomain. Anthropic asks for a rule on every subdomain you want to cover.
- If a rule in the
*group seems to be ignored, look for a group that names the crawler. That group replaces the general one. - If a rule does not match, compare its case and path with the live URL. One capital letter or one missing slash is enough to make the rule miss.
- If the file looks right in your browser and wrong to crawlers, your CDN may serve them another copy. Compare yours with the version in the Search Console report.
- If a blocked page still shows up in Google, that is how Disallow works. Google can index a disallowed URL without its content when other pages link to it. Use
noindexinstead. Leave the page crawlable so Google can see the tag. - If ChatGPT or Perplexity still fetches a page you blocked, the request may come from ChatGPT-User or Perplexity-User. Both operators say robots.txt may not govern those user requests. A firewall rule can stop them. It also stops the person who asked about your page.
Platform notes
Shopify generates a default robots.txt for every store. A theme changes it through a robots.txt.liquid file that you add to the theme's Templates folder in the code editor. Shopify strongly discourages replacing the whole file with plain text rules, since it updates the default rules regularly. Add or remove single rules instead, and run steps 3 and 4 on the result before you publish the theme.
WordPress serves a virtual robots.txt through its do_robots function. Yoast's documentation says WordPress generates it only when the site root holds no physical robots.txt file. In Yoast SEO you edit the file under Tools, then File editor, unless file editing is switched off on your site. Many owners expect the box under Settings, Reading that discourages search engines to change robots.txt. Since WordPress 5.3 it adds a noindex robots meta tag to the pages instead, so check that box separately.
Wix, Squarespace, Webflow and BigCommerce each decide where the file is edited, so take the exact place from their own help pages. The steps above still work on any platform, because they read the file your site serves.
What Visibility Mesh checks and what it does not
The Visibility Mesh scan reads your robots.txt for eight crawlers, one name at a time. It reports which of them may read your pages. Checking one name at a time means one crawler's group never hides another's. The scan also reports Google-Extended separately, as a control for Gemini training and grounding. The page on what the scan reads on each platform lists all of them. Some firewalls respond differently when a request comes from a crawler's real network addresses. Your own logs show that, and step 8 covers them. On Shopify and WordPress, AI Visibility Setup: Foundation does this work for you. Its first job is opening the site to the crawlers you choose (how the setup works).
Questions people ask
What is the difference between noindex and a robots.txt Disallow rule?
A Disallow rule asks crawlers not to fetch a URL. A noindex rule asks Google not to show the page in results. Combined, they work against each other. Google has to crawl a page to see its noindex tag, so a page blocked in robots.txt hides its own noindex. Google can also index a disallowed URL without its content when other pages link to it. To keep a page out of Google, use noindex and leave the page crawlable.
What happens if my robots.txt file returns a 404 or a server error?
A 404, or any 4xx status other than 429, makes Google crawl as if the site had no robots.txt file. Nothing is blocked. A 5xx server error makes Google stop crawling the site for 12 hours while it keeps asking for the file. For the next 30 days Google uses the last good copy it has. Server errors on this file can pause crawling of the whole site, so fix them before anything else.
Which rule wins when Allow and Disallow both match a URL?
The rule with the longer path wins, since Google treats it as the more specific one. When both rules have the same length, Google applies the less restrictive rule, which is Allow. The rules only compete inside the group the crawler follows. A rule in the * group never applies to a crawler that has its own group.
Does noindex work inside a robots.txt file?
Google does not support a noindex line in robots.txt. It retired all code for that rule on September 1, 2019. Put noindex in a robots meta tag in the page head, or send it in an X-Robots-Tag HTTP header. Either one works only if Google may crawl the page.
Should I block GPTBot and the other AI crawlers?
Decide crawler by crawler, based on what each one does. The search crawlers are OAI-SearchBot, Claude-SearchBot and PerplexityBot. They let AI search products find and link to your pages. The training crawlers are GPTBot and ClaudeBot, and blocking them does not block the search crawlers. Google-Extended is a token for Gemini training and grounding. Google says it does not affect a site's inclusion in Google Search and is not a ranking signal.
Related reading: If robots.txt looks fine and crawlers are still refused, we explain whether a WordPress security plugin is blocking AI crawlers.
Sources
- Google, robots.txt specification, Google Search Central, last updated August 31, 2026, read September 30, 2026.
- Google, How HTTP status codes affect Google's crawlers, Google Crawling Infrastructure, last updated February 4, 2026, read September 30, 2026.
- IETF, RFC 9309, Robots Exclusion Protocol, published September 2022, read September 30, 2026.
- Google, Introduction to robots.txt, Google Search Central, last updated December 10, 2025, read September 30, 2026.
- Google, Block Search indexing with noindex, Google Search Central, last updated December 10, 2025, read September 30, 2026.
- Google, A note on unsupported rules in robots.txt, Google Search Central Blog, July 2, 2019, read September 30, 2026.
- OpenAI, Overview of OpenAI crawlers (no update date shown), read September 30, 2026.
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?, Claude Help Center, dated April 7, 2026, read September 30, 2026.
- Perplexity, Perplexity crawlers (no update date shown), read September 30, 2026.
- Google, Google's common crawlers, last updated July 14, 2026, read September 30, 2026.
- Google, robots.txt report, Search Console Help (no update date shown), read September 30, 2026.
- Cloudflare, Block AI bots, last updated July 1, 2026, read September 30, 2026.
- Shopify, Customize robots.txt, Shopify developer documentation (no update date shown), read September 30, 2026.
- WordPress, do_robots() and Settings Reading Screen, last updated June 8, 2024, read September 30, 2026.
- Yoast, How to edit robots.txt through Yoast SEO (no update date shown), read September 30, 2026.
- Google, robotstxt, the open source parser on GitHub, read September 30, 2026.
- Microsoft, Bing Webmaster Tools makes it easy to edit and verify your robots.txt, Bing Webmaster Blog, September 4, 2020, read September 30, 2026.
Once the crawlers may fetch your pages, check what the page sends them. The crawler names and their purposes are in the crawler list.
Run the free scan to see which of the eight crawlers your robots.txt lets in.