robots.txt for the AI Era: Who You're Blocking Without Knowing It

robots.txt for the AI Era: Who You're Blocking Without Knowing It

robots.txt is the one file that decides which crawlers get into your store, and a single stray line in it can turn away the exact AI agents you want shopping your catalog. Most merchants have never opened theirs. In the AI era, that one file is the difference between being read and being skipped.

By Margareta Petrovic, founder of Visibility Mesh. We measure how legible ecommerce stores are to AI, and publish what we find. Updated July 2026.
Key takeaways
  • robots.txt is the one file that decides which crawlers get into your store, a single stray line can turn away the AI agents you want.
  • Amazon’s robots.txt blocks GPTBot, OAI-SearchBot and ChatGPT-User. OpenAI says sites that block OAI-SearchBot are not shown in ChatGPT search answers. A Shopify store that lets it in can appear there.
  • Shopify manages much of your robots.txt, but you can still add rules that silently block the wrong user-agent.
  • Read yours now: a misplaced Disallow under an AI user-agent is enough to make you invisible to that engine.

robots.txt is the bouncer at your front door. It is the first thing a crawler checks, and it answers one question: am I allowed in, and where. For two decades the only bot anyone worried about was Googlebot. Now the same file quietly governs whether GPTBot, ClaudeBot, and PerplexityBot get to read the store at all, and almost nobody has reread it with the new guests in mind.

Why one file decides reach
<3%
Less than 3% of Amazon’s referral traffic came from ChatGPT in August 2025 (Similarweb data reported by Digiday). When we read Amazon’s robots.txt on September 27, 2026, it blocked GPTBot, OAI-SearchBot and ChatGPT-User.
20%+
More than 20% of Etsy’s referral traffic came from ChatGPT in August 2025, and about 20% of Walmart’s (Similarweb data reported by Digiday). On September 27, 2026, neither site’s robots.txt shut OpenAI’s crawlers out.
−25%
Projected fall in traditional search by 2026 (Gartner). Raising the cost of blocking AI crawlers by mistake.

What it does, in one breath

The file lives at yourstore.com/robots.txt. It lists user-agents and the paths they may or may not crawl. A line like Disallow: / under a bot’s name slams the door entirely. The danger is not malice. It is inheritance: a rule added years ago, or shipped by an app or theme, that now blocks an agent you would happily welcome.

robots.txt: who you are blocking without knowing. One file decides who gets in. AMAZON robots.txt User-agent: OAI-SearchBot Disallow: / ChatGPT search answers leave it out. read September 27, 2026 YOUR SHOPIFY robots.txt User-agent: OAI-SearchBot Allow: / ChatGPT search can show your pages. reachable + readable CHATGPT SEARCH It skips sites that opt out. Amazon closed that door; your store can still walk in. A single stray Disallow line can turn away the exact agent you wanted seen by. VISIBILITY MESH THE AMAZON CASE VM-S-P1-03 · r1.0 CAN AI FIND YOU?

The Shopify-specific part

Shopify generates a sensible default robots.txt for you, and for a long time you could not touch it. You can now, through the robots.txt.liquid template. That is power and risk in the same lever: edit it carefully and you can welcome the AI crawlers explicitly; edit it carelessly and you can lock out the ones driving discovery. If you have apps that inject their own rules, this is where their fingerprints show up.

robots.txt line What it means Effect on AI
User-agent: GPTBot / Disallow: / Keep OpenAI’s training crawler out It tells OpenAI your content should not be used in training. The separate OAI-SearchBot line decides whether ChatGPT search can show your pages.
User-agent: * / Disallow: /checkout Block a path for all crawlers Fine. Keeps private paths out
User-agent: * / Disallow: / Block everything Catastrophic. Nobody can read you
Allow: / (default) Let well-behaved crawlers in Reachable, the readable layer still has to deliver
The difference between appearing in AI shopping and being invisible can be one line.

How to read yours right now

Open yourstore.com/robots.txt in a browser. It is public, on every store, no tools required. Look for any Disallow lines, and for named user-agents like GPTBot or CCBot. If you see a bot you meant to allow sitting under a block, you just found why it is not describing you. Then make sure the doors that are open lead somewhere a machine can read, that is where your sitemap and the broader crawler-access picture come in. And know exactly who the named crawlers are before you decide who to let through.

A scan is a snapshotlegibility drifts

A correct robots.txt is never settled. A theme update rewrites your robots or templates, an app injects a script, a migration spawns redirects, and the layer a crawler reads regresses silently while the storefront still looks perfect to you. Your store changes weekly, so being reachable and readable is a moving target, not a one-time pass. That is why serious stores measure, fix, and re-measure, and why we re-scan our own store on a schedule, in public.

Questions people actually ask

Where is my robots.txt file?

At yourstore.com/robots.txt. It is public on every store, so you can open it in any browser right now without logging in. If you want to change it on Shopify, that is done through the robots.txt.liquid template in your theme.

Should robots.txt block AI crawlers?

For most ecommerce stores, you want it to allow the crawlers that feed AI shopping answers, not block them. Blocking is a deliberate choice for specific cases, not a safe default. The common error is blocking a bot by accident through an old or app-added rule.

Can a robots.txt mistake really make my store invisible to AI?

It can, for the crawler the rule names. A disallowed crawler never fetches the page, so it cannot see anything else you did there. What you lose depends on that crawler’s job. A stray rule against OAI-SearchBot keeps your pages out of ChatGPT search answers. A rule against GPTBot only tells OpenAI not to train on them. Anthropic and Perplexity also run separate crawlers for separate jobs, so read every name in your file before you decide what a rule blocks.


See what a machine sees

You cannot tell from your browser whether AI can read your store. You can find out in a few minutes. Run a free scan and see the exact layer the machine reads, and where you are losing the shortlist.

Run my free scan →

Sources: Similarweb referral data reported by Digiday on September 25, 2025; the robots.txt files of amazon.com, walmart.com and etsy.com and OpenAI’s crawler overview, all read September 27, 2026; Gartner (2024) on traditional search decline. Figures are third-party and current as of mid-2026; we publish our own benchmark data as our scan volume grows.

Your buyers are already asking AI. This is how you make your website readable to the assistants they ask.

Everything in this article is measurable on a live storefront. The Full Assessment and Roadmap reads your website the way AI crawlers receive it and hands you every fix in plain English, in the order we would make them.

Full Assessment and Roadmap See all assessments

See whether this applies to your site

This article is about whether AI can find you; the first category of the scan checks exactly that on your site.

The free scan reads 5 pages of any website as AI crawlers receive them and returns a scorecard with 3 complete findings, each naming the page and the fix. No install, no call, no card.

Run the free scan