There is a new fleet of crawlers reading ecommerce, and they do not behave like Googlebot. GPTBot, ClaudeBot, PerplexityBot and a handful of others are the agents deciding whether your store gets described to a shopper who never types a single search query. Know their names before you decide who gets in.
- A new fleet of crawlers reads ecommerce, GPTBot, ClaudeBot, PerplexityBot and others, and they do not behave like Googlebot.
- Each reads differently: some skip JavaScript, some ignore images, some only sample a few pages before moving on.
- You decide who gets in, by user-agent name, in robots.txt. Block the wrong one and you turn away the engine you wanted to be seen by.
- Allowing a crawler is necessary but not sufficient: it lets you be read, it does not guarantee you are quoted.
For twenty years there was effectively one robot that mattered, and an entire industry grew up around keeping it happy. That era is over. The crawlers that now shape whether your products get recommended answer to different companies, follow different rules, and announce themselves with different names in your server logs. If you are still managing for one bot, you are managing for the wrong decade.
Meet the fleet
These are the user agents to know by name. Each is a real, declared crawler you can allow or block:
- GPTBotOpenAI’s crawler for training. OAI-SearchBot is the separate one that powers ChatGPT’s live search answers. Blocking one does not block the other.
- ClaudeBot collects content that may be used to train Anthropic’s models. Anthropic runs Claude-SearchBot for Claude’s search results and Claude-User for pages Claude fetches when someone asks it a question.
- PerplexityBotPerplexity’s crawler, which leans heavily on live retrieval, so being reachable matters more here than almost anywhere.
- Google-Extendednot a crawler at all, but a robots.txt token that controls whether Google may use your content to train Gemini models and to ground answers in Gemini Apps and Vertex AI. It does not affect whether your site appears in Google Search.
- CCBotCommon Crawl, the open dataset that quietly feeds dozens of downstream models.
The list grows every quarter. The point is not to memorise it. It is to stop treating “AI crawlers” as one anonymous mass you either welcome or fear.
Why they do not act like Googlebot
Googlebot is patient. It renders JavaScript, comes back repeatedly, and has years of context on your store. The newer agents are often the opposite: lighter, faster, less forgiving. Some take a single look at the raw HTML and move on. If your product grid only appears after a script runs, a crawler that does not run scripts sees an empty page and describes it as empty. That is covered in our piece on the rendering trap that hides your products.
| Crawler | Operator | What it feeds | Typical user-agent |
|---|---|---|---|
| GPTBot | OpenAI | Training OpenAI models | GPTBot |
| OAI-SearchBot | OpenAI | ChatGPT search results | OAI-SearchBot |
| ClaudeBot | Anthropic | Training Anthropic models | ClaudeBot |
| PerplexityBot | Perplexity | Perplexity answers | PerplexityBot |
| Google-Extended | Gemini training and grounding (you can opt out) | Google-Extended |
You decide who gets in
Access starts at one file: your robots.txt. A single stray line there can wave off the exact crawler you wanted shopping your store. The mistake almost nobody catches is blocking a bot you meant to welcome, or assuming a block applies to a bot it never named. This is the front door, and most merchants have never read it. The full picture of who reaches your store lives in our cornerstone on whether AI crawlers can even get in.
A scan is a snapshot; legibility drifts
A correct crawler policy is never settled. A theme update rewrites your robots or templates, an app injects a script, a migration spawns redirects, and the layer a crawler reads regresses silently while the storefront still looks perfect to you. Your store changes weekly, so being reachable and readable is a moving target, not a one-time pass. That is why serious stores measure, fix, and re-measure, and why we re-scan our own store on a schedule, in public.
Questions people actually ask
Should I block AI crawlers from my store?
For most ecommerce stores, no. Blocking the crawlers that power AI shopping answers removes you from the exact place buyers are starting to ask for recommendations. There are narrow cases for blocking training crawlers while allowing search and retrieval ones, which is why the distinction between, say, GPTBot and OAI-SearchBot matters.
How do I see which AI crawlers are visiting me?
They identify themselves in your server logs by user-agent string, so a log review shows who is reaching you and how often. Many merchants are surprised to find the bots they assumed were visiting are being turned away at robots.txt, or never arriving because the store is too slow to finish reading.
Is allowing AI crawlers the same as ranking in AI answers?
No. Access is the entry ticket, not the prize. A crawler can reach your store and still find nothing it can confidently describe, because the page is unreadable to a machine. Getting in is step one; being legible once inside is the rest of the work.
See what a machine sees
You cannot tell from your browser whether AI can read your store. You can find out in a few minutes. Run a free scan and see the exact layer the machine reads, and where you are losing the shortlist.
Sources: Similarweb referral data reported by Digiday on September 25, 2025; the robots.txt files of amazon.com, walmart.com and etsy.com, read September 27, 2026; OpenAI (early 2026) on weekly ChatGPT usage; operator documentation for crawler user-agents. Figures are third-party and current as of mid-2026; we publish our own benchmark data as our scan volume grows.
Your buyers are already asking AI. This is how you make your website readable to the assistants they ask.
Everything in this article is measurable on a live storefront. The Full Assessment and Roadmap reads your website the way AI crawlers receive it and hands you every fix in plain English, in the order we would make them.