How ChatGPT, Perplexity, Copilot and Gemini Each Find Your Store

How ChatGPT, Perplexity, Copilot and Gemini Each Find Your Store. A Visibility Mesh infographic for ecommerce AI visibility.

ChatGPT, Perplexity, Gemini, Copilot and Claude each read the web through their own search crawler or a search index. A page that crawler cannot reach has nothing to offer the assistant. The five differ in which crawler you allow, what else you control and where citations get reported, and none of them publishes how it chooses a source.

By Margareta Petrovic, founder of Visibility Mesh. I am CISA and CISSP certified, and we measure how readable websites are to AI assistants and publish what we find. I last updated this article on .
Key takeaways
  • Each assistant reads the web through its own search crawler or a search index, so access comes before anything else.
  • Google-Extended controls more than training: Google says it also governs grounding in Gemini Apps.
  • Of the five, only Google and Microsoft document a report of AI impressions or citations for site owners.
  • None of the five publishes how it chooses which source to cite.

How do ChatGPT, Perplexity, Gemini, Copilot and Claude differ?

Every cell in the table comes from the platform’s own documentation, read on September 26, 2026. Where a platform says nothing, the cell says so.

Assistant Where it gets web content What you control Where citations are reported Source
ChatGPT OAI-SearchBot and outside search providers. Allow OAI-SearchBot in robots.txt and its IP ranges at your CDN. GPTBot is the separate training crawler. We found no citation report in OpenAI’s documentation. Visits from ChatGPT search carry utm_source=chatgpt.com. OpenAI crawlers, Publishers FAQ
Perplexity PerplexityBot, which gathers and indexes pages for Perplexity search. Allow PerplexityBot and its IP ranges. Perplexity-User fetches pages when a user asks and generally ignores robots.txt. We found no citation report in Perplexity’s documentation. Perplexity crawlers
Gemini and Google AI features The Google Search index. Gemini Apps ground answers with Search content at prompt time. Keep pages indexable, keep the site included under Search generative AI in Search Console, and leave Google-Extended open if you want Gemini grounding. Search Console reports impressions in AI Overviews and AI Mode. Google crawlers, Search Console
Copilot Web search results. Microsoft says Copilot centers its response on high ranking web content. Keep the site crawlable. Bing says it respects robots.txt and other owner controls in its AI experiences. Bing Webmaster Tools reports citations in Copilot (AI Performance, public preview). Copilot transparency note, Bing blog
Claude Claude-SearchBot and a web search tool. Claude-User fetches a page when a user asks. Allow Claude-SearchBot and Claude-User. ClaudeBot is the separate training crawler. We found no citation report in Anthropic’s documentation. Claude crawlers, Claude web search

Sources are linked in the last column. Each page was read on September 26, 2026.

ChatGPT

OpenAI says OAI-SearchBot surfaces websites in ChatGPT search, and sites that opt out are not shown in search answers. ChatGPT also rewrites questions into queries for outside search providers. How to get your website to show up in ChatGPT covers the details.

Perplexity

Perplexity documents that PerplexityBot is designed to surface and link websites in Perplexity search results and is not used to crawl content for AI foundation models. Perplexity-User fetches pages when a user asks, and Perplexity says it generally ignores robots.txt, since a user requested the fetch. Perplexity’s help center says every answer includes numbered citations to the original sources. It does not publish how it ranks or selects them.

Gemini

Google’s page on Google-Extended (last updated July 14, 2026) describes grounding in Gemini Apps as providing content from the Google Search index to the model at prompt time. Gemini Apps Help says Gemini may show a Sources button with links, and that not every response includes sources.

The same Google page adds a detail that changes the robots.txt decision. Google-Extended governs whether Google may use your content for training future Gemini models, and also for grounding in Gemini Apps. A site that blocks it to stay out of training also keeps its content out of grounded Gemini answers. Google adds that Google-Extended does not affect inclusion in Google Search and is not a ranking signal there.

Copilot

Microsoft’s transparency note for Copilot (August 18, 2026) says Copilot decides whether a request needs grounding data from the web. When it does, it aligns the response with search results and provides citations to the pages it used. The Bing Webmaster Tools blog says its AI Performance report shows how often a site is cited in Microsoft Copilot and Bing’s AI summaries.

Claude

Anthropic says Claude invokes a search tool when a question benefits from current information, and its help page says every response that uses it includes citations. Claude can also fetch a specific page when a user gives it a URL.

Anthropic runs three bots, and its crawler page describes each. ClaudeBot collects content for training, Claude-User fetches pages for user requests, and Claude-SearchBot improves search results. Anthropic says disabling Claude-SearchBot or Claude-User may reduce a site’s visibility in Claude’s answers.

How does each assistant choose what to cite?

Google and Microsoft describe a ranking step. Google retrieves pages with its core Search ranking systems, and it may run several related searches at once, which its guide calls query fan out. Microsoft says Copilot centers its answer on high ranking web content. OpenAI says only that ChatGPT ranks search results on multiple factors meant to find relevant, reliable information. Anthropic and Perplexity describe the output rather than the choice: Claude’s search responses include citations, and Perplexity says every answer does.

None of the five lists its citation factors. A factor list you read elsewhere is an observation or a guess, however confident it sounds.

What should you fix first?

  1. Allow the search crawlers in robots.txt: OAI-SearchBot, PerplexityBot, Claude-SearchBot and Claude-User. Decide on the training crawlers, GPTBot and ClaudeBot, as a separate question.
  2. Treat Google-Extended as a training and Gemini grounding decision, not a training decision alone.
  3. Let the IP ranges that OpenAI and Perplexity publish through your CDN or firewall. A robots.txt allow does not open a firewall.
  4. Keep your pages indexable with a snippet, and keep the site included under Search generative AI in Search Console, which is the default setting.
  5. Verify the site in Bing Webmaster Tools and Google Search Console, so you can read the two reports that exist.

The user agents that fetch a page when a person asks are easy to leave out. Any crawler the file does not name falls under the group marked with an asterisk. On Shopify that group is the platform default, and it closes the /policies/ and /search paths. So a file can allow OAI-SearchBot, PerplexityBot and Claude-SearchBot by name and still put ChatGPT-User, Claude-User and Perplexity-User in that default group.

Anthropic says its bots honor robots.txt, so a shopper who asks Claude to read a store’s refund policy would see Claude-User turned away. OpenAI and Perplexity say their user agents may not follow robots.txt at all. Each of the three user agents needs its own named group, and that takes a few lines. Our own robots.txt carries all three.

What our scan checks

The free scan checks the robots.txt rule for each page it reads against eight AI crawler names, each on its own. They are GPTBot, OAI-SearchBot and ChatGPT-User from OpenAI, ClaudeBot, Claude-SearchBot and Claude-User from Anthropic, and PerplexityBot and Perplexity-User from Perplexity. It reports Google-Extended on its own, since Google documents it as a control token with no crawler of its own.

The eight names do not include Googlebot or Bingbot. Search Console and Bing Webmaster Tools report on those two crawlers directly, and they are the better tools for it.

Common questions

Do I need a different strategy for each assistant?

You need one strategy and a few platform checks. The shared work is clear, crawlable pages that state the same facts everywhere. The differences are access and reporting. Allow OAI-SearchBot for ChatGPT, PerplexityBot for Perplexity, and Claude-SearchBot and Claude-User for Claude. Keep the site included in Search generative AI features for Google, and verify it in Bing Webmaster Tools for Copilot. Then read Search Console and Bing Webmaster Tools, the two places that report AI impressions or citations.

How do I get my business recommended or cited by Gemini?

Google does not publish a separate guide for being cited in the Gemini app. It does document that Gemini Apps ground answers with content from the Google Search index at prompt time. So the practical path is the same as for Google Search: an indexed page with clear, specific content and consistent business facts. Do not block Google-Extended if you want your content used for grounding in Gemini Apps. How Gemini chooses between sources is not documented, so treat any list of ranking factors as a guess.

Do AI assistants answer from training data or from a live web search?

They do both, depending on the question. OpenAI says ChatGPT may search the web when a question would benefit from current information. Without search, the answer comes from what the model learned in training. Anthropic says Claude invokes a search tool for topics that benefit from current information. Microsoft says Copilot decides whether a request needs grounding data from the web. Google says its AI features highlight content from its Search index. An outdated fact about a business can come from either route.

Does Perplexity have its own search index?

Perplexity says it collects data with web crawlers that gather and index information from the internet. Its developer documentation adds that PerplexityBot is designed to surface and link websites in Perplexity search results. It does not say whether that index is its only source. The help page we read describes searching the internet in real time and names no provider. The practical step is the same either way: allow PerplexityBot and its published IP ranges.

Which search engine does Claude use?

Anthropic’s help pages, read on September 26, 2026, do not name the provider behind Claude’s text search. They do say that image results in Claude are powered by Bing, and that Anthropic runs its own Claude-SearchBot to improve search results. Some community posts say Claude’s web search uses Brave Search, and we have not confirmed that from an Anthropic source. The documented step for a site owner is to allow Claude-SearchBot and Claude-User.

How do I help Claude find and cite my website?

Start with access, which is the part Anthropic documents. Anthropic says disabling Claude-SearchBot stops its system from indexing your content for search, and disabling Claude-User stops Claude from retrieving your pages when a user asks. It says both may reduce your visibility. So allow both in robots.txt, and treat ClaudeBot, the training crawler, as a separate decision. Anthropic does not publish how Claude chooses sources, so clear pages whose facts other sources confirm are the reasonable bet.

Which assistant matters most for ecommerce?

The one your buyers use, and your own data can show you which that is. ChatGPT tags its referral links with utm_source=chatgpt.com, Search Console reports impressions in Google’s AI Overviews and AI Mode, and Bing Webmaster Tools reports citations in Copilot. We found no citation report for Perplexity or Claude in their documentation, so for those two, watch referral visits and ask the buyer questions yourself.

How do I know it is working?

Measure in the places that report it. Search Console has a generative AI performance report for Google, and Bing Webmaster Tools has an AI Performance report for Copilot. For ChatGPT, count the visits tagged utm_source=chatgpt.com in your analytics. For every assistant, ask the same buyer questions each month and record who gets named. Scan the site again after each change so you can see what changed in what crawlers can read.

Sources

I opened every source below on September 26, 2026. Perplexity’s help center refused one automated request that day with HTTP 403, and a second help page opened normally.

Related reading

For the whole subject, start with AI visibility: what it is, what the work covers and how to measure it. Platform guides cover how Perplexity picks products, how Copilot and Bing find a site, Google AI Mode and how Grok and X search find a store.

Check AI accessibility.
See whether the crawlers of OpenAI, Anthropic and Perplexity are allowed into your pages. The free scan reads up to five of your key pages, and you do not need a card or a call. Run the free scan.
See whether this applies to your site

This article is about what AI can quote; the scan checks whether your pages answer the questions buyers ask.

The free scan reads 5 pages of any website as AI crawlers receive them and returns a scorecard with 3 complete findings, each naming the page and the fix. No install, no call, no card.

Run the free scan