How do I find AI crawler visits in my server logs?

Export at least a week of raw access logs from your web server or CDN and filter them for the user agents the AI operators publish. Then check each IP address against the operator's published list or reverse DNS, and count only the verified requests by crawler, status code and page.

By Margareta Petrovic, founder of Visibility Mesh. We checked every statement in this guide against its primary source on September 30, 2026.

Google Analytics will not show you these visits. Google excludes known bots and spiders from Analytics automatically, and it does not let you see how much it excluded. The record of a crawler visit lives in the access log of your server or your CDN.

Any client can send a crawler's user agent, so the IP address is what proves a visit came from the operator. Common Crawl's own page warns that other crawlers pretend to be CCBot, which tells you how popular the costume is.

Before you start

  • Get raw access logs from the web server or the CDN, covering at least seven days. Ask your host how many days it keeps.
  • Check that each line holds five fields: the time, the client IP address, the user agent, the requested URL and the status code. The Apache combined log format has all five, and nginx writes that format by default.
  • Keep the crawler list open. It says what each crawler is for.
  • Write down the pages that matter to the business, such as your top products, services or articles.

Steps

Step 1. Find where the logs live

A request can stop at the CDN, at a firewall or at the origin server. A CDN answers cached pages itself, and the origin log never sees those requests. If a CDN sits in front of your site, its logs are the fuller record. Check whether the IP column in the origin log holds the visitor's address or the CDN's address. Step 4 needs the visitor's.

Step 2. Export a fixed window

Pick a window of at least seven days and export every line in it. Keep every field for now. Write down the start, the end and the time zone. You will compare the next export with this one.

Step 3. Filter for the documented user agents

Search the user agent field for the names the operators publish. OpenAI documents GPTBot, OAI-SearchBot and ChatGPT-User. Anthropic documents ClaudeBot, Claude-SearchBot and Claude-User. Perplexity documents PerplexityBot and Perplexity-User. Add Googlebot and bingbot for search, and Applebot, Amazonbot and CCBot if Apple, Amazon or Common Crawl matter to you.

grep -Ei "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Googlebot|bingbot|Applebot|Amazonbot|CCBot" access.log > crawlers.log

OpenAI says its crawlers may add a robots.txt marker to the user agent when they fetch robots.txt. Count those lines on their own, since they are rule checks and not page reads. Google-Extended will never appear. Google documents it as a robots.txt token with no user agent string of its own.

Step 4. Verify the addresses

Check each address against the operator's own source. Put every request that fails into an unverified pile, and keep that pile out of your counts.

Sources: each operator's own crawler page, read September 30, 2026. The Bing method comes from the Bing Webmaster Blog post of August 31, 2012, and the Verify Bingbot tool page.
Crawler How the operator says to verify it
GPTBot, OAI-SearchBot, ChatGPT-User Published lists at openai.com/gptbot.json, openai.com/searchbot.json and openai.com/chatgpt-user.json
ClaudeBot, Claude-SearchBot, Claude-User Anthropic's list at claude.com/crawling/bots.json
PerplexityBot, Perplexity-User Published lists at perplexity.com/perplexitybot.json and perplexity.com/perplexity-user.json
Googlebot and other Google crawlers Reverse DNS to googlebot.com, google.com or googleusercontent.com, then a forward lookup, or Google's published lists
bingbot Reverse DNS to a name that ends in search.msn.com, then a forward lookup, or Bing's Verify Bingbot tool
Applebot Reverse DNS in applebot.apple.com, or Apple's published list
Amazonbot Amazon's published address pages
CCBot Reverse DNS in crawl.commoncrawl.org, or index.commoncrawl.org/ccbot.json

The DNS check takes two commands. Google publishes this pair as its example.

host 66.249.66.1
host crawl-66-249-66-1.googlebot.com

The first command should return a name in the operator's domain. The second should return the address you started with. For the operators that publish lists, match the address against the ranges in the operator's published file. For Anthropic, a match against its published list is proof enough.

Step 5. Count by crawler, status and page group

Group the verified requests by crawler, by status code and by page group. Products, categories, articles and checkout work well as groups. A pivot table is enough. Keep the robots.txt fetches in their own row.

Step 6. Flag blocked responses to verified crawlers

Look for 403, 429 and 5xx responses to verified crawlers. Google says 5xx and 429 responses make its crawlers slow down. It also says 403 should not be used to limit crawling. When a verified search crawler gets 403 on a public page, look first at the firewall or bot rule in front of the site. The guide on how to test what a crawler receives covers the next check.

Step 7. Compare the pages requested with the pages that matter

Put your list of important pages next to the requested URLs. A product line that no verified crawler requested in seven days is a gap. Look for the cause in robots.txt first, then in the internal links and the sitemap. The robots.txt guide shows how to check what robots.txt allows.

A ChatGPT-User or Perplexity-User line is as close as a log gets to a person's question. Both operators say a person's request triggered that fetch.

Step 8. Repeat after any change

Run the same export over the same length of window after every robots.txt, firewall or CDN change. OpenAI, Perplexity and Amazon each allow about 24 hours for a robots.txt update to reach their systems, so wait a day before you compare.

Check that it worked

Pick one crawler and one day, and count its verified lines by hand. The total should match your summary table. Verified search crawlers should also have received 200 on the public pages you meant to open. Keep the summary with its window and time zone, so you can compare next month's export on the same terms.

If it did not work

  • If the host keeps only a few days of logs, ask for longer retention. An export on a schedule builds the window up over time.
  • If the origin log shows fewer crawler requests than you expected, a CDN may be answering from cache, and its logs hold the rest. Cloudflare marks a cached response HIT and an origin response MISS in its CF-Cache-Status header.
  • If the counts jumped overnight, check the addresses before you trust the jump. Common Crawl says some crawlers falsely identify themselves as CCBot, and the same trick works with any name.
  • If a report treats crawler visits as proof of citations, keep the two apart. A log shows a request and the response it got, and it cannot show whether an assistant quoted the page later.
  • If the firewall blocks a crawler by IP address, put the policy in robots.txt instead. Anthropic says blocking its addresses may not work as an opt out. The block gets in the way of Anthropic reading your robots.txt.

Platform notes

On WordPress, the access log belongs to the web server and not to WordPress. Ask your host where it keeps the file and for how long.

We found no Shopify document that gives merchants raw server access logs, so ask Shopify support whether you can get them before you plan around it.

On any platform where you control the CDN, its logs can fill the gap. Every Cloudflare plan includes AI Crawl Control. It shows which AI services access your content and whether they follow robots.txt.

What Visibility Mesh checks and what it does not

The Visibility Mesh scan reads your robots.txt for eight crawlers, one at a time. They are GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User. The scan also reads your pages the way a crawler receives them. Your own logs hold the record of who visited, and the steps above show you how to read them. If you would rather have a person read them with you, describe the problem on our contact page. We will reply within one business day.

Questions people ask

Why do AI crawler visits not show up in Google Analytics?

Google Analytics excludes traffic from known bots and spiders automatically. Google says you cannot turn that off or see how much was excluded. So crawler visits live in server and CDN logs. People who click a link in an AI answer are a different matter. The guide to AI referral visits in Google Analytics shows where those visits land.

How can I see AI crawler activity if my host does not give me server logs?

Ask the host whether it keeps access logs that the dashboard does not show, and ask for an export. If the site sits behind a CDN you control, read the CDN's logs or its bot analytics, such as Cloudflare's AI Crawl Control. Without either source you cannot verify crawler visits. A scan can still tell you whether crawlers are allowed in.

Which fields do I need in my server logs to analyze crawler visits?

You need five: the time, the client IP address, the user agent, the requested URL and the status code. Keep the IP address above all, since every verification step depends on it. The Apache combined log format carries all five, and nginx writes that format by default.

Can my server logs show whether ChatGPT cited my page?

They cannot. A log shows that a crawler requested a page and what status it received. A ChatGPT-User line tells you that a person's question led ChatGPT to fetch the page at that moment. Treat it as a sign of demand. Whether the answer then cited the page is not in the log.

Sources

Every page below was read on September 30, 2026.

The Academy has the crawler list, with what each crawler feeds, and a guide to AI referral visits in Google Analytics for the people who click through.

Run the free scan, and we read five of your key pages the way an AI crawler reads them, with no card and no call.

See whether this applies to your site

This article is about whether AI can find you; the first category of the scan checks exactly that on your site.

The free Visibility Mesh scan checks whether AI crawlers can reach your website and reads the structured data and text of 5 key pages. The scorecard has 3 complete findings, and each one names the page and the fix. No install, no call, no card.

Run the free scan