How do I look at the raw HTML a page sends to a crawler?

Open the page source in your browser, or fetch the URL with curl, and save what comes back. That saved file is the raw HTML your server sent before any script ran. A crawler that does not run JavaScript receives exactly that file. The browser's Inspect panel shows the page after scripts changed it, so it cannot show you what that crawler receives.

By Margareta Petrovic, founder of Visibility Mesh. We checked every statement in this guide against its primary source on September 30, 2026.

Most of the checks in this Knowledge Base start from this saved file. Before you fix anything, look at what the crawler receives. The browser will not show you this, because it runs every script and fills every gap before you look.

Google renders JavaScript, and its own documentation says not all bots can. OpenAI, Anthropic and Perplexity do not say on their crawler pages whether their crawlers run it. We read all three on September 30, 2026. So plan for the crawler that reads only the first response. A product page whose raw HTML says only Loading is telling that crawler the truth, and nothing else.

Before you start

  • Pick one live URL per template. For a store that means the home page, a product page, a collection page and an article. For a service business, use a service page instead of the product page.
  • Open a browser. A command line with curl helps, but it is optional.
  • If you have Search Console access, sign in. Step 3 uses its URL Inspection tool to read Google's own copy of the page.
  • Keep two terms straight. Raw HTML is the document your server sends in answer to the request. Rendered HTML is what the browser builds from it after the scripts run.

Steps

Step 1. Open the page source and save it

In Chrome, press Ctrl+U on Windows or Command+Option+U on a Mac. In Firefox, right click the page and choose View Page Source. The shortcut is Ctrl+U, or Command+U on a Mac. Firefox also accepts view-source: typed in front of the address. Safari needs one setting first. Open Settings and then Advanced, and select Show features for web developers. Show Page Source is then in the Develop menu.

Select all the text and save it as a file named after the template and the date, for example product-2026-09-30.html.

Step 2. Fetch the same URL with curl

A developer can save the headers and the body in one command.

curl -sL -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64)" -D headers.txt -o page.html https://www.example.com/products/blue-coat

The -L option follows redirects and -A sets the user agent. The -D option writes the response headers to headers.txt, and -o writes the page to page.html. The command sends a browser user agent on purpose, since some firewalls answer curl's own user agent with a block page. A block page tells you about the firewall rather than the page.

Step 3. Read Google's copy in URL Inspection

In Search Console, paste the URL into the inspection bar at the top. View crawled page shows the HTML Google received when it last crawled the page. Test live URL, followed by View tested page, fetches the page again. It shows the rendered HTML and a screenshot. Google offers the screenshot only in the live test.

Google's help page on rendered source says the browser's page source shows the code before scripts run, and the live test shows it after. When a fact appears in Google's rendered copy and not in your saved file, that fact depends on JavaScript.

Step 4. Search the saved file for six facts

Open the saved file in a plain text editor and search for each item in the table. On a service site, use the main offer and the contact details in place of the price.

Some editors search case sensitively, so try both cases before you mark a fact as missing.
Fact What to search for Why a crawler needs it
Title <title It names the page in results and in AI answers.
Main heading <h1 It states what the page is about.
Main text A sentence copied from the page It proves the body text arrives without scripts.
Price and availability The price digits and the stock wording A shopping answer quotes them.
Canonical link rel="canonical" It names the address you want credited.
Structured data application/ld+json It carries the facts in machine form.

Step 5. Read the response basics in the headers

Open headers.txt. The status line should say 200. If the page redirected, look for the Location lines, since each one shows where a hop went. The Content-Type header should say text/html. Look for an X-Robots-Tag header as well. Google accepts a noindex rule in that header, and it applies even when the HTML looks clean.

Step 6. Record each fact in one of three columns

For every template, mark each fact as present in the raw HTML, present only after rendering, or missing. Present only after rendering means you found it in Google's rendered copy or in the browser's Inspect panel and not in the saved file. That column is your repair list. Theme updates and new apps can move content out of the first response, so rerun this check after each one.

Check that it worked

The check passes when every fact on the list is in the saved file. Google's rendered copy from the live test should hold the same facts. If both are true for every template, your pages do not depend on scripts for the facts a crawler needs.

If it did not work

  • If your saved file has content that curl did not receive, you were probably logged in. Use a private window or curl. Curl sends no cookies unless you give it some.
  • If the saved file is mostly a cookie banner or a consent wall, the site may hold back content until a visitor clicks. A crawler never clicks, so treat that content as missing.
  • If curl receives a page from another country or market, the site redirects by location. The Location header shows where you were sent. URL Inspection shows the page Google received.
  • If curl receives a challenge page, your firewall was answering your command and not a verified crawler. A challenge page proves nothing about what Googlebot or an AI crawler receives. Your server logs show that.
  • If Google's copy looks complete and your saved file does not, go by the saved file for crawlers that do not run scripts. Google's copy is the page after rendering.
  • If you copied the page from the Inspect panel, you saved the rendered page. Go back to step 1.

Platform notes

On Shopify, test the theme's own text and each app's text separately. Reviews, size charts and bundles often come from apps. When an app loads its content with JavaScript, that content will be missing from the saved file even when the theme's text is there.

On WordPress, do the same for each page builder section and each plugin widget. Search the saved file for one sentence from each section.

On a site built as a single page application, the first response can be an empty shell. Google's JavaScript documentation calls this the app shell model, where the initial HTML does not contain the content. If your saved file holds little more than a script tag, the site needs server side rendering or pre rendering for crawlers that do not run scripts.

What Visibility Mesh checks and what it does not

The Visibility Mesh scan renders each page in a browser engine before it scores it, with a plain fetch as a backup. The file you saved with this guide shows the other view, the first response a crawler without JavaScript receives. Run the scan and save the raw HTML with this method, and you have both views of the same page. Hand the differences to your developer as the fix list.

Questions people ask

What is the difference between View Source and Inspect?

View Source shows the HTML your server sent before any script ran. Inspect shows the page as the browser holds it now, after scripts added, removed or changed things. A crawler that does not run JavaScript receives what View Source shows. Use View Source, or a saved curl fetch, to check what a crawler gets.

What does initial HTML mean?

Initial HTML is the first document the server returns for a URL, before the browser runs any script. People also call it raw HTML or the server response. Google renders pages after it crawls them, while a crawler that does not run scripts stops at the initial HTML. Put the facts every crawler must see in the initial HTML.

Does text inside an image count?

Text that exists only inside an image is not in the HTML, so this check marks it as missing. Put the fact in the page text as well. Give the image alt text that describes it.

Can JavaScript set the page title?

Google says you can use JavaScript to set or change the title element and the meta description, and Google reads them after rendering. A crawler that does not run scripts sees whatever title the raw HTML holds. Put the final title in the HTML the server sends.

Do AI crawlers run JavaScript?

OpenAI, Anthropic and Perplexity do not say on their crawler pages, which we read on September 30, 2026. Google documents that it renders pages, and it also says not all bots can run JavaScript. Until the AI companies publish an answer, treat any fact that appears only after rendering as invisible to their crawlers.

Sources

Once you have the saved file, compare the raw and the rendered page. For the reasons behind all of this, read why rendering decides what AI sees. If a crawler cannot fetch the page at all, check what the crawler is allowed to fetch first.

Run the free scan to see what a machine reads on five key pages of your site.

See whether this applies to your site

This article is about whether AI can find you; the first category of the scan checks exactly that on your site.

The free Visibility Mesh scan checks whether AI crawlers can reach your website and reads the structured data and text of 5 key pages. The scorecard has 3 complete findings, and each one names the page and the fix. No install, no call, no card.

Run the free scan