Related Products, Faceted Nav, and Pagination: Where AI Gets Lost in Big Catalogs

Related Products, Faceted Nav, and Pagination: Where AI Gets Lost in Big Catalogs

Faceted navigation and pagination hurt AI crawling when every filter and sort order becomes its own URL, and page numbers add more on top. On a big catalog those combinations multiply into thousands of near duplicate pages. A crawler can spend its budget on them while your real products wait unread behind the flood.

By Margareta Petrovic, founder of Visibility Mesh. We measure how legible ecommerce stores are to AI, and publish what we find. Updated September 2026.
Key takeaways
  • Filters and infinite pagination can quietly spawn thousands of near-duplicate URLs that drown your real pages in noise.
  • On a big catalog, faceted nav is where crawl budget goes to die, a machine spends it on filter permutations, not products.
  • Traffic control is the answer: canonical tags, noindex on filter pages, and robots rules to contain the sprawl.
  • On a large Shopify store this matters most; small catalogs rarely generate enough combinations to hurt.

Faceted navigation, the filters for color, size, price and brand, is great for shoppers and dangerous for crawlers. Every combination of filters can generate its own URL, and the combinations multiply fast. A catalog with a handful of filters can produce tens of thousands of filtered URLs, most of them thin, overlapping, and near-identical. A crawler with a finite budget can drown in them before it reaches your real pages.

Why URL sprawl loses the machine
41%
of 216 Shopify stores failed Front Door in our 2026 study, the most basic test of whether an AI crawler can read the page at all.
58%
failed Mesh Integrity, so their collections and products did not form a graph a machine can follow.
97.7%
of reachable homepages in our study of non ecommerce websites served a robots.txt, the file where crawl rules for filter URLs live.

Why it loses the machine

A crawler does not know in advance which of those filtered URLs matter. It can spend its limited budget wandering combinations such as red + size 9 + under $100 + page 4 instead of reading the products and collections you actually care about. Pagination chains add to it: page after page of the same collection, each a separate URL competing for attention. The signal-to-noise ratio collapses.

Faceted nav & pagination: the duplicate-URL explosion. Filters and pagination can spawn thousands of near-duplicate URLs. COLLECTION /boots /boots?color=brown /boots?size=9&color=brown /boots?sort=price /boots?page=7 /boots?color=brown&size=9&sort=price /boots?size=10 /boots?color=black&page=3 /boots?sort=new&size=11 /boots?page=12&color=brown /boots?size=8&sort=price /boots?color=tan /boots?color=brown real pages drown in filter-generated noise TRAFFIC CONTROL canonical → /boots noindex filters · robots rules On a big catalog, faceted nav is where crawl budget goes to die. Tame the URLs. VISIBILITY MESH URL SPRAWL CONTROL VM-S-P3-09 · r1.0 CAN AI FOLLOW YOU?

Traffic control for big catalogs

The fix is telling crawlers which of these URLs to ignore so they spend their budget on the pages that matter. That is the job of canonical tags (pointing filtered variants back at the real collection) and careful crawl directives. It connects straight to sitemap health, because your sitemap should list the real pages and leave the filtered noise out.

Faceted-nav pattern URLs it can spawn The control
Each filter as a crawlable URL One per filter combination. Thousands Canonical to the base collection
Sort orders as separate URLs A copy per sort option noindex the sort variants
Deep numbered pagination One URL per page, indefinitely Robots rules / sensible limits
Unmanaged, filters bury your real pages under near-duplicates. Canonicals and noindex are the traffic control.

On a large Shopify store

If your catalog is small, this is rarely a problem. If it is large, with lots of filters, it is one of the biggest hidden drains on crawlability, and exactly the kind of thing that looks completely fine to you while quietly starving the machine of your real pages. Keeping the path clear is part of being a store AI can follow without getting lost.

What our own Shopify store does with filter, sort and page URLs

We checked our own store on September 26, 2026, with plain requests, the way a crawler fetches a page. The robots.txt it serves tells every crawler to skip collection URLs that contain sort_by and collection URLs that combine 2 or more filters. It also blocks tag combinations joined with a plus sign.

A URL with a single filter is not blocked, so a crawler can still reach it. On our store that page carries a canonical tag pointing at the plain collection, and the sort_by version does the same. Page 2 of our catalog behaves differently. It returns a normal page with a canonical pointing at itself and a prev link back to page 1, so each page number stands as its own URL.

A store can customize its robots.txt and its theme can change these tags, so check yours rather than assume it matches ours.

  1. Open your store's robots.txt at /robots.txt and look for the lines that mention sort_by and filter. If they are missing, someone has customized the file, so find out who changed it and why.
  2. Open a collection, apply 1 filter, and view the page source of the address you land on. Search it for canonical and confirm it points at the plain collection address.
  3. Do the same for page 2 of your largest collection. A canonical pointing at page 2 itself is what our store shows, and page 2 should list different products from page 1.

A scan is a snapshot. Legibility drifts

Contained faceted nav is never settled. A redesign reshuffles the menu, an app changes URL patterns, a bulk edit orphans a page, and the structure a crawler follows regresses silently while the storefront still looks perfect to you. Your catalog and navigation change weekly, so a clean, followable structure is a moving target, not a one-time pass. That is why serious stores measure, fix, and re-measure, and why we re-scan our own store on a schedule, in public.

Questions people actually ask

Why is faceted navigation a problem for AI crawlers?

Because each combination of filters can generate its own URL, and the combinations multiply into thousands of thin, near-duplicate pages. A crawler with a finite budget can exhaust it wandering those filtered URLs before it reaches your real products and collections.

How does pagination affect crawlability?

Long pagination chains create page after page of the same collection, each a separate URL competing for a crawler's attention. Combined with faceted URLs, they lower the ratio of meaningful pages to noise, which can leave important pages under-read.

Do I need to worry about this on a small store?

Usually not. With few products and few filters, the duplicate-URL explosion stays manageable. The problem becomes serious on large catalogs with many filters, where it is one of the biggest hidden drains on how well AI can crawl your store.

Does Shopify block filter and sort URLs from crawlers?

On our own Shopify store, checked in September 2026, the robots.txt blocks collection URLs with sort_by and URLs that combine 2 or more filters. A single filter URL stays open to crawlers, but its canonical tag points at the plain collection. Check your own robots.txt, because a store can customize it.


See what a machine sees

You cannot tell from your browser whether AI can read your store. You can find out in a few minutes. Run a free scan and see the exact layer the machine reads, and where you are losing the shortlist.

Run my free scan →

Sources: Visibility Mesh, The State of AI Visibility 2026 (216 Shopify stores, scored June 2026), The State of AI Visibility on the Non Ecommerce Web (July 2026) and our own store's robots.txt and collection pages, read September 26, 2026. Every figure on this page was measured by us.

Your buyers are already asking AI. This is how you make your website readable to the assistants they ask.

Everything in this article is measurable on a live storefront. The Catalog Assessment reads every product in your Shopify or Matrixify export, up to the plan limit, and returns catalog wide figures on what AI systems can read, with a fix list in priority order.

Catalog Assessment See all assessments

See whether this applies to your site

This article is about what AI can read; the scan checks your structured data page by page.

The free scan reads 5 pages of any website as AI crawlers receive them and returns a scorecard with 3 complete findings, each naming the page and the fix. No install, no call, no card.

Run the free scan