Faceted navigation: which filters should you let Google index?

Most ecommerce platforms generate thousands to tens of thousands of filter URLs, yet only a fraction of them match real search demand. I use demand, product depth and page uniqueness to decide which combinations to index, which to consolidate under a canonical URL and which not to let be crawled at all. It is a core part of effective ecommerce SEO.

Cover illustration for the article on which filter URLs to let Google index in faceted navigation.

Before you touch the first robots.txt setting, you need two baselines. A full crawl of the site shows which URLs exist at all and which status code each one returns. I also recommend setting up a Search Console export to BigQuery, because it gives you a data history longer than 16 months. Then I can clearly see which of those URLs get impressions and clicks from search. The first baseline shows what really exists on the site. The second shows which of it people search for and click on.

Why filters flood the index and how to deal with it

Faceted navigation multiplies the URLs a site already has. Every filter you create (color, size, brand, price, availability and so on) is a parameter, and parameters combine with one another. An online store category with ten filters and a few values each does not produce dozens of variants but potentially thousands of possible combinations, which search engines have to crawl and evaluate to decide whether to index the page.

If this is set up badly, it can become a problem and a heavy load for search engine bots, which keep revisiting those pages. Ideally, the online store’s category structure, filters and product parameters are defined up front. That lets you decide in advance which combinations deserve their own landing page.

Crawl budget

The load on crawl budget is easy to underestimate. In audits of mid-sized online stores, I regularly find that a large share of crawled URLs are filters with no reason to be indexed. A focused technical SEO analysis protects the site’s crawl capacity and keeps weak variants from weakening the pages that generate revenue.

Google itself writes on the Search Central blog that faceted navigation is by far the most common source of overcrawling issues site owners report to it, and that in the vast majority of cases it could have been avoided by following a few basic rules.

How I decide what gets indexed

Before I change robots.txt or a single meta tag, every filter pattern goes through four checks. Search demand comes first, because every later decision depends on whether people search for that combination at all.

  1. Demand. Use keyword data and Search Console to confirm that people actually search for the combination. A pattern without proven demand usually has no reason to get its own indexed page. That is also why a blanket noindex before you have mapped demand is a mistake. You can unknowingly cut off pages that used to bring in traffic.
  2. Product depth. Set a minimum number of products a combination has to meet. A filter that returns only a few items creates a thin landing page.
  3. Distinct value. Compare the filter page with its parent category. A near-duplicate page competes with the stronger category instead of covering something the category does not.
  4. Technical rule. Assign one treatment to each pattern and build it into the templates. The practical options are index, noindex, canonical or a crawl block in robots.txt.
Decision framework for indexing faceted URLs by demand, product depth and page uniqueness.

Recommended treatment by URL pattern

PatternRecommended treatmentWhy
Single high-demand attribute (e.g. brand, size)indexReal demand, distinct results
Two-attribute combinationindex, depending on the combination (black T-shirt, size S)Some value, but search volume may be lower
Three or more attributesDepends on the combination: index if it makes senseThere may be combinations nobody searches for
Sort and view parametersblock or canonicalizeNo unique content
Paginationself-canonical, keep crawlablePage two carries different products, so it is not a duplicate
Empty filter resultsreturn 404Google recommends this directly, see below
Sold-out productsdepends on whether they come back, otherwise noindex or 410A short-term outage is no reason to remove the page

This table is my decision rule, not Google’s. Google does not say anywhere how many attributes are too many. It only says which tools you have and what they do. Where you draw the line is up to you, and it should be based on your own demand data.

Careful: noindex does not save crawling. Google fetches the URL and only then drops it. What directly saves crawl capacity is a disallow in robots.txt. And once a page has been noindexed for a long time, Google stops crawling it altogether, so the links on it stop helping.

The rule belongs in the templates, not in an analysis document. It also has to carry through to the links. A pattern you have excluded should not be linked from the filter in every category, and this is exactly where internal linking rules and indexation rules meet.

Shared template rule for processing faceted URLs.

Four things Google says outright

Once you decide to let filter URLs be crawled, Google has specific technical recommendations for them. They are dull, easy to overlook, and yet they decide whether your decision has any effect at all.

  • Use the standard parameter separator, which is &. Google may not recognize a comma, semicolon or brackets as a separator.
  • Keep the filter order in URLs consistent. When the same combination generates two different parameter orders, you have created duplication where none had to exist.
  • Return a 404 status code for combinations with no results. Google treats an empty page returning 200 as a soft 404 and will keep crawling it.
  • Do not redirect empty results to a generic “not found” page. Return the 404 on the URL itself. The only exception is when that is technically impossible, for example in a single-page application.

And one alternative that often gets forgotten. Filters can be handled through the URL fragment after the # sign, because search engines usually ignore fragments. It is a different route from robots.txt, and on new builds it is worth considering before you start producing thousands of parameter URLs and then blocking them again.

What to avoid

  • Do not put a blanket noindex on all filters without demand data. One online store put a blanket noindex on all its filters and lost its “brand + category” pages, which were bringing in orders. It is the fastest way to quietly lose pages that make money. Data and a Search Console export first, rules second.
  • Do not block a URL pattern in robots.txt once it is already indexed. Google then will not see the noindex, and the URL may stay in the index. The order is noindex, wait, and only then disallow crawling in robots.txt.
  • Do not let the price filter get indexed. A “from–to” price filter generates a URL for every slider value. The result is tens of thousands of URLs with the same product range. Do not index the price filter, and do not let it be crawled either.
  • Do not index sorting and items per page as separate URLs. Sort order and the number of items per page get their own URLs and end up indexed as duplicates of the category. The fix is a canonical to the default view, or a fragment.
  • Do not let the parameter order vary. The same combination appears in two orders (color + size, size + color) depending on what the user clicked first. The template must always order parameters the same way.
  • Do not send conflicting signals. A filter has a canonical to the category, yet every category links to it and it sits in the sitemap. Google gets conflicting signals. What should not be in the index should not be in the sitemap or in internal links either.
  • Do not write copy for filter pages before the canonical is resolved. You will make the duplication bigger, not solve it.
  • Do not treat indexation and internal linking as two separate decisions. It is one decision, written into two places in the code.
  • Do not take this on for a brochure site with a few dozen URLs. Google itself says crawl budget matters for sites with a million pages, or with ten thousand pages that change daily. Faceted navigation, however, produces that many URLs even for a smaller online store, which is why this approach is for online stores, catalogs and marketplaces. On a small site, the problem is almost certainly somewhere else.

Questions I get before an audit

Does robots.txt remove filter URLs from the index?

No. It only blocks crawling, so a URL that is already in the index can stay there even after you disallow it. The order matters. If a pattern is already indexed, noindex it first, wait for the URLs to drop out of the index, and only then block the pattern in robots.txt. If you block first, you risk freezing the URLs in the index, because Google will no longer see the noindex that would have taken them out.

How quickly will I know I blocked too much?

Within roughly 48 hours or more, from your own repeat crawl. This threshold comes from my audit practice, not from a Google number, and it exists for two reasons. First, you compare it with the baseline crawl and immediately see which indexable URLs you blocked by mistake. A drop in rankings would only show up weeks later. Second, Search Console is much slower here, because the Page indexing report updates with a delay and a noindex only takes effect after Google’s next crawl. When you run the crawl yourself, you catch an overly broad rule while the fix is still cheap.

Should filter pages have their own copy?

Not before the canonical is resolved. Mass-generating filter copy at this stage only makes the duplication problem bigger instead of solving it.

Three numbers that show it worked

  1. Page indexing by pattern in Search Console. Is each pattern handled the way you intended?
  2. Crawl stats. Is crawling shifting toward the URLs that matter?
  3. Impressions and clicks for the patterns you kept indexed, to confirm the pages you protected still perform.

If crawled and indexed URLs drop for the patterns you excluded, and impressions and clicks on the protected pages hold, the cleanup has done its job. Googlebot and other bots stop visiting combinations nobody searches for, and the pages that make money lose nothing.

That is the whole difference between cleaning up filters and switching them off across the board. Switching them off is fast, looks like work done, and sometimes costs you revenue you never find out about. A data-based decision takes a few days longer.