·11 min read
Crawl budget waste: does it affect your site at all?
Crawl budget is not about how many URLs your site has. It is about where Google spends its time when crawling. Googlebot has a certain crawl capacity for each site and at the same time evaluates which URLs are worth visiting and how often. So the problem does not start with having a large site, but at the point where a substantial part of the crawl is used up by duplicates, filters, parameters or other URLs with no real value. Important pages may then get their turn later than you need.

Before you read on, run one test. If your pages are usually crawled the same day you publish them, you can stop here. It is a test Google itself recommends, and it ends most crawl budget debates before they start.
Is crawl budget your problem at all?
Google’s crawl budget documentation is unusually direct about who it is for. It names three site profiles. Large sites with more than a million unique pages whose content changes roughly weekly. Medium sites with ten thousand or more pages that change daily. The third group is sites where a large share of URLs stays in the “Discovered – currently not indexed” state.
A company website with two hundred pages, whose new articles get indexed within hours, does not have a crawl budget problem, whatever the health score in an audit tool says. If content is not getting indexed there, the cause is almost always quality, internal linking or duplication, and “crawl budget optimization” is a way to spend a quarter fixing something that does not hurt you.
The complication is what counts as the number of pages. The thresholds refer to unique URLs Google can reach, not products in your database. An online store with 30,000 products and unmanaged faceted navigation can expose millions of crawlable URLs. Whether crawl budget concerns you depends on how many URLs your site actually exposes, not how many products it has.
Two mistakes happen at this stage. Do not optimize crawl budget on a site where new pages get indexed the same day. You will spend the effort and the business will not notice any difference. And do not read the “Discovered – currently not indexed” state purely as a crawling problem. On smaller sites it is far more often a verdict on quality or internal linking than on capacity, even though Google lists a large share of URLs in this state as one of the three profiles this work is meant for.
Why Search Console alone will not show you the waste
I start with the Crawl stats report in Search Console. It shows total requests, the response time trend and a breakdown by status code, file type and Googlebot type. It tells you that something is wrong, whether that is a spike in requests, a growing share of 404s or a rising average response time.
What it will not reliably tell you is where crawling goes at the level where you make decisions, URL pattern by URL pattern. The report shows examples, not the full stream of requests. For a site large enough for crawl budget to matter, the only complete record of Googlebot’s behavior is your server access log. Everything else is a sample or an estimate.
This is a mistake I see over and over. The analysis is often run on your own crawler’s output instead of the logs. A crawl shows what Googlebot could fetch by following your links. Logs show what it actually fetched, how often and with what response. The findings I care about usually sit in the gap between the two lists, the URLs Google keeps requesting that your own crawl never discovers at all. Typically these are leftovers from old structures, feeds and long-canceled campaigns.
What log analysis looks like
The process I use on technical SEO projects is the same whether you do it in BigQuery, in Python or in a log analysis tool, and it has four steps.
- Isolate verified Googlebot traffic. Filter by user agent, then verify against Google’s published crawler IP ranges. Anyone can send a fake Googlebot user agent, and on some sites scrapers posing as Googlebot make up a significant part of “bot” traffic.
- Assign every request to its URL group. Products, categories, filter parameters, internal search, pagination, static files, feeds, everything else. This classification is the core of the whole analysis, because a log with a pattern column answers questions that a raw log cannot.
- Aggregate by pattern. Request count, share of the total, status code mix and trend over time. The sentence “Googlebot made 400,000 requests last month” turns into a sentence like “38% of fetches went to sort parameters that are canonicalized anyway”.
- Join it with value data. Match the patterns with Search Console impressions and, where you have them, with revenue per landing page. The result is two columns side by side, what Google spends crawling on versus what makes money.
From my own practice. I do this join in BigQuery, logs on one side, the Search Console bulk export to BigQuery on the other, both joined on URL pattern. The reason is not sophistication but repeatability. A one-off log analysis tells you where the waste was last month. The same query run every month tells you whether the fixes actually changed how Googlebot crawls, and that is the question the client pays for.

Where crawl is wasted most often
Server log data at pattern level usually brings the same categories of waste to the surface. I list them roughly by how often they appear.
- Filter and parameter URLs. Sorting, view switches, session and tracking parameters, and filter combinations nobody searches for. On online stores this is usually the biggest item.
- Internal search results. Crawlable
?q=URLs, sometimes created by spam queries linked from outside the site. - Redirect chains and loops. Every hop is a separate request, so a chain that passes through three URLs before reaching the final one costs four fetches instead of one. Googlebot can follow a chain of up to ten hops, so long chains usually get followed to the end. It just costs you something every time. Google itself recommends redirecting straight to the final destination and, if that is not possible, keeping the chain ideally to three hops and definitely under five. Most chains appear after migrations, when a new redirect map is stacked on top of the previous one, and they are cheapest to find during a redirect check right after launch.
- Soft 404s. Pages that return 200 while telling the user nothing was found. Google says explicitly that they keep getting crawled and really do waste crawl budget. Note the difference from real 404s, because in its overview of crawling myths and facts Google states that pages returning 4xx codes (except 429) do not waste crawl budget, because the crawler got a status code and no content. If your log analysis counts ordinary 404s as waste, it is counting the wrong thing.
- Host and protocol duplicates. HTTP next to HTTPS, www next to non-www, variants with and without a trailing slash, staging subdomains nobody ever closed.
- Calendar and pagination traps. Templates that endlessly generate a “next” link and send the crawler on forever.

How to fix it: each pattern has its own remedy
The remedies are nothing groundbreaking. The discipline lies in applying one remedy per pattern and knowing what each tool actually does. Robots.txt prevents fetching, which makes it the right tool for pure waste such as internal search or endless calendars, but it does not remove URLs that are already indexed, and Googlebot will not see a noindex rule on a page blocked in robots.txt.
Noindex removes pages from the index, but it uses crawling to do so. Canonical tags consolidate near-duplicates. And what has worked best on the sites I have worked on is to stop generating internal links to URLs you never wanted crawled. A pattern that nothing links to stops being a problem on its own.
Two traps come up at this stage. Do not block patterns in robots.txt as the first step for URLs that are already indexed. Googlebot will not see the noindex on them, and the URLs may stay in the index. And do not expect sitemap compression to save much crawling, because compressed sitemaps still have to be downloaded from the server.
One limitation is worth knowing. Google calculates the crawl capacity limit per host, and it is shared across all of its crawlers. If a site is hitting that shared capacity limit, high demand from one of them, for example AdsBot on dynamic ad targets or Google Shopping on merchant feeds, can reduce the capacity available to Googlebot. On top of that, every site starts at the same conservative default and only gets more capacity if it handles the load without errors or slowdowns.
If you want to go deeper, Google has published a whole Crawling December series on crawling. It is the most complete official material on the topic and worth reading in full, not just in parts.
What to avoid
- Do not work on crawl budget on a small site. A site with a few hundred pages gets a “crawl budget issue” from a tool, and the team spends three months tuning robots.txt. Meanwhile, new articles get indexed within hours. Start with the test: “does it get indexed the same day?”.
- Do not analyze logs without verifying Googlebot. The log analysis runs on all requests with a Googlebot user agent, and half of them are a scraper. Without verifying the IP ranges, you end up analyzing someone else’s traffic.
- Do not redirect 404s to a category. A report counts ordinary 404s as waste, and the team “fixes” thousands of 404s by redirecting them to a category. The result is soft 404s and even more waste. Leave 404s alone, deal with soft 404s.
- Do not block in robots.txt what first needs to drop out of the index. Indexed filters get blocked in robots.txt “to save crawl” and may stay in the index, because Google will not see the noindex. The order is noindex, wait, block.
Questions I get before an audit
Will Google crawl my good pages more if I cut the waste?
Not by itself. Google says so directly. Google does not shift newly freed crawl budget to other pages unless it is already hitting your site’s crawl capacity limit. Crawl demand matters as much as capacity. Even if the capacity limit is not reached, low demand means Google crawls your site less.
Does better crawling improve rankings?
No, and Google lists this among the crawling myths. Crawling is necessary for a page to appear in search results, but it is not a ranking signal. What this work changes is how quickly new and updated pages show up, and that is exactly where the business value lies for large, fast-changing catalogs. If someone sells you crawl budget optimization as a lever for rankings, that is a warning sign.
Do my subdomains share one crawl budget?
No. Here Google treats a site as a unique hostname, so, for example, www.example.com and shop.example.com have separate crawl budgets. That is why the answer has to come from data, not from an overall score a tool assigns to the whole site.
Three numbers that show it worked
- Crawl share by URL pattern, month by month. The waste groups should shrink, and products and categories should grow as a share of total fetches.
- Time from publication to Googlebot’s first visit for new products and articles, a metric that even management understands without explanation.
- Status code mix in Googlebot traffic. The share of 200s on indexable URLs should rise, and the number of unnecessary 404s and of redirects in chains should fall as redirect chains get shorter.
If these three metrics move in the right direction, the crawl work is done and the discussion can go back where it belongs, to content and links.
And if the first test showed you that crawl budget does not concern you, that is the cheapest finding in this whole article. It saved you a lot of work on a problem you do not have.
Not sure where your crawl budget is really going?
Send me your site and access to the logs. I will tell you whether crawl budget is your problem at all, and if it is not, I will tell you straight. If it is, the next step is a website SEO audit.