Crawl budget is real, documented and almost certainly irrelevant to your site. Google’s own guidance aims it at sites with over a million pages, or hundreds of thousands of pages that change daily. I’m writing this guide as much to tell you when to close the tab as to explain the mechanics.
The 2 dials: capacity and demand
Google decides how much to crawl your site from 2 inputs. Crawl capacity is what your server can take: fast, healthy responses raise it, errors and slowdowns cut it. Crawl demand is how much Google wants your URLs: popularity, how often your content genuinely changes, and how stale its copy of you has become. Budget is the product of the two. You influence capacity with infrastructure and demand with being worth revisiting; direct control doesn’t exist.
Who actually has a crawl budget problem
- Ecommerce with faceted navigation minting URL combinations by the million.
- Classifieds, job boards and listings sites with heavy churn and expiring pages.
- Publishers with decades of archives plus a news velocity that needs fast pickup.
- Sites where the page indexing report shows ‘Discovered, currently not crawled’ stacking up in the tens of thousands.
A 300-page business site does not have a crawl budget problem. It might have a crawling problem (server errors, blocked resources) or an indexing problem (quality), and those have their own guide. Renaming them ‘crawl budget’ just imports fixes designed for million-page sites.
Reading the crawl stats report
Search Console → Settings → Crawl stats. Look at total requests against your indexable URL count, the response-code split (a rising 4xx/5xx share eats capacity), average response time trend, and the by-purpose split of discovery versus refresh. A healthy pattern shows most crawls refreshing known URLs, quick responses, and errors staying rare. Host status warnings are the report shouting; the trends are it talking.
Fixes that move the needle at scale
- Cut URL proliferation at source: canonicalise or block faceted combinations you’d never want indexed, per Google’s faceted navigation guidance.
- Collapse redirect chains; every hop spends a fetch.
- Return 404 or 410 quickly for gone content instead of soft-404 pages that invite re-crawling.
- Speed the server up. Capacity follows response time more directly than any other lever you hold.
- Keep sitemaps clean and dated: only canonical, indexable URLs, with honest lastmod values.
Sources
Google’s crawl budget management guide, which contains the million-page threshold, and the crawl stats report documentation. The advice to close the tab is mine and I stand by it.