Crawl budget is how much crawling Google is willing to do on your site. It is real, it is documented, and for almost every site reading this it is irrelevant.
Who it applies to
Google’s own guidance aims crawl budget management at sites with more than a million pages, or sites with hundreds of thousands of pages that change daily. Below that, Google says most sites do not need to think about it.
That threshold gets ignored constantly, because crawl budget sounds technical and important and gives an audit something to say. Working on it at 400 pages is a way of looking busy.
The two dials
| Dial | What it is | What moves it |
|---|---|---|
| Crawl capacity | What your server can handle without strain | Response time, error rates, server health |
| Crawl demand | How much Google wants your URLs | Popularity, genuine change frequency, staleness of its copy |
Budget is roughly the product of the two. You influence capacity with infrastructure and demand by being worth revisiting. Direct control does not exist, and no setting in Search Console grants it.
Who genuinely has a problem
- Ecommerce sites where faceted navigation mints URL combinations by the million.
- Classifieds, job boards and listings with heavy churn and expiring pages.
- Publishers with decades of archive plus a news velocity that needs same-day pickup.
- Any site where ‘Discovered, currently not crawled’ is stacking up in the tens of thousands in Search Console.
That last one is the real test. If that bucket isn’t growing, you don’t have a crawl budget problem regardless of your page count.
What actually helps at that scale
- Cut URL proliferation at source. Canonicalise or block filter combinations you’d never want indexed. This is the biggest lever by a distance.
- Make the server faster. Capacity follows response time more directly than any other input you control.
- Collapse redirect chains. Every hop spends a fetch that could have been a real page.
- Return 404 or 410 promptly for gone content, rather than soft 404s that invite re-crawling forever.
- Keep sitemaps honest. Canonical, indexable URLs only, with truthful lastmod dates. Lying in lastmod trains Google to ignore it.
- Fix the error rate. A rising share of 5xx directly reduces capacity.
How to check whether you have one
Search Console, Settings, then Crawl stats. Look at total requests against your indexable URL count, the response code split, average response time, and the by-purpose split between discovery and refresh. A healthy pattern is most crawls refreshing known URLs, fast responses, and errors rare.
Common mistakes
- Working on it at all on a small site. If pages aren’t crawled there, the cause is almost always crawlability or quality.
- Blocking crawling to save budget on pages you want indexed. Self-defeating.
- Assuming crawl frequency is a quality score. It reflects change and demand, not approval.
- Submitting enormous sitemaps of low-value URLs and wondering why the good pages get crawled less.
Related
The full guide is when crawl budget matters and when it’s a distraction. See also crawlability and indexing.