Crawlability is whether a search engine can find and fetch your pages. It’s the first gate in the pipeline, and a page that fails it is invisible no matter how good it is.
It’s routinely confused with indexability, which is the gate immediately after. Keeping them separate is the difference between fixing a problem and fixing the wrong problem.
Crawlability and indexability are not the same thing
| Crawlability | Indexability | |
|---|---|---|
| Question it answers | Can Googlebot fetch this page? | Will Google store it in the index? |
| Controlled mainly by | Internal links, robots.txt, server responses, rendering | Content quality, duplication, noindex, canonicals |
| Typical failure | Page never fetched | Page fetched, then discarded |
| Search Console status | Discovered, currently not crawled | Crawled, currently not indexed |
| Fix | Technical: links, config, server | Editorial: make the page worth keeping |
A crawlable page can still be left out of the index. A page blocked from crawling can still appear in results as a bare URL with no snippet, because Google learned it exists from links elsewhere and was never permitted to fetch it and see your noindex.
What decides crawlability
Internal links, first and foremost
Google finds most URLs by following links from pages it already knows. This makes your own linking the primary discovery mechanism, ahead of sitemaps and everything else.
The consequence gets missed constantly: a page sitting 5 clicks deep with a single link pointing at it has been told, by your own architecture, that it barely matters. Orphan pages, with no internal links at all, are discovery’s first casualty.
Robots.txt
Tells crawlers which paths they may fetch. It is a crawling control and nothing else. It does not remove pages from the index and it never has.
Server responses
Timeouts, 5xx errors and slow responses reduce crawling quickly, because Google backs off rather than risk overloading a struggling server. A site that gets slower gets crawled less, which becomes a compounding problem.
Rendering
Google fetches the HTML, then renders the page in a recent version of Chrome and runs its JavaScript. For most sites the gap is short. For sites where navigation and content only exist after scripts run, it’s a dependency worth understanding. My rule with clients: anything you’d be upset to lose belongs in the server HTML, which means primary content, internal links and canonical tags.
A worked example
A retailer with 40,000 products was convinced they had a content problem: decent descriptions, almost no organic traffic. Their category navigation was rendered entirely client-side, so the only internal links Googlebot could reliably follow were the 12 in the footer.
Around 90% of the catalogue had never been crawled. No amount of content work would have touched it. The fix was a server-rendered category navigation, and the crawl followed within weeks.
How to diagnose it
- Page indexing report in Search Console. Read the excluded buckets by trend across months rather than reacting to a single day. Rising ‘Discovered, currently not crawled’ is the crawlability warning light.
- Crawl stats (Settings, then Crawl stats). Look at response codes, average response time and host status. A rising share of 4xx and 5xx eats capacity that would otherwise fetch real pages.
- Crawl your own site with any crawler and compare the result to your sitemap. Anything in the sitemap that a link-following crawl never reached is effectively an orphan.
- URL Inspection for specific pages, which shows last crawl date and lets you see the rendered HTML Google actually got.
Common mistakes
- Blocking CSS and JavaScript in robots.txt. Google renders pages and needs those files to see what a user sees. This was common advice a decade ago and is actively harmful now.
- Using robots.txt to remove a page from results. Wrong tool for the job. Blocking prevents the crawl, so Google never sees the noindex you added.
- Unchecked faceted navigation generating millions of filter combinations that swallow crawling on large sites.
- Orphan pages, particularly on ecommerce sites where discontinued products fall out of navigation but stay live.
- Assuming a sitemap solves discovery. Sitemaps are a hint sheet. Links are the road network.
- Calling it a crawl budget problem on a 300-page site. See crawl budget for who genuinely has one.
Related
Crawling and indexing is the full guide, including the 3 controls and which job each does. See also indexing, crawl budget and internal linking.