Crawling and indexing: how pages get into Google (or don’t)

Before a page can rank, it has to get into the index. More pages fail at that step than most site owners realise, and the failure is usually silent: no error, no warning, just a page that never shows up. This guide covers the whole journey, discovery through indexing, the 3 controls you have over it, and how to check each step in Search Console. It pairs with the wider view in how Google Search works.

How Google finds URLs in the first place

Google discovers URLs mainly by following links from pages it already knows. Everything else supplements that: XML sitemaps declare what you think matters, redirects hand over old equity, and Search Console’s URL inspection tool lets you request a specific page directly.

The practical consequence gets missed constantly. If a page sits 5 clicks deep with one link pointing at it, you’ve told Google it barely matters, whatever the sitemap says. Sitemaps are a hint sheet, and links are the road network. Orphan pages, ones with no internal links at all, are discovery’s first casualty; my internal linking guide covers finding and fixing them.

Crawling and rendering are 2 different events

Googlebot first fetches your raw HTML. Rendering happens separately: the page is loaded in a recent version of Chrome, scripts run, and Google sees what a user would see. For most sites the gap between the two is short. For sites that assemble their content and navigation entirely in JavaScript, the gap is a dependency you’ve chosen to take on.

My rule for clients: anything you’d be upset to lose, put in the server HTML. That covers primary content, internal links and canonical tags. JavaScript enhancement on top is fine and normal. Whole sections of a site existing only after hydration is where indexing problems breed.

How much Google crawls is a budget set by demand for your content and your server’s health. Genuine crawl budget problems belong to sites with a million-plus URLs or rapid churn; if that’s you, I’ve written up when crawl budget actually matters. Most sites can ignore it.

‘Crawled, currently not indexed’, decoded

The Search Console status that launches a thousand forum threads. It means exactly what it says: Google fetched the page, looked at it, and chose to keep it out of the index for now. The docs confirm indexing is never guaranteed, and this status is that policy made visible.

In practice, when I audit sites carrying a lot of these, the causes cluster:

  • The page adds nothing new. Its content substantially exists elsewhere, on your site or on better ones. Thin tag pages, near-duplicate product variants and boilerplate location pages live here.
  • The site’s overall quality sets a bar the page misses. Index selectivity appears to operate sitewide, so weak sections drag on strong ones. That’s my read of the evidence, and I’ve marked it as such.
  • Weak internal signals. One distant internal link tells Google the page is expendable.
  • Timing. New pages on new sites sometimes just wait. The status can clear on its own, which is why I check trend before reacting to a snapshot.

Its sibling, ‘Discovered, currently not crawled’, means Google knows the URL exists and hasn’t bothered fetching it yet. Same underlying economics, one stage earlier.

Robots.txt, noindex and canonical: 3 controls, 3 different jobs

ControlWhat it governsWhat people wrongly expect from it
robots.txtCrawling. Tells crawlers which paths they may fetch.Removal from the index. A blocked URL can still be indexed from links alone, shown with no snippet.
noindexIndexing. Tells Google to keep a crawled page out of results.Working while blocked in robots.txt. Google has to crawl the page to see the tag.
canonicalPreference. Suggests which duplicate should represent the group.Obedience. It’s a hint, and Google picks its own canonical when signals disagree.

The classic self-inflicted wound combines the first two: block a path in robots.txt and add noindex to the same pages. The block stops Google from ever seeing the noindex, and the URLs linger in the index as bare links. Pick one control per job. If a page must vanish from results, noindex it and let it be crawled.

Checking all of this in Search Console

  • Page indexing report (Indexing → Pages): your master list of what’s in, what’s out and why. Read the excluded buckets by trend, month over month, rather than reacting to any single day.
  • URL inspection: paste one URL and get its last crawl date, whether it’s indexed, and which canonical Google chose versus the one you declared. When those two canonicals differ, that’s Google telling you your signals disagree with each other.
  • Sitemaps report: confirms the sitemap parses and shows how many of its URLs made the index. A big gap between submitted and indexed is your quality-bar warning light.
  • Crawl stats (Settings → Crawl stats): response codes, host status and fetch volume over 90 days. This is where server problems masquerading as SEO problems get caught.

Cadence-wise, monthly is plenty for a small site, weekly during a migration or after a confirmed update. Obsessive daily checking mostly teaches you how noisy the reports are.

Sources

Google Search Central: How Search works, robots.txt introduction, blocking indexing with noindex and canonicalisation. Where I’ve gone beyond what those pages state, I’ve flagged it as my own reading in the text.