Indexing

Indexing is Google storing your page so it can be returned in results. It is not automatic, it is not guaranteed, and Google’s own documentation says so in plain words.

What happens during indexing

After fetching and rendering a page, Google analyses what’s on it: text, images, video, the language it’s written in, the region it targets, whether it works on a phone. It groups near-identical pages and picks one, the canonical, to represent the group.

Then it decides whether to keep the page at all. Google crawls considerably more than it stores, and the bar has visibly risen over the last few years. Pages repeating what’s already indexed elsewhere are the usual casualties.

Reading the Search Console statuses

StatusWhat it meansUsual cause
Crawled, currently not indexedFetched, examined, not storedAdds nothing new; weak internal links; site quality bar
Discovered, currently not crawledGoogle knows the URL, hasn’t fetched itCrawl economics; slow server; too many low-value URLs
Duplicate, Google chose different canonicalYour canonical was overruledContradictory signals between links, sitemap and tags
Duplicate without user-selected canonicalDuplicates found, no canonical declaredMissing canonical tags
Excluded by noindex tagYou told it not toIntentional, or a template mistake
Soft 404Returns 200 but looks emptyEmpty category pages, out-of-stock products
Alternate page with proper canonical tagWorking as intendedNo action needed

Read these by trend across months rather than as a snapshot. A single day’s numbers tell you very little; a bucket growing steadily over a quarter tells you something is systematically wrong.

Why pages get left out

In audits, the causes cluster into 4:

  1. The page adds nothing. Its content substantially exists elsewhere, on your site or better ones. Thin tag archives, near-duplicate product variants and boilerplate location pages live here.
  2. The site’s overall quality sets a bar this page misses. Assessment appears to operate sitewide, so weak sections drag on strong ones. Google’s own description of the helpful content classifier as site-wide supports this reading.
  3. Internal signals are too weak. One distant internal link says the page is expendable.
  4. It’s simply new. New pages on new sites sometimes wait. This clears on its own, which is why trend beats snapshot.

What actually helps

  • Publish fewer, better pages. If thin pages are dragging the bar down, adding more thin pages makes it worse. Consolidation beats production.
  • Strengthen internal linking to pages you want indexed. It’s the cheapest signal you control.
  • Noindex the tail deliberately. Deciding what should not be indexed is a real strategy, not an admission of failure.
  • Fix soft 404s. Empty states returning 200 waste crawl and index budget on nothing.

What doesn’t help

Mass-requesting indexing in Search Console does not override a quality judgement. Nor do indexing services, which mostly do the same request at scale. If Google has looked at a page and declined it, the answer is to change the page rather than ask louder.

Related

Crawling and indexing is the full guide. Crawlability covers the gate before this one, and canonical tags the signal most likely to confuse it.