A canonical tag tells Google which URL you consider the master version of a page. It’s a hint rather than an instruction, and Google overrules it more often than most people realise.
The problem it solves
Duplicate and near-duplicate URLs are normal on real sites, not a sign of incompetence. Tracking parameters, filter combinations, print versions, session IDs, and the same product reachable through 3 category paths all produce multiple URLs for one piece of content.
Left alone, that splits signals across several URLs. Links point at different versions, Google picks whichever it judges best, and you get a weaker page than the sum of its parts. The canonical consolidates them: you nominate one URL, and the ranking signals should accrue there.
How to implement it
<link rel="canonical" href="https://example.com/boots/" />Rules that prevent most of the problems:
- Every page gets one, including the canonical version itself, pointing at its own URL. Self-referencing canonicals are correct and expected.
- Use absolute URLs, with the right protocol and consistent trailing slashes. Relative canonicals are valid and invite mistakes.
- One per page. Two canonical tags means Google ignores both.
- Point at a 200. A canonical to a redirecting, 404ing or noindexed URL is a contradiction Google has to resolve for you.
- Keep it consistent with everything else. Internal links, sitemap entries and hreflang should all name the same canonical URL.
Why Google overrules you
Google treats the canonical as one signal among several. It also weighs internal links, sitemap inclusion, redirects, hreflang and which version it judges better for users. When those disagree with your declared canonical, Google picks its own.
To see what it chose, use URL Inspection in Search Console. It reports both the user-declared canonical and the Google-selected canonical. When those differ, you have a signals problem somewhere else, and changing the canonical tag again will not fix it.
Common mistakes
- Canonicalising everything to the homepage. A surprisingly common bug that tells Google your entire site is one page. Usually a misconfigured plugin or template.
- Canonicalising paginated pages to page 1. Page 2 is not a duplicate of page 1, and doing this hides its content from the index.
- Combining canonical with noindex on the same URL. You’re asking Google to consolidate signals into a page and drop it simultaneously. Pick one.
- Canonical plus robots.txt block. Blocked pages can’t be crawled, so the canonical is never read.
- Treating it as a directive. It’s a strong hint. Contradict it elsewhere and it loses.
- Cross-domain canonicals set carelessly. Valid for syndication, and a good way to hand your rankings to someone else if applied wrongly.
Related
Crawling and indexing sets out how canonical sits alongside robots.txt and noindex, and which job each one does. See also indexing and keyword cannibalisation.