A canonical tag is a line in a page’s HTML head, written as <link rel="canonical" href="...">, that names the preferred URL for that content. When the same or very similar content is reachable at more than one address, the canonical tells search engines which one to index and which one should receive the ranking signals the others would otherwise split between them.
The crucial thing to understand is that a canonical is a hint, not an instruction. Google takes it as strong evidence and usually honors it, but it weighs other signals too: internal links, the sitemap, redirects, and which version looks more like the real page. If those contradict your canonical, Google may pick a different URL, and it will tell you so in Search Console.
What problem it solves
Duplicate content is rarely deliberate. It arises from the ordinary machinery of websites: tracking parameters appended to URLs, filter and sort options, session identifiers, print versions, products reachable through several categories, and trailing-slash or case variations.
Each variation is a separate URL to a search engine. Left unmanaged, they compete with each other, dilute link signals between them, and waste crawl effort on near-identical pages. The canonical consolidates that back onto one address.
Note that duplicate content is generally not a penalty. It is a dilution and efficiency problem, which is why the fix is consolidation rather than deletion.
Self-referencing canonicals
Every indexable page should carry a canonical pointing at itself. This is not redundant. It sets the definitive form of the URL, so that arrivals with tracking parameters, alternate casing, or other variations all resolve to the version you intend. Most SEO plugins add self-referencing canonicals automatically, which is one of the more valuable things they do without being asked.
Canonical, redirect, or noindex?
Three tools address overlapping problems and choosing wrongly is the usual source of trouble.
Use a 301 redirect when the old URL should no longer exist. Nobody needs to reach it. This is the strongest and cleanest signal, and it should be your first choice whenever the duplicate serves no purpose.
Use a canonical when both URLs need to stay reachable but only one should be indexed. Filtered product listings, tracked campaign URLs, and content syndicated to a partner all fit here: the variant must work for the visitor, but should not compete in search.
Use noindex when a page should be accessible but has no place in search at all, and is not a duplicate of anything. Thank-you pages, internal search results, thin archives. See indexability.
A rule of thumb: if you would be happy for the URL to stop working, redirect it. If it must keep working but should not rank, canonicalize it. If it must keep working and is not a duplicate, noindex it.
Pagination, and a mistake we made ourselves
Paginated listings deserve their own mention because the wrong approach is so widespread. Page two of an archive should carry a self-referencing canonical pointing at page two, not at page one. Canonicalizing every paginated view back to the first page tells Google those pages are duplicates of page one, which means they can be dropped from the index and the internal links they contain get discounted. On a large listing, that weakens the crawl path to everything beyond the first page.
We had exactly this on our own glossary. The listing paginated through a custom query variable that our SEO plugin did not recognize, so every paginated view emitted the same canonical as page one, and the entries beyond the first thirty were being reached through a devalued path. The fix was to rebuild the listing as a proper archive so pagination became native and self-canonicalizing. The lesson generalizes: any custom pagination or filtering mechanism your SEO plugin does not know about will quietly inherit the wrong canonical, and nothing will alert you.
Common mistakes
Canonicalizing everything to the homepage. A catastrophic and surprisingly common misconfiguration. It tells Google no other page is worth indexing.
Canonical pointing at a redirected or missing URL. The signal is contradictory, so Google ignores it and decides for itself.
Canonical pointing at a page blocked in robots.txt, which Google cannot fetch and therefore cannot confirm.
Canonical plus noindex on the same page. Contradictory instructions: one says consolidate signals here, the other says exclude this. Pick one.
Multiple canonical tags on a page, usually because a theme and a plugin both output one. Google typically ignores all of them.
Relative URLs. Always use the absolute URL including protocol and domain.
Canonicalizing genuinely different pages together because they look similar. If the content differs meaningfully, they are not duplicates and should both be indexed.
Mismatched signals. Canonicals that say one thing while internal links and the sitemap consistently point elsewhere. Google follows the weight of evidence, not the tag alone.
Cross-domain canonicals
Canonicals can point to a different domain, which is the correct handling for syndicated content. If a partner republishes your article, their copy should canonicalize to yours so the original gets the credit. It relies on the other site cooperating, so agree it in advance rather than assuming.
Checking what Google actually chose
Because the canonical is a hint, the only way to know what Google settled on is to ask. The URL Inspection tool in Search Console reports both the user-declared canonical (what your page says) and the Google-selected canonical (what Google decided). When they differ, something else is outweighing your tag, and that is the thing worth investigating.
The page indexing report also lists pages excluded as “Alternate page with proper canonical tag,” which is normal and expected, and “Duplicate, Google chose different canonical than user,” which is not.
Beyond that, view the page source rather than the rendered page to confirm exactly one canonical is present, and crawl the site to find pages where the canonical points somewhere unexpected.
Where this fits
Canonical errors are among the most common issues we find and correct, partly because they are invisible without looking and partly because plugins, themes and custom templates can each introduce one. They are a standing check in our technical SEO services, alongside crawlability and indexability. If pages are missing from search or the wrong version is ranking, book a discovery call.