skip to content

In HTML, what does `<link rel="canonical">` tell a search engine, and what is a self-referencing canonical?

level: juniorimportance: must knowfreq 62%

answer

  1. one address wins among duplicates
  2. parameter variants pool their signals
  3. absolute URL, exactly one tag
  4. points at its own clean URL
  5. a strong hint, never a redirect

basics

~20 s

A canonical link names the URL you consider authoritative for content that is reachable at several addresses, so duplicates consolidate onto one page instead of competing. A self-referencing canonical points a page at its own clean URL.

solid answer

~50 s

`<link rel="canonical" href="…">` goes in the `<head>` and tells search engines which URL you consider the authoritative address for this content. It exists because the same page is usually reachable at many URLs — tracking-parameter variants, `?sort=` variants, http vs https, with and without a trailing slash — and without a canonical those variants compete with each other and split their ranking signals. A self-referencing canonical is a page pointing at its own clean URL; it costs nothing and means any variant that gets crawled or shared still carries a pointer back to the version you want indexed. Rules that make it work: use an absolute URL, emit exactly one canonical per page, and point it at a URL that returns 200 and is itself indexable. The important caveat is that a canonical is a strong hint, not a command — a search engine can choose a different canonical if your other signals disagree.

code

html · 12 lines
html
<!doctype html>
<html lang="en">
  <head>
    <meta charset="utf-8">
    <title>Running shoes</title>
    <!-- Same tag is rendered for /shoes and /shoes?utm_source=newsletter -->
    <link rel="canonical" href="https://shop.example/shoes">
  </head>
  <body>
    <h1>Running shoes</h1>
  </body>
</html>

go deeper

for a junior

Be able to write the tag from memory, say it belongs in the head, and state plainly that it names the preferred URL among duplicates. Mention that it is not a redirect and users never notice it.

for a middle

Explain how duplicate URLs arise in practice — parameters, host and scheme variants, trailing slashes — and why a self-referencing canonical built from the route rather than the request URL handles them all at once. Know the correctness rules: absolute URL, one tag, indexable 200 target.

for a senior

Show that you treat the canonical as one signal among several. Be ready to describe how internal links, sitemap entries and redirects must agree with it, and what you would check when a declared canonical is not the one the engine picked.

for a principal

Own it as a site-wide invariant rather than a per-page tag: one host form, one URL shape, canonicals generated centrally in the rendering layer, and a test that asserts the head of each page type. Be ready to argue when a duplicate should be canonicalised versus redirected outright.

## The problem canonicalization solves A single piece of content on a real site is almost never reachable at exactly one URL. A product page might answer at `https://shop.example/shoes`, at `https://shop.example/shoes/`, at `http://` instead of `https://`, at `https://www.shop.example/shoes`, and at `https://shop.example/shoes?utm_source=newsletter` after a campaign link is shared. To a search engine these are four or five distinct URLs that happen to serve the same bytes. Left alone, the engine has to guess which one to show, and whatever reputation the page earns (links pointing at it, engagement) is spread across the set rather than pooled on one address. `rel="canonical"` is the markup you use to answer the question for it. ## The markup It is a `<link>` element in the document `<head>`: ```html <head> <link rel="canonical" href="https://shop.example/shoes"> </head> ``` Points worth knowing: - **Use an absolute URL**, including scheme and host. Relative values are technically resolved against the document's base URL, but a wrong base or a page served under an unexpected path turns a relative canonical into a pointer at a URL that does not exist. - **Exactly one per page.** If a template and a CMS plugin each inject a canonical and they disagree, the usual outcome is that all of them get ignored, which is worse than having none. - **The target must be a real, indexable page** — status 200, not `noindex`, not itself redirecting somewhere else. A canonical pointing into a redirect chain or a 404 is a broken signal. - The same declaration can be sent as a response header for files that have no HTML head, such as PDFs. The `<link>` element is the form you use for HTML pages. ## Self-referencing canonicals A self-referencing canonical is simply a page whose canonical points at its own clean, parameter-free URL: ```html <!-- served at /shoes AND at /shoes?utm_source=newsletter --> <link rel="canonical" href="https://shop.example/shoes"> ``` Because the same template renders both URLs, the variant that carries the tracking parameter also ships the canonical pointing back at the clean address. That is the whole trick: you do not have to enumerate the variants, you only have to make every rendering of the page state its own preferred URL. Building the canonical from a known route rather than from the incoming request URL is what makes this work — if you echo back the current location you will happily canonicalise a page to its own tracking-parameter variant. ## It is a hint, not a directive This is the part interviewers probe. A canonical does not force anything. Search engines weigh it against other signals: which URL your internal links point at, which URL appears in your sitemap, where redirects lead, and whether the two pages really do carry the same content. If those signals contradict the tag, the engine may select a different canonical than the one you declared. A canonical also does not redirect users — the browser does nothing with it; visitors stay exactly where they are. And it does not remove anything from the index by itself. If your goal is "this URL must not appear in search results at all", that is a robots directive, not a canonical. ## Common mistakes - **Canonicalising every paginated page to page one.** Pages 2, 3 and 4 hold different products; declaring them duplicates of page one asks the engine to ignore them and the items they list. Paginated pages should self-canonicalise. - **Canonicalising a filtered or faceted view to the unfiltered page when the content genuinely differs.** Canonical means "duplicate of", not "related to". - **Combining `noindex` with a canonical to another URL.** One says "drop this page", the other says "merge it into that one"; the pair is contradictory, so pick the one that matches your intent. - **Emitting the canonical only from client-side JavaScript.** It may eventually be seen, but the tag in the initial HTML response is the version you can rely on. - **Mixing hosts inconsistently** — canonicals to `https://www.` while every internal link uses the bare host. Pick one host form and make canonicals, internal links, redirects and sitemap agree.

  • Can a page carry both a `noindex` robots directive and a canonical pointing at another URL?
    Technically yes, but the two ask for opposite things: `noindex` says drop this URL from the index, while the canonical says merge its signals into another URL. Conflicting signals get resolved unpredictably. Decide which you mean — canonical for a genuine duplicate you want consolidated, `noindex` for a page that should not be in the index at all.
  • Does a canonical have to point at a URL on the same domain?
    No. Cross-domain canonicals are supported and are the normal answer for syndicated content: the syndicating site canonicalises to the original publisher. The target still has to serve equivalent content and be reachable and indexable, otherwise the signal is ignored.
  • Where else can a canonical be declared for a file that has no HTML head, like a PDF?
    As a `Link` response header with `rel="canonical"`, which the server or CDN attaches to that path. It carries the same meaning as the `<link>` element. The single-declaration rule still applies — a header and a tag that disagree tend to cancel each other out.

saying these in an interview costs you the question

  • Says a canonical redirects visitors to the other URL
  • Claims a canonical guarantees the duplicate is removed from the index
  • Emits several canonical tags from different templates on one page
  • Canonicalises every paginated page back to page one
  • Builds the canonical from the incoming request URL including its parameters

context