Search Console shows that Google selected a different canonical URL than the one your page declares in `<link rel="canonical">`. What causes a declared canonical to be overridden, and how would you investigate?
answer
- a hint competing with other evidence
- links and sitemap vote too
- target must be a clean 200
- raw HTML, not the rendered DOM
- declared versus selected, side by side
basics
~20 sA canonical is a hint weighed against other signals, so it loses when internal links, sitemap entries, redirects or hreflang point elsewhere, when the target is not indexable, or when the pages are not actually duplicates. Investigate by reading the raw HTML and comparing every signal.
solid answer
~60 sA canonical is one signal, not a command, and an engine overrides it when the rest of your site disagrees. The usual causes: the pages are not genuinely equivalent, so the "duplicate" claim is rejected; internal links and the sitemap point at a different URL than the canonical does; the declared target redirects, 404s, is `noindex`, or itself canonicalises somewhere else, forming a chain; several conflicting canonical tags are injected by different parts of the template; or the tag only appears after client-side rendering. To investigate, start with the raw response — fetch the page with `curl` and confirm exactly one canonical ships in the initial HTML with an absolute URL. Then check the target's status code and its own canonical. Then compare signals: which URL do internal links use, which is in the sitemap, do http/https and www variants redirect consistently. Search Console's URL Inspection tool shows the user-declared and Google-selected canonical side by side, which tells you which URL won and gives you the thing to reconcile.
code
bash · 8 lines# 1. Exactly one canonical, in the initial HTML, absolute?
curl -s https://shop.example/shoes | grep -io '<link[^>]*canonical[^>]*>'
# 2. Does the declared target answer 200 without a redirect chain?
curl -sIL https://shop.example/shoes | grep -E '^HTTP/'
# 3. Does the target canonicalise somewhere else again?
curl -s https://shop.example/shoes | grep -io 'rel="canonical"'go deeper
Know the headline fact: a canonical is a hint, so a search engine can pick a different URL. Being able to say that other signals such as internal links and redirects also count is enough here.
Name the concrete causes — a non-indexable or redirecting target, conflicting duplicate tags, a script-injected tag, pages that are not really duplicates — and be able to check the raw HTML for exactly one absolute canonical.
Walk a full diagnosis: raw response first, then the target's status and its own canonical, then the internal link graph, sitemap and host/scheme redirects, then URL Inspection to see declared versus selected. Say explicitly why blocking the crawl is the wrong reaction.
Frame it as signal consistency across the whole site: one host form, canonicals generated centrally from the route, sitemaps built from the same source of truth, and automated checks that fail a deploy when a page type emits zero or multiple canonicals.
## Why an override is possible at all `rel="canonical"` is defined as a preference. Search engines accept that site owners misconfigure it constantly — templates that canonicalise everything to the home page, plugins that fight each other — so they treat it as strong evidence rather than an instruction, and they combine it with everything else they know about the URL. When the tag and the rest of the evidence disagree, the tag can lose. Seeing a different "Google-selected canonical" in the URL Inspection tool is therefore a report that your signals are inconsistent, not a bug in the tag. ## The causes, roughly in order of frequency **The pages are not actually duplicates.** Canonical means "this is the same content as that". If a filtered listing, a size variant, or a paginated page carries materially different content from the URL it names, the claim is rejected and each page is judged on its own. The fix is usually to stop making the claim. **Internal links contradict it.** If every navigation link, breadcrumb and related-product link on the site points at `/shoes?ref=nav` while the canonical says `/shoes`, the link graph — which is a much larger body of evidence — is arguing for the other URL. **The sitemap lists the non-canonical URL.** A sitemap is a statement of the URLs you want indexed. Listing variants that canonicalise elsewhere is a direct contradiction. **The target is not a clean 200.** A canonical pointing at a URL that redirects, 404s, or serves `noindex` is unusable. So is a chain: A canonicalises to B, B canonicalises to C. Point directly at the final URL. **Duplicate or conflicting tags.** A base template emits one canonical, a CMS plugin emits another, and a page-level override emits a third. Multiple conflicting `<link rel="canonical">` elements in one head are typically all discarded. **The tag is client-rendered.** If the canonical is written into the head by JavaScript after load, the initial HTML has none. It may be picked up when the page is rendered, but you have made the signal contingent on rendering succeeding. **Host and scheme inconsistency.** Canonicals to `https://www.example.com/...` while redirects, links and the sitemap use the bare host. Pick one and make everything agree. **hreflang disagreement.** Localised pages that canonicalise across locales rather than to themselves collapse the alternate set and hand the engine contradictory instructions. ## The investigation, in order **1. Read the raw response, not the rendered DOM.** ```bash curl -s https://shop.example/shoes | grep -i 'rel="canonical"' ``` You are checking three things: that a canonical is present at all, that there is exactly one, and that its value is the absolute URL you expect. Doing this with `curl` rather than DevTools distinguishes "in the HTML" from "added by script". **2. Follow the target.** Fetch the declared canonical URL and check its status code, whether it redirects, and what its own canonical says. A chain or a redirect here explains most overrides on its own. **3. Compare the two candidates' content.** If the URL the engine chose serves noticeably different content from the one you declared, the duplicate claim was never valid. **4. Audit the supporting signals.** Which URL form do internal links use? Which is in the sitemap? Do `http://`, `www.` and trailing-slash variants all redirect to the same final URL in one hop? Every one of these that disagrees is a vote against your tag. **5. Use URL Inspection.** For a verified property, it reports the user-declared canonical and the Google-selected canonical for a URL. That is the authoritative statement of what happened; everything above is how you make the two agree. ## What not to do Do not respond by adding a robots.txt `Disallow` on the losing URL — blocking the crawl means the canonical on it can never be read either, and the URL can still be indexed bare. Do not stack `noindex` on top of a canonical hoping one will stick; they ask for different outcomes and the combination is a contradiction, not a belt-and-braces. And do not treat one override as an emergency: the correct move is to make the whole signal set consistent and re-check after the pages are recrawled.
- Why does it matter whether the canonical is in the initial HTML rather than added by JavaScript?Because the signal then depends on the page being rendered before it exists. The initial response is what every crawler sees immediately and identically; anything script-injected is best-effort. If a template can emit the tag server-side, it should, and a canonical added late is a prime suspect when the declared one is being overridden.
- Your canonical points at a URL that permanently redirects to a third URL. What is the effect?You have declared a preference for a URL that itself says "go elsewhere", so the engine follows the redirect and is likely to select the final destination — or to distrust the chain altogether. Always point the canonical at the final, non-redirecting URL rather than at an intermediate hop.
- Would adding a `Disallow` for the losing URL fix the override?No, it makes it worse. Blocking the crawl means the canonical on that URL can never be read again, so the consolidation you wanted stops happening, and the blocked URL can still be indexed from links with no snippet. Reconcile the signals instead of hiding one of them.
saying these in an interview costs you the question
- Treats the canonical as a directive engines must obey
- Blocks the losing URL in robots.txt as the fix
- Checks the rendered DOM instead of the raw HTML
- Ignores that internal links point at a different URL
- Adds a second canonical tag to reinforce the first