skip to content

Canonical and Robots Directives

How a page declares which URL is authoritative and whether crawlers may index or follow it. The distinction interviewers probe is that robots.txt controls crawling while a noindex directive controls indexing — blocking the crawl can prevent the noindex from ever being seen.

part ofHTMLoverview, primer and where to startread it →
on this pageshow

questions

6

A URL is listed under `Disallow` in robots.txt and the page also serves `<meta name="robots" content="noindex">`. Why can that URL still turn up in search results?

level: middleimportance: must knowfreq 58%

basics

~20 s

robots.txt controls crawling, not indexing. A crawler that obeys the Disallow never fetches the page, so it never reads the noindex directive inside it — while the URL itself, discovered from links elsewhere, can still be indexed without a description.

open as a page

Your site must keep a set of PDF files out of search results. Why does `<meta name="robots" content="noindex">` not help there, and what do you use instead?

level: middleimportance: should knowfreq 36%

basics

~20 s

A robots meta tag only exists inside an HTML head, and a PDF has none, so there is nothing for a crawler to parse. Send the same directive as an X-Robots-Tag response header for those paths instead; it works for any file type.

open as a page

Search Console shows that Google selected a different canonical URL than the one your page declares in `<link rel="canonical">`. What causes a declared canonical to be overridden, and how would you investigate?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A canonical is a hint weighed against other signals, so it loses when internal links, sitemap entries, redirects or hreflang point elsewhere, when the target is not indexable, or when the pages are not actually duplicates. Investigate by reading the raw HTML and comparing every signal.

open as a page

A large catalogue site generates many URL variants — filter combinations, sort orders, pagination and tracking parameters. How do you decide, per surface, between disallowing the crawl, serving a noindex directive, canonicalising to a base URL, or leaving it indexable?

level: principalimportance: should knowfreq 28%

basics

~20 s

Match the tool to the intent: canonical for true duplicates that must stay crawlable, noindex for pages that should exist but never be listed, robots.txt only for near-infinite spaces where saving crawl effort outweighs losing visibility, and indexable for variants with genuine standalone demand.

open as a page

A shop serves the same product page at /en-gb/kettle and /de-de/wasserkocher. How should `<link rel="alternate" hreflang="…">` be set up across those pages, and how does it interact with each page's canonical?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

Every page in the language set lists the full set of alternates, including a link to itself, using matching language-region codes plus an optional x-default. Each page keeps a canonical pointing at itself — canonicalising one localisation to another collapses the set.

open as a page