skip to content

A large catalogue site generates many URL variants — filter combinations, sort orders, pagination and tracking parameters. How do you decide, per surface, between disallowing the crawl, serving a noindex directive, canonicalising to a base URL, or leaving it indexable?

level: principalimportance: should knowfreq 28%

answer

  1. same content, or merely unwanted?
  2. two axes: duplicate, and crawlable
  3. only one tool saves crawl budget
  4. pagination is not a duplicate
  5. encode it in the rendering layer

basics

~20 s

Match the tool to the intent: canonical for true duplicates that must stay crawlable, noindex for pages that should exist but never be listed, robots.txt only for near-infinite spaces where saving crawl effort outweighs losing visibility, and indexable for variants with genuine standalone demand.

solid answer

~60 s

I start from what each surface is actually for. A tracking-parameter or sort-order URL is the same content, so it self-canonicalises to the clean route and stays crawlable — the canonical is the only tool that consolidates signals rather than discarding them. A page that must exist for users but should never be a landing page, like internal search results, gets `noindex` and stays crawlable so the directive can be read. Pagination is neither: pages two onward hold different items, so they self-canonicalise and stay indexable, because canonicalising them to page one hides the products they list. robots.txt is the last resort, reserved for combinatorial spaces — three or more stacked facets — where the crawl cost is the actual problem; it saves crawl budget but blinds the engine to everything on those URLs, including any directive. Above all I try not to generate or link the bad URLs in the first place, and I encode the policy in the URL and head-rendering layer so it cannot drift per template.

go deeper

for a junior

Focus on the vocabulary first: know what a canonical, a noindex directive and a robots.txt Disallow each do, and that only one of them stops a fetch. The per-surface decision comes later.

for a middle

Be able to pick the right tool for a given example and justify it — tracking parameters canonicalise, internal search results get noindex — and to explain why a page must stay crawlable for its directive to be read.

for a senior

Show the tradeoffs explicitly: what each option costs in crawl volume and in lost link value, why pagination is not a duplicate, and how you would sequence a change so URLs are removed before any crawl block is added.

for a principal

Own it as an architectural policy, not a per-page tag: which surfaces exist at all, generating canonicals from routes, one declarative head layer, tests that fail the build on a violation, and the coverage metrics you watch to know the policy is holding.

## Frame the decision, not the tag Every one of these four options is cheap to apply and hard to undo, so the useful discipline is to ask two questions per surface before reaching for markup: 1. **Is this URL the same content as another one?** If yes, it is a canonical case. If no, canonical is off the table regardless of how much you would like the variant to disappear. 2. **Do I want the engine to see this URL at all?** If yes, it must remain crawlable, and the only lever is a directive on the page. If no — if the problem is the *volume of fetching* — robots.txt is the only tool that helps, and you accept blindness as the price. Those two axes give you the whole matrix. Everything else is applying it. ## Surface by surface **Tracking parameters.** Identical content, arrives from outside, unbounded in number. Self-canonicalise to the clean route, keep crawlable. Do not disallow: those URLs collect real inbound links, and blocking the crawl means the canonical that would have consolidated those links is never read. **Sort orders and view toggles (`?sort=price`, `?view=grid`).** Same item set, different presentation. Canonical to the base URL. If sorting changes *which* items appear, it is not a duplicate and you are in the facet case instead. **Single-facet filters with real demand (`/shoes/waterproof`).** Different content, and people search for it. Leave indexable with a self-canonical, give it a distinct title and description, and link it from navigation. This is the case people over-suppress: a filter with genuine search demand is a landing page, not noise. **Stacked facet combinations (three or more, or free-range value combinations).** The URL space is combinatorial, the content is thin and near-duplicated, and nobody searches for it. Two moves in order: stop linking them internally (crawlers mostly follow links, so an unlinked space largely does not exist), and if the crawl volume is still a problem, disallow the pattern in robots.txt. Accept that you have then given up on any directive being read on those URLs. **Pagination.** Pages two onward list *different* products. They should self-canonicalise and stay indexable so those products remain discoverable. Canonicalising them to page one is the single most common self-inflicted wound in this area. If you also offer a view-all page and it is genuinely usable, canonicalising the paginated set to it is defensible; otherwise leave the sequence alone. **Internal search results.** Must work for users, should never be a landing page, and the URL space is user-generated and unbounded. `noindex`, crawlable. Some sites additionally disallow once the pages have dropped out — sequencing matters, because a disallow first freezes whatever is already indexed. **Non-production environments.** Not a markup problem. Authentication or an IP allowlist. A `noindex` on staging is one deploy away from being copied to production, and robots.txt is a public list of paths. ## What you are trading - **Canonical** preserves the value of the variant (its inbound links flow to the base URL) but costs crawl: the engine keeps fetching the variants to re-read the tag. - **noindex** removes the page from results cleanly and, unlike a block, is verifiable — but it also costs crawl, and the value of any links to that URL is largely lost. - **robots.txt** is the only lever that reduces fetching, and it is the bluntest: no canonical, no noindex, no redirect on those URLs will ever be seen again, and bare URLs can still be listed from links. - **Indexable** is the right answer more often than defensive instincts suggest, but only where the page has content and demand of its own; otherwise you are competing with yourself. ## Making it stick A policy that lives in a wiki decays. Encode it where URLs and heads are produced: - Generate the canonical from the **route**, never from the request URL, so parameters cannot leak into it. - Centralise the head so a page type declares its intent (indexable, noindex, canonical-to) and the rendering layer emits the tags. - Assert it in tests: for each page type, render it and check there is exactly one canonical, that it is absolute, and that the robots directive matches the declared intent. This is the check that catches a plugin adding a second tag. - Generate the sitemap from the same source of truth that decides indexability, so it can never list a URL you canonicalise away. - Watch the index-coverage reporting for the two symptoms that mean you got it wrong: URLs indexed though blocked (a disallow that should have been a noindex), and a rising count of pages excluded as duplicates with an engine-chosen canonical (signals that disagree with each other). ## The failure mode to name in an interview The classic bad answer is to reach for robots.txt for everything, because it feels like the strongest tool. It is the only one of the four that removes your ability to say anything else about those URLs. Reserve it for the case where the fetching itself is the cost you are trying to cut, and use the page-level tools everywhere the engine's understanding matters more than its request volume.

  • Why is canonicalising paginated pages to page one a mistake?
    Because pages two onward list different products. Declaring them duplicates of page one asks the engine to disregard them and everything they link to, so deep catalogue items lose their discovery path. Paginated pages should self-canonicalise and stay indexable; only a genuinely usable view-all page is a defensible canonical target.
  • When is robots.txt genuinely the right tool rather than a noindex?
    When the cost you are cutting is fetching, not listing: combinatorial facet spaces, calendar-style infinite URLs, generated exports nobody links to. You are trading away the ability to have any directive read on those URLs, so it only pays where you never want the engine to look, and where the URLs are not accumulating inbound links.
  • How would you stop this policy drifting as teams add page types?
    Make the head declarative — each page type states its indexing intent and one shared layer renders the tags — and add a test per page type asserting exactly one absolute canonical and the expected robots directive. Build the sitemap from that same source so it can never list a URL the policy excludes.
  • What signal tells you the policy has gone wrong in production?
    Two coverage symptoms. URLs reported as indexed though blocked by robots.txt mean you used a crawl block where a noindex was needed. A growing set of pages excluded as duplicates with an engine-selected canonical means your canonicals disagree with your internal links or sitemap, not that the tags are missing.

saying these in an interview costs you the question

  • Reaches for robots.txt as the default for every variant
  • Canonicalises paginated pages back to page one
  • Applies noindex to URLs already blocked from crawling
  • Suppresses filter pages that have real search demand
  • Treats robots.txt as a way to hide non-production environments

context