skip to content

Your site must keep a set of PDF files out of search results. Why does `<meta name="robots" content="noindex">` not help there, and what do you use instead?

level: middleimportance: should knowfreq 36%

answer

  1. the file has no head to parse
  2. move the directive onto the response
  3. same vocabulary, different transport
  4. works for PDFs, images, exports
  5. still needs to be fetchable

basics

~20 s

A robots meta tag only exists inside an HTML head, and a PDF has none, so there is nothing for a crawler to parse. Send the same directive as an X-Robots-Tag response header for those paths instead; it works for any file type.

solid answer

~50 s

The robots meta tag is HTML markup — it lives in a document's `<head>`, and a PDF, an image or a CSV has no head to put it in, so there is no way for a crawler to read a directive from the file itself. The equivalent for non-HTML resources is the `X-Robots-Tag` response header, configured on the web server or CDN for the matching paths: `X-Robots-Tag: noindex`. It carries the same vocabulary as the meta tag — `noindex`, `nofollow`, `none`, `noarchive`, `nosnippet`, `noimageindex`, `max-image-preview`, `unavailable_after` — and it can be scoped to one crawler by prefixing a user-agent token, as in `X-Robots-Tag: googlebot: noindex`. The same crawl caveat applies as for the meta tag: the URL must stay fetchable for the header to be seen, so do not also disallow it in robots.txt. For HTML pages either form works; when a page carries both and they conflict, the more restrictive directive is the one applied.

code

bash · 5 lines
bash
# Confirm the directive is actually on the response for a non-HTML file
curl -sI https://example.com/downloads/datasheet.pdf | grep -i 'x-robots-tag'

# Confirm the same path is not blocked from being fetched in the first place
curl -s https://example.com/robots.txt | grep -i 'downloads'

go deeper

for a junior

Know that the robots meta tag is HTML-only and that non-HTML files need the directive delivered on the response instead. Being able to name X-Robots-Tag and say it does the same job is enough at this level.

for a middle

Explain that it is a transport difference, not a semantic one — same directive vocabulary, applied to any content type — and repeat the crawl caveat: a disallowed path means the header is never received.

for a senior

Show judgment about placement: page-level tags where the directive depends on the page's own state, an edge or server rule where it is a property of a path or environment. Be ready to say how you would verify the header ships in production and past a CDN cache.

for a principal

Decide where indexing policy lives across an estate: which rules belong in infrastructure so no template can opt out, which belong to the application, and how conflicts between the two are resolved and monitored as environments and paths multiply.

## Why the meta tag runs out of road `<meta name="robots" content="…">` is an HTML element. Its entire delivery mechanism is "a crawler parses an HTML document and finds a `<meta>` in the head". That works for pages and for nothing else. Real sites publish a great deal that is not a page: PDF datasheets, generated CSV exports, images, plain-text files, ZIP archives. Each of those is a URL a search engine can find, fetch and index — Google indexes PDF content and will show it in results — and none of them has anywhere to put a meta tag. ## X-Robots-Tag The answer is to move the directive out of the document and onto the response. `X-Robots-Tag` is a response header carrying the same directive vocabulary as the meta tag: ``` X-Robots-Tag: noindex X-Robots-Tag: noindex, nofollow X-Robots-Tag: googlebot: noindex X-Robots-Tag: unavailable_after: 2027-01-01T00:00:00Z ``` Because it rides on the HTTP response rather than in the body, it applies to any content type. You configure it where responses are produced — a rule in the web server or CDN matching, say, `/downloads/*.pdf`, or a header set by the application for a particular route. Multiple `X-Robots-Tag` headers may be sent; a crawler-specific one is scoped by prefixing the user-agent token and a colon. The directive names are the same ones you already know from markup: - `noindex` — keep this resource out of the results. - `nofollow` — do not follow links found in this resource. - `none` — shorthand for both. - `noarchive` — no cached copy link. - `nosnippet` / `max-snippet` — suppress or bound the text excerpt. - `noimageindex` / `max-image-preview` — control image indexing and preview size. - `unavailable_after` — stop showing the resource after a date. ## The crawl caveat is identical A header is only observed by a client that made the request. If robots.txt disallows `/downloads/`, the crawler never issues the request, never receives the response, and never sees the `X-Robots-Tag`. The directive is as invisible as a meta tag behind the same block. Keep the path crawlable until the resource has actually left the index. ## Which form to use for HTML For HTML pages both mechanisms work and mean the same thing. Choose based on who owns the decision: - The **meta tag** is right when the directive depends on the page's own content or state — a thin profile page, a search-results page, a page the CMS marked private. The rendering code already knows. - The **header** is right when the directive is a property of a path or an environment — an entire `/preview/` tree, everything served by a staging host, every file under `/exports/`. Setting it at the edge means no template can forget it. When a single HTML page ends up with both and they disagree, the more restrictive rule is the one that applies, so a header saying `noindex` is not overridden by a page-level tag that omits it. ## Verifying it The common failure is that the header is configured but not actually being sent — the rule did not match, a proxy stripped it, or a cached copy predates the change. Inspect the real response for the exact URL: ```bash curl -sI https://example.com/downloads/datasheet.pdf | grep -i x-robots-tag ``` The browser DevTools network panel shows the same thing for a request made in context. If a CDN sits in front, re-check after a purge — an object cached before the rule was added will keep serving without the header until it is evicted. And confirm the path is not simultaneously disallowed in robots.txt, which is the other half of the same class of bug. ## A related markup-level control For HTML, when you want to suppress only part of a page from being used as a snippet rather than exclude the page, the `data-nosnippet` attribute on a `<span>`, `<div>` or `<section>` marks that subtree as ineligible for snippet text. It is a finer instrument than a page-wide `nosnippet` and it is the only one of these controls that is expressed inside the body rather than the head or the response.

  • Can `X-Robots-Tag` be scoped to a single crawler?
    Yes — prefix the directive with a user-agent token, as in `X-Robots-Tag: googlebot: noindex`. Without a prefix the rule applies to every crawler that honours the header. You can send several headers to give different crawlers different rules on the same resource.
  • How would you check that the header is really being sent?
    Request the exact URL and read the response headers — `curl -sI <url>` or the DevTools network panel. Two things commonly go wrong: the server rule did not match the path, or a CDN is serving an object cached before the rule existed, so re-test after a purge.
  • If an HTML page has both a `noindex` header and a meta tag without `noindex`, which applies?
    The more restrictive one, so the page is treated as noindex. That is why an environment-wide header is a safer place to put a blanket rule than a template — an individual page cannot accidentally opt itself back in.

saying these in an interview costs you the question

  • Suggests putting a meta robots tag inside a PDF
  • Thinks PDFs are never indexed so need no directive
  • Adds a robots.txt Disallow so the header is never read
  • Assumes X-Robots-Tag works on HTML pages only
  • Believes a header directive is weaker than a meta tag

context