skip to content

A URL is listed under `Disallow` in robots.txt and the page also serves `<meta name="robots" content="noindex">`. Why can that URL still turn up in search results?

level: middleimportance: must knowfreq 58%

answer

  1. two different stages of the pipeline
  2. one blocks the fetch, one blocks the listing
  3. the directive is inside the file
  4. links alone can surface a URL
  5. result with no description

basics

~20 s

robots.txt controls crawling, not indexing. A crawler that obeys the Disallow never fetches the page, so it never reads the noindex directive inside it — while the URL itself, discovered from links elsewhere, can still be indexed without a description.

solid answer

~50 s

The two mechanisms operate at different stages. A `Disallow` line in robots.txt asks crawlers not to *fetch* the URL; `<meta name="robots" content="noindex">` asks them not to *index* the page. The noindex lives inside the HTML, so it can only be honoured by a crawler that downloaded the HTML — and the Disallow is precisely what stopped it from doing so. Meanwhile a URL that has never been fetched can still be indexed if the engine discovers it from links on other pages; it typically appears with no snippet, since the engine has no content to describe. So the combination produces the opposite of what was intended. The fix is to leave the URL crawlable and serve `noindex` until it drops out of the index; only then, if you also want to save crawl budget, add the Disallow. Note also that robots.txt has no supported `noindex` directive — Google ignores one.

code

html · 12 lines
html
<!doctype html>
<html lang="en">
  <head>
    <meta charset="utf-8">
    <title>Order confirmation</title>
    <!-- Keep this URL crawlable so the directive can actually be read -->
    <meta name="robots" content="noindex, follow">
  </head>
  <body>
    <h1>Thanks for your order</h1>
  </body>
</html>

go deeper

for a junior

Be able to state the one-line distinction: robots.txt is about fetching, the robots meta tag is about appearing in results. Know where each one lives — a file at the site root versus a tag in the page head.

for a middle

Explain the failure mechanically: the directive lives inside the document, so blocking the fetch hides it, while link discovery can still put the bare URL in the index. Give the correct ordering — crawlable plus noindex first, Disallow only later if at all.

for a senior

Show you have diagnosed this in production: recognising the indexed-though-blocked state, knowing that a Disallow also hides canonicals and redirects, and being explicit that neither mechanism is a security control.

for a principal

Own the policy rather than the tags: which surfaces are blocked for crawl-budget reasons, which are noindexed, which are behind authentication, and how that is enforced in the rendering and infrastructure layers so no team can reintroduce the contradiction by hand.

## Two different jobs Search engines do two separable things with a URL. They **crawl** it — issue an HTTP request and download the bytes — and they **index** it, deciding whether it belongs in the results and what to show for it. Every robots control belongs to exactly one of those stages, and confusing them is the single most common SEO bug in production sites. **robots.txt** is a plain-text file at the root of an origin (`https://example.com/robots.txt`, per scheme, host and port). It is a *crawl* control: ``` User-agent: * Disallow: /admin/ Allow: /admin/public-notice.html Sitemap: https://example.com/sitemap.xml ``` A compliant crawler reads this before requesting anything and simply does not fetch matching URLs. That is the entire effect. It says nothing about the index. **The robots meta tag** is an *index and serving* control that lives in the HTML head of the page it applies to: ```html <meta name="robots" content="noindex, nofollow"> ``` Common values are `noindex` (keep this page out of the results), `nofollow` (do not use the links on this page for discovery or ranking), `none` (both), plus serving controls such as `noarchive`, `nosnippet`, `max-snippet`, `max-image-preview` and `noimageindex`. Using `name="googlebot"` instead of `name="robots"` targets one crawler rather than all of them. ## Why the combination silently fails The meta tag is *content*. To obey it, a crawler has to download the document, parse the head, and find it. If robots.txt forbids the fetch, that never happens. From the engine's point of view the page has no directives at all — it has no page. But a URL does not need to be fetched to be known. If any other page anywhere links to it, the engine has the URL, and it may index that URL on the strength of the link alone. What you get is a result with the URL and perhaps some anchor text, and no description — Google labels this state in Search Console as indexed though blocked by robots.txt. The page you tried hardest to hide is now in the index in its least attractive form, and your noindex is unreachable behind the very wall you built. ## The correct sequences - **"This page must not appear in search results."** Leave it crawlable. Serve `noindex` (or the equivalent response-header directive for non-HTML files). Wait until the URL actually drops out. Only afterwards, if crawl volume is a concern, consider adding a Disallow — and understand that from that moment the engine can no longer re-confirm the noindex. - **"This section is worthless to crawl and nobody links to it."** robots.txt Disallow is the right tool. It saves crawl budget on infinite parameter spaces, internal search results, and calendar-style URL explosions. - **"This must not be visible to the public at all."** Neither tool is a security control. robots.txt is public and readable by anyone, and it advertises the paths you named. Put the resource behind authentication. ## Related traps - **`noindex` in robots.txt is not a thing.** It was never part of the standard and Google does not support it. A line like `Noindex: /private/` does nothing. - **A blocked URL also hides its canonical, its hreflang set and its redirect.** Disallowing a URL that redirects elsewhere means the redirect is never seen, so the old URL lingers. - **Blocking CSS and JavaScript** prevents the engine from rendering the page as users see it, which can degrade how the page is understood. Blocking assets wholesale is rarely worth it. - **robots.txt is per-origin.** `https://example.com/robots.txt` does not govern `https://cdn.example.com/` or the http:// variant. - **`Disallow` is voluntary.** Well-behaved crawlers obey it; scrapers and malicious bots do not, and treat the file as a map of interesting paths. ## How to verify Fetch the page yourself and confirm the directive really ships in the initial HTML response rather than being injected later by script. Check the robots.txt rules that match the path — the longest matching rule wins, and an `Allow` can carve an exception out of a broader `Disallow`. In Search Console, the URL Inspection tool reports both whether the URL is blocked and what indexing decision was made, which is the fastest way to tell "never fetched" apart from "fetched and deliberately excluded".

  • What does a robots.txt-blocked URL actually look like when it does get indexed?
    Usually a bare result: the URL, sometimes anchor text from the pages linking to it, and no snippet — because the engine has never downloaded any content to summarise. Search Console reports it as indexed though blocked by robots.txt, which is the signal to remove the Disallow and serve a real `noindex` instead.
  • How do you keep a staging environment out of search results?
    Put it behind HTTP authentication or an IP allowlist. robots.txt is public and advertises paths, and a `noindex` meta tag only works for crawlers that choose to obey it. Authentication is the only mechanism that also stops people and scrapers, and it removes the whole class of problem rather than politely asking about it.
  • What is the difference between `nofollow` in the robots meta tag and `rel="nofollow"` on an individual link?
    The meta tag applies to every link on the page; `rel="nofollow"` applies to the one anchor it sits on and is the right granularity for user-generated links, alongside `rel="ugc"` and `rel="sponsored"`. Neither prevents the target from being discovered by other routes — they are signals about link endorsement and discovery, not access control.

Disallow is a locked door; noindex is a note pinned to the wall inside the room. Lock the door and nobody ever gets in to read the note — but they can still see the door and write it down.

saying these in an interview costs you the question

  • Thinks robots.txt Disallow removes a page from the index
  • Adds Disallow and noindex together for faster removal
  • Believes robots.txt supports a Noindex directive
  • Treats robots.txt as a way to hide private URLs
  • Says a blocked URL can never appear in search results

context