skip to content

Your exported site answers every unknown URL with the app's not-found screen and a 200 status - what breaks, and how do you fix it?

level: seniorimportance: should knowfreq 44%

answer

  1. body says missing, status says fine
  2. a rewrite carries the served file's status
  3. status is chosen before client routing
  4. monitors and crawlers read the status
  5. a missing asset returns HTML instead

basics

~20 s

That is a soft 404: the page says missing while the status says fine. Crawlers, link checkers, monitors and caches all read it as healthy, and a missing asset returns HTML. Serve the not-found document with a real 404.

solid answer

~50 s

The usual cause is a catch-all rewrite that sends every unmatched path to the shell document so client-side routes survive a reload - and a rewrite answers with the rewritten file's success status. The result is a **soft 404**. Machines that read the status rather than the words - search crawlers, link checkers, uptime monitors, caches - conclude the URL exists, so dead links are never reported and junk URLs stay indexed. Worse, a missing hashed asset also returns HTML, and the browser reports a parse error instead of a 404. The fixes are: use the host's not-found document setting, which serves a chosen document with a real `404`; **scope** the rewrite to the path prefixes that genuinely have client-only routes instead of `/*`; keep asset directories out of it; and prerender a document per known route so unknown paths genuinely miss.

code

http · 12 lines
http
GET /produtcs/42 HTTP/1.1
Host: example.com

HTTP/1.1 200 OK
Content-Type: text/html

<!doctype html>...We could not find that page...

--- what it should have been ---

HTTP/1.1 404 Not Found
Content-Type: text/html

go deeper

for a junior

Learn that the status line and the page's wording are separate signals, and that a page reading "not found" can still be delivered as a success.

for a middle

Explain why a rewrite carries the served file's status and why a status must be chosen before any client-side route matching can happen.

for a senior

Show the diagnosis and the ordered fix: prerender what you can enumerate, scope the fallback, exclude asset paths, use the host's real not-found setting, and assert it in CI.

for a principal

Name the tension - surviving deep links versus reporting misses honestly - and decide how large the ambiguous URL space is allowed to be, since it is the part monitoring cannot see.

## What a soft 404 is A **soft 404** is a response whose body says the resource is missing and whose status line says the request succeeded. Humans read the body and are satisfied. Every automated consumer reads the status line, and all of them are now wrong about the site. The status code is not decoration: it is the only part of the response that non-human consumers agree on. When it lies, a whole layer of tooling silently stops working. ## Why a files-only export produces them so easily Three paths lead here, and they compound: 1. **The catch-all rewrite.** Routes resolved only by the client router have no file behind them, so a reload or a pasted link would miss. The standard remedy is a rule rewriting unmatched paths to the shell document - and a *rewrite* answers with the status of the file it served, which is normally success. The rule cannot distinguish a valid client route from a typo, because matching happens later, in the browser. 2. **Status is chosen before routing.** On a server-backed deployment the app matches the route and then picks the status. In a files-only export the status is fixed when the host picks the file, which is strictly before any client-side matching. The client can render a not-found screen; it cannot retroactively change a status line that was sent with the first byte. 3. **The error document configured as a normal document.** Some hosts let you nominate a not-found document *and* a status; others serve a fallback with 200 unless told otherwise. It is easy to configure the document and never notice the status. ## What actually breaks - **Search.** Crawlers keep dead URLs in the index and keep re-crawling them, and typo'd or generated URLs can accumulate as apparently valid pages, all with near-identical content. - **Link checking.** Every internal and external link checker reports the site as clean. A broken link is indistinguishable from a working one. - **Uptime and synthetic monitoring.** A check that asserts "status is 200" passes against a site serving nothing but its error screen. - **Caching.** Intermediaries happily store a successful response, so an error document can be cached under a real URL, occasionally outliving the mistake that produced it. - **Assets, and this is the sharp one.** If the rewrite covers asset paths too, a request for a hashed script that is missing - a stale reference, a partial upload - returns the HTML shell with a success status. The browser tries to parse HTML as JavaScript and reports a syntax error, so the failure looks like a bundler bug rather than a missing file. | Consumer | What it reads | What it concludes | |---|---|---| | A search crawler | the status line | the URL is a real page worth indexing | | A link checker | the status line | every broken link on the site is fine | | An uptime check | the status line | the service is healthy | | An intermediary cache | status and headers | this response is worth storing | | A human | the body | the page is missing, as intended | ## The fix, in order 1. **Prerender a document per route you can enumerate.** The more routes have real files, the fewer requests reach a fallback at all, and the fallback's behaviour stops being load-bearing. 2. **Use the host's not-found setting**, the one that serves a nominated document *with* a `404` status, rather than a rewrite, for genuinely unmatched paths. 3. **Scope the catch-all.** Rewrite only the prefixes that actually own client-resolved routes. A blanket rule over the whole site is what drags assets and typos into the success path. 4. **Exclude asset directories explicitly**, so a missing hashed file returns a real miss and the failure reads as what it is. 5. **When the client router decides a record does not exist**, present the not-found state honestly in the UI and avoid canonical or index signals for that URL - the status is already spent, so the page should not claim to be content. 6. **Assert the status in CI.** One probe against a deliberately absent path, asserting `404`, prevents the regression from ever shipping twice. ## The trade-off to state out loud There is a real tension between "deep links must survive a reload" and "unknown URLs must report a miss", and on a files-only target you cannot fully have both for client-resolved routes: the host must choose a file before anyone knows whether the route exists. The engineering answer is to shrink the ambiguous set - prerender what can be enumerated, scope the fallback tightly - so the remaining soft 404s cover a small, known set of paths rather than the entire URL space.

  • Why can the client router not simply set the status once it decides the route is unknown?
    The status line was sent before the document body, which was sent before the bundle ran. By the time the router matches, the response is finished. The client can change what is displayed and what the page signals to indexers, but not the status that was already transmitted.
  • Does a catch-all rewrite hurt routes that were prerendered?
    No - those paths match a stored file first, so the rule never fires for them. That is why prerendering everything enumerable is the cheapest part of the fix: it shrinks the set of requests whose status depends on the fallback.
  • How would you detect this on a site that already shipped?
    Request a path that certainly does not exist and read the status line, not the page. Then request a missing file under the asset prefix and confirm the response is not HTML. Both probes belong in CI afterwards, because the defect returns whenever hosting rules are edited.

saying these in an interview costs you the question

  • Saying the status does not matter if the page reads as an error
  • Assuming crawlers infer a miss from the page's wording
  • Believing client-side routing requires answering unknown URLs with 200
  • Expecting uptime monitoring to catch it, when the check asserts success
  • Routing asset paths through the same fallback as documents
  • Treating a canonical link as a substitute for the status code