In open-source nginx, a backend in an `upstream` group crashes and some clients still get 502s for the next several seconds. How do `max_fails`, `fail_timeout` and `proxy_next_upstream` decide that, and what can open-source nginx not do here?
answer
- nginx OSS never probes on its own
- failures come from real traffic
- fail_timeout means two different things
- the failure set is proxy_next_upstream's list
- a lone server is never marked down
basics
~10 sOpen-source nginx marks backends down only passively, from real requests: max_fails failures within fail_timeout take a server out for fail_timeout seconds. Failures are whatever proxy_next_upstream counts, and requests already partly answered cannot be retried.
solid answer
~50 sNginx OSS has no active probing of HTTP upstreams; a dead server is discovered when a client request hits it. The `server` parameters `max_fails` (default 1) and `fail_timeout` (default 10s) do double duty: `max_fails` failures inside a `fail_timeout` window mark the server unavailable, and it stays unavailable for `fail_timeout` before nginx tries it again with a live request. What counts as a failure is defined by `proxy_next_upstream`, whose default is `error timeout` — so a backend returning HTTP 500 is not a failure unless you add `http_500`. Retrying onto another server is only possible while nothing has been sent to the client yet, and since 1.9.13 non-idempotent methods such as POST are not retried unless you add the `non_idempotent` option. Two facts explain the residual 502s: a group with a single server ignores max_fails entirely, and the detection window is by definition paid for in failed requests. Active health checks and slow start are nginx Plus features.
code
nginx · 15 linesupstream backend {
zone backend 64k;
server 10.0.0.11:8080 max_fails=3 fail_timeout=15s;
server 10.0.0.12:8080 max_fails=3 fail_timeout=15s;
}
server {
location / {
proxy_pass http://backend;
proxy_next_upstream error timeout http_502 http_503;
proxy_next_upstream_tries 2;
proxy_next_upstream_timeout 5s;
proxy_connect_timeout 2s;
}
}go deeper
Know that max_fails and fail_timeout take a bad backend out of rotation for a while, and that nginx learns it is bad from real requests rather than from a health probe.
Explain fail_timeout's two roles, the defaults of 1 and 10s, and that the failure set is whatever proxy_next_upstream lists — which by default excludes every HTTP status code.
Reason about the trade in production: detection costs failed requests, aggressive settings eject healthy servers, retries on 5xx amplify load during an incident, and POSTs are excluded for a reason. Mention the per-worker counters and the shared zone.
Decide where health decisions belong for the platform: passive marking in a proxy you already run, versus a tier that probes actively, versus discovery-driven membership — and set a fleet-wide policy on which failures may be retried so retries never amplify an outage.
## Passive marking is the only mechanism in OSS Open-source nginx does not poll HTTP upstreams. It learns a server is bad the way a user does — by sending it a request and failing. That single fact explains the shape of every symptom here, and it is what the `health_check` directive in nginx Plus (and third-party modules) exists to change. ```nginx upstream backend { server 10.0.0.11:8080 max_fails=3 fail_timeout=15s; server 10.0.0.12:8080 max_fails=3 fail_timeout=15s; server 10.0.0.13:8080 backup; } ``` ## The two parameters, and their double meaning `fail_timeout` is used twice, which is the part people get wrong: 1. It is the **window** in which `max_fails` failures must occur for the server to be considered unavailable. 2. It is the **duration** the server is then considered unavailable. Defaults are `max_fails=1` and `fail_timeout=10s`: one failure and the server is out for ten seconds. `max_fails=0` disables the marking entirely, so the server is always tried — occasionally useful, usually a mistake. After the period elapses, nginx does not probe. It routes a **real client request** to the server; success restores it, failure starts the clock again. So the recovery path also costs a request. Two behaviours worth knowing precisely: if the group contains only one server, `max_fails` and `fail_timeout` are ignored and the server is never taken out — there would be nowhere else to go. And when every server has been marked unavailable, nginx resets their state and tries again rather than failing everything outright. ## What counts as a failure Not obvious, and configurable: the failure set is exactly what `proxy_next_upstream` lists. The default is: ```nginx proxy_next_upstream error timeout; ``` `error` covers connection failures and errors while sending or reading; `timeout` covers the connect, send and read timeouts. Notably absent are status codes. A backend that answers every request with `500` is, by default, perfectly healthy from nginx's point of view — requests are neither retried elsewhere nor counted toward `max_fails`. If you want application errors to trigger both behaviours, list them: ```nginx proxy_next_upstream error timeout http_502 http_503 http_504; ``` Be deliberate: adding `http_500` means a genuine application bug that returns 500 for one request will be replayed against every remaining backend, multiplying load exactly when the system is unhealthy. `proxy_next_upstream_tries` bounds how many servers are attempted (0 = unlimited, the default) and `proxy_next_upstream_timeout` bounds the total time spent across attempts. Setting at least `_tries` is good hygiene. ## When a retry is impossible A request can only be passed to the next server **while nothing has been sent to the client yet**. Once nginx has begun streaming a response — which happens early if `proxy_buffering` is off or the response is large — a mid-stream upstream failure cannot be masked; the client gets a truncated response. Since nginx 1.9.13, requests with non-idempotent methods (POST, LOCK, PATCH) are **not** passed to the next server by default. The `non_idempotent` parameter re-enables it, and you should only add it if the backend deduplicates, because a timeout does not tell you whether the first server processed the request. ## Why 502s persist for several seconds Put it together for the interview answer: - Detection is paid for in failed requests — with `max_fails=3`, three clients per worker eat an error before the server is out. - **Marking is per worker process.** Each worker keeps its own failure counters in the default configuration, so with eight workers the pool learns about the dead backend eight times over. The `zone` directive puts the group's state in shared memory so workers agree, and is what makes marking behave the way people assume it already does. - After `fail_timeout` expires, nginx probes with a real request — another potential 502 if the server is still down. - POSTs are not retried, so those users see the error even when a retry would have succeeded. The honest framing is that passive marking trades a small number of failed requests for zero probing traffic, and open-source nginx gives you no way to buy detection without those failures. Tuning `max_fails` down shortens the loss but makes a transient blip eject a healthy server, and `proxy_next_upstream` is what actually hides most single failures from clients.
- A backend returns HTTP 500 for every request. Does nginx take it out of the pool?Not by default. proxy_next_upstream defaults to `error timeout`, so a valid HTTP response — whatever its status — is a success as far as the upstream module is concerned. It is neither retried elsewhere nor counted toward max_fails. You would have to add http_500 explicitly, weighing that against replaying genuine application errors across every remaining backend.
- Why are POST requests not retried on another upstream, and when would you change that?Since nginx 1.9.13, non-idempotent methods are excluded from proxy_next_upstream by default because a timeout does not reveal whether the first server already processed the request; a retry risks a duplicate side effect. The `non_idempotent` parameter re-enables it, and it is only safe when the endpoint deduplicates — typically via an idempotency key the client supplies.
- Why can the same dead backend be discovered separately by each worker process?Without a `zone` directive, an upstream group's failure counters live in each worker's own memory, so every worker must independently accumulate max_fails failures before it stops using the server. Declaring the group with a shared-memory zone puts the state where all workers see it, so one worker's detection benefits the rest and the pool converges much faster.
saying these in an interview costs you the question
- Believes open-source nginx actively probes upstream health
- Thinks a 5xx response automatically counts as a failure
- Reads fail_timeout as only the ejection duration
- Expects POST requests to be retried on the next server by default
- Assumes a single-server upstream group can be marked unavailable