A reverse proxy is configured to try the next upstream when a request fails, allowing up to two extra attempts with a 10-second read timeout each, while the calling client gives up after 15 seconds. Walk through what the client and the backends actually experience.
answer
- multiply attempts by per-attempt timeout
- compare that to the client's budget
- last attempt runs into a closed connection
- retries add load to a slow backend
- retry only what never reached the app
basics
~20 sWorst case is 30 seconds of proxy attempts against a 15-second client, so later attempts run after the client has already left. The backends do up to three times the work per user request, and slow backends turn retries into a load multiplier.
solid answer
~50 sDo the arithmetic first: three attempts at a 10-second read timeout is up to 30 seconds, and the client leaves at 15. So attempt two can barely land and attempt three cannot possibly reach anyone — it occupies a backend worker for ten seconds producing a response into a closed connection. Meanwhile the client, having timed out, very likely retries, and each of *its* attempts fans back out to three of yours. The dangerous case is not a dead backend but a **slow** one: retrying a timeout adds load to a system that is already struggling, and each attempt holds a worker for the full timeout. The fix is to make the numbers consistent — attempts times per-attempt timeout, plus margin, must fit inside the client's budget — and to restrict which failures retry at all, ideally to ones where the request provably never reached the application.
go deeper
Know that a proxy retry repeats the whole request against another backend, so the total time can be several times the single-attempt timeout.
Be able to do the arithmetic out loud — attempts times per-attempt timeout against the caller's budget — and say which failures are safe to retry versus which may already have taken effect.
Demonstrate the load reasoning: retries multiply work exactly when the fleet is least able to absorb it, and each attempt against a slow backend pins a worker for the full timeout.
Own the defaults across tiers — a narrow retry condition, an attempt budget that fits the caller's, and a cap on the share of traffic that may be retries so no single failure multiplies fleet-wide load.
## The arithmetic nobody does Retry settings and timeout settings are usually configured by different people at different times, and each looks reasonable alone. Multiplied, they are not: ``` attempts (1 initial + 2 retries) = 3 per-attempt read timeout = 10s proxy worst case = 30s client budget = 15s ``` The first thing to establish is whether the retries can *ever* produce a result the client will see. Here, only if the failing attempts fail fast. A connection refused returns in milliseconds, so three quick failures plus a successful attempt land comfortably. But a retry triggered by a **read timeout** consumes the full ten seconds before it even begins, which means retry number two starts at t=20s — five seconds after the client stopped listening. So: retries on connect failures are useful here, retries on timeouts are not. That distinction is the whole design. ## What the backends experience From the backend fleet's point of view, one user request is up to three executions. Add the client's own retries — if it tries three times too — and a single user action can be nine backend executions. That is fine when the fleet is healthy and catastrophic when it is not, because the condition that triggers retries is precisely the condition where extra load is most harmful. The slow-backend case deserves spelling out. If a backend is not dead but merely saturated, every attempt holds one of its workers for the full ten-second timeout before being abandoned. Retrying *adds* concurrent work to a system whose problem is too much concurrent work. And if one member of the pool is slow, retries redirect its overflow onto the healthy members, which is helpful in small doses and is how a partial degradation becomes a total one in large ones. ## Which failures are safe to retry at a proxy A proxy has less information than the client does. It cannot see application semantics; it sees a TCP connection and some bytes. The conservative classification: - **Never wrote a byte to the upstream** — connection refused, connect timeout, DNS failure. The application demonstrably did not see this request, so another member is safe for any method. - **Wrote the request, no response** — a read timeout or a reset mid-flight. The application may have completed the work. Retrying is only safe for idempotent methods, or for requests carrying an idempotency key the application honours. - **Got a 5xx** — the backend answered, so it processed something. Whether that is retryable is an application-level question the proxy cannot answer for itself. The practical setting is to enable retries for the first class by default, and to enable the second only where the request semantics justify it. nginx expresses these conditions in `proxy_next_upstream`; HAProxy's `option redispatch` covers the reconnect-to-another-server case. The default sets in these products are not always as narrow as you would choose, so read them rather than assuming. ## Making the numbers fit Work backwards from the client: ``` client budget 15s safety margin 3s usable proxy budget 12s attempts 3 per-attempt read timeout 4s ``` Four seconds per attempt is a real change to behaviour — a legitimately slow endpoint will now fail — so this number has to come from the endpoint's latency distribution, not from the retry maths alone. If p99 is eight seconds, three attempts do not fit and the honest answer is fewer attempts, not a shorter timeout that fails healthy requests. Two more guards belong here. First, a cap on the *proportion* of traffic that may be retries, so a fleet-wide failure cannot multiply total load — retry budgets as a discipline are covered under resilience patterns, but the proxy is where you enforce them for this tier. Second, backoff between attempts, which matters less at a proxy than in a client because the proxy is usually moving to a *different* upstream rather than hammering the same one. ## How it shows up in an incident The signature is a backend request rate that rises when the error rate rises, in a fixed ratio close to your attempt count, while the client-observed request rate stays flat. If your dashboards only show requests *received* by the proxy, the multiplication is invisible; instrument attempts, not just requests, or you will diagnose the wrong thing.
- Why is a slow backend more dangerous to retry than a dead one?A dead backend refuses the connection in milliseconds, so retrying is cheap and usually correct. A slow one accepts the connection and holds a worker for the full timeout on every attempt, so each retry adds concurrent load to a system whose problem is already too much concurrency. That is how a degraded service becomes an unavailable one.
- If the proxy retries and the client also retries, what does the backend see?The product of the two. Three proxy attempts under three client attempts is up to nine executions of one user action, arriving in a burst precisely during a failure. Any retry policy has to be reasoned about across the whole chain rather than per hop, and the total is bounded by capping the share of traffic that may be retries rather than by tuning any single tier.
- How would you tell from monitoring that retries are amplifying an incident?Compare requests entering the proxy with attempts leaving it toward upstreams. In a healthy period those track each other; during amplification the upstream attempt rate rises toward your attempt multiplier while the incoming rate is flat. If you only graph received requests, the multiplication is invisible and you will misread the backend's load as organic.
saying these in an interview costs you the question
- Sets retry count and timeouts without multiplying them
- Thinks retries always improve availability
- Retries a read timeout on a POST as if it were safe
- Assumes the proxy knows the client's remaining deadline
- Measures requests received rather than upstream attempts