skip to content

Why can a request-level (layer 7) proxy retry a failed attempt on a different backend when a connection-level (layer 4) proxy generally cannot, and what must the layer 7 proxy do to make that retry possible?

level: seniorimportance: nice to knowfreq 38%

answer

  1. you can only resend what you kept
  2. request boundaries exist only at layer 7
  3. connect-time retry versus request retry
  4. the body has to sit somewhere
  5. headers forwarded means no going back

basics

~20 s

A layer 7 proxy owns the request as a unit — it parsed it, can hold the body, and opened the upstream connection itself — so it can send the same request to another backend before the client sees anything. A layer 4 proxy forwards an uninterpreted byte stream and has kept nothing to resend.

solid answer

~50 s

Retrying means resending a request, and only a proxy that knows what a request is can do that. A layer 7 proxy has parsed the request, can buffer its body, and opened the upstream connection on its own behalf, so if the upstream refuses the connection or resets before response headers arrive, it can pick another backend, replay the request, and the client never knows. A layer 4 proxy is copying an opaque stream: it has no request boundaries and has not retained the bytes, so once data has flowed it can only propagate the failure as a reset — it can re-select a backend on a failed *connect*, before anything was sent, and that is all. The layer 7 proxy pays for the capability: it must buffer the body within a size cap, must confine automatic retries to requests that are safe to repeat, and must stop retrying the moment it has forwarded response headers.

go deeper

for a junior

Know that only a proxy which understands requests can resend one, and that a byte-forwarding proxy can react to a failed connection but not re-send data it has already passed on.

for a middle

Explain the requirement chain: the request must be parsed and its body buffered, the upstream connection must be the proxy's own, and the retry must happen before any response is relayed to the client.

for a senior

Demonstrate judgment about when replay is safe — connect failures and resets before any response versus an ambiguous drop after the body was sent — and about the buffer-size and streaming limits that quietly disable retries in production.

for a principal

Own the policy question of where retries live at all: proxy, client library or mesh, how attempts are capped so a degraded pool is not amplified, and which endpoints are declared repeatable in the first place.

## Retry is a request-level idea The word "retry" presupposes a unit of work with a beginning and an end. That unit exists at layer 7 and does not exist at layer 4. A connection-level proxy sees an unbroken stream of bytes; there is no point in it where one thing finished and another can be attempted, and it did not keep the bytes it already forwarded. If the upstream dies mid-stream, the only honest thing it can do is tear the client connection down and let the client decide. The one exception is the moment before anything has been sent. If the upstream connection attempt itself fails — refused, unreachable, timed out on connect — the proxy has forwarded nothing, so it can choose another backend and try again transparently. That is genuinely useful (it is what removes a dead backend from the picture during a rolling restart) but it is a connect-time retry, not a request retry. ## What the layer 7 proxy actually holds A request-level proxy is in a different position: - it has the request **parsed and in memory**: method, path, headers; - it can hold the **body**, either fully buffered or at least the part read so far; - it created the upstream connection itself, so nothing about that connection is visible to the client; - it knows exactly how far the exchange got, because it is parsing the response too. So when the upstream connection is refused, or resets after the request was sent but before response headers came back, or the per-attempt timeout expires, the proxy can select another backend and re-send. The client sees one request and one response. ## The buffering requirement Replaying a request means having the request. For a GET with no body that is free. For a POST or PUT it means the proxy must hold the body until the attempt succeeds, and that has hard edges: - **A size cap.** Buffering is memory (or disk) proportional to body size times concurrency. Every real proxy has a limit, and above it the choice is to stop buffering — losing retry ability — or reject the request outright. - **Streaming is incompatible.** If the proxy is streaming the body through as it arrives to keep latency low or to support uploads of unknown size, it cannot replay what it has already handed over. - **Latency.** Full buffering delays the upstream's first byte until the client's last byte, which is fine for small JSON and wrong for large uploads. ## The point of no return Once the proxy has forwarded response headers to the client, retrying is over. The client has already seen a status line; you cannot substitute a different attempt's response behind it. This is why per-attempt behaviour matters: the proxy must decide to retry while it still owns the whole exchange, which in practice means before it starts relaying the response. ## Safety, not just capability Being able to replay does not make replaying correct. A request that the upstream actually processed before the connection dropped will be processed twice. Proxies therefore restrict automatic retry to conditions where the upstream provably did nothing (connection refused, reset before any response) or to requests whose method is safe to repeat, and leave anything else to explicit configuration by someone who knows whether that endpoint is idempotent. "The connection dropped after we sent the body" is precisely the ambiguous case, and treating it as automatically retryable is how duplicate charges happen. ## And retries are load Retrying into a failing backend pool multiplies traffic exactly when the pool is least able to take it, so proxies bound the behaviour — a maximum number of attempts, and often a cap on what fraction of traffic may be retries. Worth naming in an answer as the reason retries are always bounded, even though the arithmetic of layered timeouts and budgets is a topic of its own. ## Summing up the altitude difference Layer 4 can pick a different backend when the connection never got established. Layer 7 can pick a different backend for a request that was sent and failed — but only because it kept the request, only until it has committed to a response, and only where repeating the work is acceptable.

  • What can a connection-level proxy still retry, and when?
    Only the upstream connection attempt. If the connect is refused, unreachable or times out before any bytes were relayed, the proxy can select another backend transparently, because nothing has been forwarded and the client is still waiting. Once the stream is flowing, a failure can only be surfaced as a reset.
  • An upload endpoint has retries configured but never retries in practice. What is the likely reason?
    The body is too large to buffer, or buffering is disabled to keep the upload streaming. Once the proxy has relayed body bytes upstream it cannot replay them, so it silently loses the retry capability. Either raise the buffer cap for that route and accept the memory and latency cost, or accept that this route is not retryable.
  • Why is 'the connection dropped after the request was sent' the dangerous case to retry automatically?
    Because the upstream may have completed the work and died before answering. Replaying then performs it twice. It is safe only for operations that are idempotent by construction or protected by a deduplication key; otherwise the proxy should surface the failure and let the caller decide.

saying these in an interview costs you the question

  • Claiming a layer 4 proxy retries requests it has already forwarded
  • Confusing TCP retransmission with retrying a request
  • Assuming any request can be replayed without buffering the body
  • Retrying after response headers have reached the client
  • Treating a drop after the request was sent as safe to repeat

context