skip to content

Users get a 504 after about 30 seconds on a request that crosses a CDN, an edge load balancer and an ingress proxy, but the application log shows that same request completing successfully after 45 seconds. Explain what happened, and how you would set the timeouts across those hops.

level: seniorimportance: must knowfreq 62%

answer

  1. three stopwatches, no shared deadline
  2. effective timeout is the shortest hop
  3. the 504 was invented by one tier
  4. work continues after the caller gives up
  5. increase timeouts outward from the app

basics

~20 s

One hop's timeout fired at 30 seconds and synthesised the 504 while the application kept working to 45. Proxy hops run independent fixed timers with no shared deadline, so the effective timeout is the shortest hop, not the sum.

solid answer

~50 s

Each proxy in the path runs its own timer, and nothing carries a deadline between them: there is no header saying "you have 12 seconds left". So the effective end-to-end timeout is simply the **shortest** hop, and the hop that fires first is the one that manufactures the 504 and returns it to the client. Everything downstream of it may keep running — your application burned a worker for the remaining 15 seconds producing a response nobody read. The diagnosis is to find which tier emitted the error: correlate the tiers' access logs on a request ID and look for the one that logged an upstream timeout rather than a passed-through 504. For the fix, order the timeouts so they *increase outward* — the innermost hop shortest, each outer hop a small margin longer — so the tier closest to the failure is the one that times out and reports the precise cause. Then make sure the whole budget is shorter than the client's own patience.

code

bash · 3 lines
bash
curl -o /dev/null -s \
  -w 'code=%{http_code} connect=%{time_connect}s ttfb=%{time_starttransfer}s total=%{time_total}s\n' \
  https://example.com/api/report

go deeper

for a junior

Know that every proxy in the path has its own timeout, and the shortest one decides when the user sees an error, whatever the others are set to.

for a middle

Be able to explain why a 504 can appear while the backend later succeeds, and why nothing propagates a remaining-time deadline between ordinary HTTP hops.

for a senior

Show the diagnosis: correlate tiers on a request ID, read per-tier timing fields to find which hop generated the status, then order the ladder so the innermost hop fires first.

for a principal

Own the ladder as policy — a documented budget per tier including the client, monotonic from the application outward, with the total derived from what the product can actually wait for.

## Why the numbers disagree A request that crosses three proxies is governed by three independent stopwatches. None of them knows about the others. HTTP has no standard deadline field a proxy propagates, so each hop starts its own timer when it forwards the request and gives up when its own number is reached, regardless of how much budget the hops above or below thought they had. That produces two consequences worth stating out loud: 1. **The effective timeout is the minimum, not the sum.** Hops of 60s, 30s and 45s give you a 30-second service, and the other two numbers are decoration. 2. **A timeout upstream does not stop work downstream.** When the 30-second hop gives up it responds 504 toward the client and stops reading. Whether the application is *told* depends on whether the abandoned connection is actually closed and whether the application checks — so the common outcome is exactly what the logs show: a successful 45-second response into a void. ## Finding which hop fired The 504 the user sees was invented by some tier. Identify it before changing any number: - **Correlate on a request ID.** Every tier should stamp and log the same identifier. The tier that logged a *generated* 504 while its own upstream logged nothing (or logged a later success) is your culprit; tiers above it merely relayed the status. - **Read the tier-specific timing fields.** nginx's log can include `$upstream_response_time` next to `$request_time`, which separates "the backend was slow" from "the proxy itself was slow". Envoy's `%RESPONSE_FLAGS%` marks an upstream request timeout as `UT`. HAProxy encodes an equivalent in its termination-state field. - **Look at the duration.** A cluster of failures at a suspiciously round wall-clock duration — always 30.0s, always 60.0s — is a configured timer, not a backend. - **Reproduce with timing on the client.** `curl -w` breaks the request into connect, TTFB and total, so you can see whether the failure lands on the connect phase or waiting for the first response byte. ## Setting the budget The design rule is that timeouts should **increase as you move outward from the application**, each layer allowing a little more than the layer beneath it: ``` application internal deadline 10s ingress proxy read timeout 12s edge load balancer 15s CDN / origin timeout 20s client (browser or SDK) 25s ``` The reason is diagnostic, not arithmetic. The hop nearest the failure has the most specific information about it. If the innermost timer fires first, the application or the ingress reports a precise error you can act on. If the outermost fires first, every inner tier is still happily working and all you get is a generic gateway timeout at the edge, with wasted capacity behind it. The margins should be small and deliberate. A layer that allows *much* more than the one beneath it is not adding safety, it is adding a window in which work continues past the point anyone cares. And the total must sit inside the client's own limit: if the browser or SDK gives up at 10 seconds, a beautifully layered 25-second budget is irrelevant, and the user experience is decided by a timer you do not control. ## Where people get this backwards - **Raising the edge timeout to "fix" 504s.** It converts fast failures into slow ones and holds connections open at the tier with the least capacity to spare. - **Assuming a deadline propagates.** Some protocols do carry one — gRPC has a deadline that intermediaries can honour — but a chain of ordinary HTTP proxies does not, and a header your team invented is only honoured by tiers you configured to read it. - **Setting the same number everywhere.** Equal timeouts make the firing hop effectively random, which makes the error message random too. - **Ignoring the shape of the timers.** "Read timeout" usually means time without receiving *any* bytes, not total request duration, so a slowly streaming response can outlive it comfortably while a stalled one dies quickly. If you need a bound on the whole exchange, you need the timer that actually measures the whole exchange. ## The takeaway Write the whole ladder down, client included, as one column of numbers. If it is not monotonic from the inside out, some tier is doing work whose result will be discarded, and some incident will be diagnosed at the wrong layer.

  • Why not simply raise the edge timeout until the 504s stop?
    Because it treats the symptom at the most expensive layer. Longer edge timeouts hold client connections and edge capacity open through failures, turn fast errors into slow ones, and still leave the real question — why a request takes 45 seconds — unanswered. Raise a timeout only when you have decided that latency is legitimate, and then raise the whole ladder coherently.
  • What is the difference between a read timeout and a total request timeout at a proxy?
    A read timeout measures the gap since the last byte arrived, so a response that trickles bytes steadily can run far longer than the number suggests. A total timeout bounds the whole exchange regardless of activity. Streaming endpoints need the first; anything where you must guarantee an upper bound needs the second, and mixing them up is why some requests outlive their stated limit.
  • Does the application stop working when the proxy times out?
    Not automatically. The proxy stops reading and may close the upstream connection, but whether the application notices depends on it observing that closure and on the framework surfacing it as cancellation. In practice much abandoned work runs to completion, which is why an aggressive outer timeout under load can leave the backend saturated with requests nobody will ever read.
  • Some protocols do carry a deadline — how does that change the picture?
    gRPC carries a deadline that intermediaries and the server can observe, so remaining time travels with the call and each hop can fail fast rather than guess. Plain HTTP through a chain of proxies has no equivalent, so the layered-ladder discipline is the substitute: independent timers arranged so the most informed hop fires first.

Nested egg timers in a kitchen: whichever one rings first ends the meal, no matter how much cooking is still happening at the stove.

saying these in an interview costs you the question

  • Adds the hops' timeouts together to get the budget
  • Raises the edge timeout to make 504s go away
  • Assumes a deadline is propagated between proxy hops
  • Thinks the application stops when the proxy gives up
  • Sets the same timeout at every tier and expects a clear error

context