Your load balancer's access logs show a burst of HTTP 502 and 504 responses, while the application servers behind it log almost no errors. How do you work out where the failure actually is?
answer
- Codes come from the proxy — start there
- 502 = crash / reset / keep-alive race / header buffer
- 504 = latency vs read timeout
- Timeout budget must shrink inward
- Request id: did the origin ever see it?
basics
~20 sSplit by code: 502 means the proxy could not get a valid response (crashes, resets, closed keep-alives, oversized headers); 504 means upstream was too slow. Correlate proxy and origin logs on a request id, then check process restarts, connection reuse and timeout budgets.
solid answer
~60 sFirst, separate the two codes, because they accuse different things. **502** — the proxy got no usable response. Check: are origin processes restarting or OOM-killed? Is a rolling deploy terminating instances still in rotation? Are keep-alive idle timeouts mismatched, so the proxy reuses a connection the origin has just closed? Are response headers exceeding the proxy's buffer limits (a large `Set-Cookie` or auth header is the classic)? **504** — the proxy timed out. Check origin latency percentiles and the timeout budget down the chain: the proxy's read timeout versus the app's own client timeouts versus the database. A 504 with the origin logging a successful 200 seconds later is the signature. The decisive tool is **correlation**: a request id injected at the edge and logged at every hop, so you can prove whether the origin ever saw the request. "Origin has no matching log line" versus "origin logged 200 after the proxy gave up" separates 502 from 504 root causes immediately. Then look at whether the errors are spread across all instances or concentrated on a few.
code
bash · 6 linescurl -sS -o /dev/null -w '%{http_code} %{time_total}\n' \
-H 'X-Request-Id: probe-1' https://api.example.com/v1/orders/42
curl -sS -o /dev/null -w '%{http_code} %{time_total}\n' \
-H 'Host: api.example.com' -H 'X-Request-Id: probe-1' \
http://10.0.3.17:8080/v1/orders/42go deeper
Say that these codes come from the proxy, that 502 means an invalid or failed upstream response and 504 means an upstream timeout, and that you would check whether the app servers are healthy.
Give concrete causes per code — restarts, keep-alive mismatch, header buffers for 502; latency versus read timeout for 504 — and know to compare proxy and origin logs.
Drive a structured investigation: split the signal by code and instance, audit the end-to-end timeout budget, correlate on a request id, and reason about side effects and retry safety.
Turn it into policy: standard timeout budgets that shrink inward, drain-before-terminate rollouts, request-id propagation as a platform requirement, and load shedding with 503 rather than letting the edge time out.
## Why the origin's clean logs are the clue, not a contradiction 502 and 504 are emitted **by the intermediary**, not by the application. So a quiet origin error log is exactly what you expect: in the 502 case the application often never ran, and in the 504 case it ran and believed it succeeded. Any diagnosis that starts by grepping application exceptions will find nothing and stall. Start instead at the hop that generated the status. ## Step 1: split the signal Chart 502 and 504 separately over time, per upstream instance and per route. The shapes tell you a lot: - **502 spikes aligned with deploys or restarts** — lifecycle problem. - **502 at a steady low rate, evenly spread** — connection reuse or header-size problem. - **504 rising with latency percentiles** — a saturation or dependency problem. - **Errors concentrated on one or two instances** — a sick host, not a systemic fault. ## Step 2: the 502 checklist 1. **Process health.** Are origin processes crashing, being OOM-killed, or restarting? Correlate 502 timestamps with restart counts and kernel OOM messages. 2. **Deploy and drain.** During a rolling update, instances must be removed from rotation and allowed to finish in-flight requests before the process is signalled. If termination beats deregistration, the proxy forwards to a socket that is closing — refused or reset connections, and 502s. 3. **Keep-alive races.** If the origin's idle keep-alive timeout is shorter than the proxy's, the proxy will eventually pick a pooled connection the origin has just closed and get a reset mid-request. The fix is to make the upstream's idle timeout comfortably longer than the proxy's, and to allow retry on idempotent requests that fail before any bytes of the response arrive. 4. **Header and buffer limits.** A response whose header block exceeds the proxy's buffer is rejected as invalid. Large cookies, long redirect targets and verbose auth headers are the usual culprits. 5. **Protocol mismatch.** Speaking plaintext to a TLS listener, or an HTTP/2 expectation against an HTTP/1.1-only upstream, yields bytes the proxy cannot parse. ## Step 3: the 504 checklist 1. **Compare latency to the timeout.** Pull the origin's response-time distribution for the affected route. If p99 is brushing the proxy's read timeout, the 504 rate is simply the tail crossing the line. 2. **Audit the timeout budget end to end.** Each hop should time out sooner than the hop in front of it, so failure surfaces at the layer that can act. When an inner timeout is longer than the outer one, the outer hop always gives up first and you get 504s while the inner call keeps working — and keeps holding a connection and a thread. 3. **Find the slow dependency.** Traces beat logs here: which span dominates? Lock waits, a saturated connection pool, an N+1 query, a slow third-party call. 4. **Check for queueing.** If requests are waiting in an accept queue or a thread pool before work begins, origin-side timers that start at handler entry will look healthy while wall-clock latency at the edge is terrible. Measure time from connection accept, not from handler start. ## Step 4: correlate Inject a request id at the edge, propagate it downstream, and log it everywhere. Then for a sample of failures ask one question: **did the origin ever see this request?** - No matching origin line → the request died before or during forwarding: pure 502 territory. - Origin line exists, status 200, completing after the edge already responded → 504 confirmed; the work continued after the client gave up. - Origin line exists with a 5xx of its own → the proxy may be translating an origin failure, and you are back to reading application logs after all. ## Step 5: consider correctness, not just availability 504 has a nasty property: the origin may still complete the work. If the route is non-idempotent — creating a payment, sending a message — a client that retries on 504 can duplicate the effect. Part of the fix is operational (raise capacity, cut latency) and part is contractual (make the operation safely repeatable). Similarly, retry-on-502 is usually safe but should be restricted to failures that occurred before any response bytes were received. ## What good answers include Interviewers listen for three things: that you know the codes come from the proxy and therefore start at the proxy; that you have concrete, named causes rather than "check the network"; and that you mention correlation and timeout budgets rather than only raising limits. Raising the proxy timeout to make 504s disappear is the answer that fails the question — it converts a fast failure into a slow one and moves the queue upstream.
- Why does raising the proxy's read timeout usually fail as a fix for 504s?It converts a fast failure into a slow one. The underlying work is still too slow, so requests now occupy proxy connections, origin threads and database sessions for longer, which increases queueing and pushes the system further toward saturation. It also delays the client's own timeout, worsening user-visible latency. The real fixes are reducing the work, adding capacity, or shedding load early with 503.
- What is the keep-alive race that produces low-level, constant 502s, and how do you fix it?The proxy pools idle connections to the origin. If the origin's idle timeout is shorter than the proxy's, the origin closes a connection at roughly the moment the proxy chooses to reuse it, and the request fails with a reset. The fix is to make the upstream's keep-alive idle timeout meaningfully longer than the proxy's, and to allow the proxy to retry when the failure happened before any response bytes were received.
- During a 504, what should you assume about side effects at the origin?Assume the work may still complete. The origin accepted the request and was merely slow, so a create or charge may commit after the edge has already answered the client. That makes blind client retries on 504 a duplication risk for non-idempotent operations, and it is why an idempotency mechanism, rather than the status code, is what makes retrying safe.
saying these in an interview costs you the question
- Starting the investigation in application exception logs, when the codes were generated by the proxy.
- Treating 502 and 504 as interchangeable "upstream problems".
- Fixing 504s by simply increasing the proxy timeout.
- Assuming a 504 means the origin did no work, and retrying non-idempotent requests.
- Never correlating proxy and origin logs, so "did the origin see the request" is left unanswered.