skip to content

You own a platform where dozens of services call each other over HTTP. How would you set policy for which 5xx status code each service returns under overload, dependency failure and deploy, and how should client retry behaviour and the Retry-After header follow from that policy?

level: principalimportance: should knowfreq 38%

answer

  1. 503 = deliberate + temporary + Retry-After
  2. 500 = defect rate, keep it rare
  3. 502/504 belong to intermediaries only
  4. Fast shed beats slow timeout
  5. Retry budget, not just attempt caps

basics

~20 s

Make the codes mean something fleet-wide: 503 for deliberate, temporary refusal (overload, drain, open breaker) with Retry-After; 500 only for genuine defects; never hand-write 502/504, which belong to proxies. Then let the shared client retry on 503/502 with jitter and budgets, and require idempotency for anything else.

solid answer

~60 s

Treat the 5xx codes as a **fleet-wide vocabulary**, not per-service taste. - **503** — deliberate, temporary refusal the service chose: load shedding, queue overflow, draining during deploy, open circuit breaker, maintenance. Always with `Retry-After` when a real estimate exists. - **500** — genuine defect: an escaped exception. It should be rare and should page someone. Never use it for capacity. - **502 / 504** — reserved for intermediaries. Application code should not emit them; a service whose dependency failed should translate that into its own honest code rather than passing the shape upward. From that vocabulary the client policy follows mechanically, implemented once in a shared HTTP client: retry 503 (respecting `Retry-After`, plus jitter) and 502 when no response bytes were received; do **not** blind-retry 500 or 504 unless the operation is idempotent or carries an idempotency key. Cap retries with a **budget** — a percentage of total traffic, not just per-request attempt limits — so retries cannot amplify an outage. The payoff is that dashboards and alerts become meaningful: 500 rate is defect rate, 503 rate is saturation, and they route to different responses.

go deeper

for a junior

Focus on the vocabulary: 503 for temporary unavailability with Retry-After, 500 for unexpected errors, and that proxies are what emit 502 and 504.

for a middle

Add the client side: which codes are retryable, why Retry-After needs jitter, and why deploys should drain before terminating.

for a senior

Argue the shed-versus-collapse tradeoff concretely, describe timeout and deadline propagation, and connect codes to alerting that distinguishes defects from saturation.

for a principal

Own it as platform policy: a standard vocabulary enforced through shared middleware and a shared client, retry budgets to prevent amplification, deploy lifecycle requirements, and an explicit stance on whether shed traffic counts against the SLO.

## The problem: codes drift into noise In a large estate, each team picks its own 5xx conventions. One service returns 500 when its database is slow; another returns 503 for a null pointer; a third passes a downstream 502 straight through. Within a year the platform's aggregate 5xx metric means nothing: you cannot tell defects from saturation, alerting is either deafening or blind, and client retry logic has no principled basis. The fix is to standardise what each code *asserts* and to make the client library act on those assertions. ## The vocabulary **503 — I am refusing on purpose, and it is temporary.** This is the workhorse. It covers load shedding when a queue or pool crosses a threshold, an instance draining ahead of shutdown, an open circuit breaker in front of a sick dependency, and planned maintenance. The service is healthy enough to answer; it is choosing not to do work. Attach `Retry-After` whenever you have a defensible number: remaining drain time, breaker cooldown, queue depth divided by drain rate. **500 — something is broken that should not be.** Reserve it for unanticipated defects. If 500 only ever means "a bug fired", then the 500 rate is a defect rate and can page a human with a low false-positive rate. The moment capacity problems are also 500, that signal is destroyed. **502 and 504 — the intermediary's codes.** These describe a hop that failed between two components, and only the component in the middle can honestly assert them. When service A calls service B and B times out, A should not return 504 to its own caller — A is not acting as B's gateway from the caller's point of view. A decides what its own availability story is: shed with 503 if it cannot serve without B, or degrade and return a partial 200 if the contract allows. ## Choosing shed over collapse A fast 503 is strictly better than a slow 504. Once a service is saturated, every additional accepted request makes latency worse for the requests already in flight; the queue grows, timeouts fire at the edge, and the work is discarded after the cost has been paid. Explicit admission control — bounded queues, concurrency limits, shedding at a threshold with 503 — converts an unbounded latency collapse into a bounded, well-labelled error rate. This is the single most consequential item in the policy, and it is why the platform should make shedding easy: a shared middleware, not a per-team project. ## Client policy that follows The policy is only real if the shared client implements it. 1. **Retry 503**, honouring `Retry-After` as a floor, adding randomised jitter, and capping the delay so a synchronous path is not parked indefinitely. 2. **Retry 502** when the failure occurred before any response bytes arrived, which strongly implies the request never took effect. 3. **Do not blind-retry 500 or 504.** Both mean application code ran and may have committed. Retry only when the method is idempotent or the request carries an idempotency mechanism. 4. **Enforce a retry budget.** Per-request attempt caps are insufficient: during a partial outage, every client retrying three times triples load exactly when the system can least afford it. A budget expressed as a fraction of successful traffic — retries allowed only while they remain a small share of requests — bounds amplification. 5. **Propagate deadlines.** Each hop passes its remaining time budget downstream, so no service starts work that its caller has already abandoned. ## Deploy behaviour Rollouts are the biggest avoidable source of 502. The required sequence is: fail the health check or deregister first, let in-flight requests drain, then signal the process. Skipping deregistration means the proxy keeps forwarding to a socket that is closing, and the estate sees 502s on every deploy — which teams then learn to ignore, which is worse than the errors themselves. ## Observability contract Standardising codes pays off in dashboards. Split 5xx by code and by emitter (edge, mesh sidecar, application), alert on 500 rate as a defect signal and on 503 rate as a saturation signal, and treat sustained 502 as a lifecycle or connection-management defect. Require a request id propagated across hops so any 5xx can be walked end to end. Where an error budget or SLO exists, decide explicitly whether deliberate load shedding counts against availability — it usually should, or shedding becomes a way to hide capacity problems. ## The judgement being tested There is no single right answer here; the interviewer is listening for whether you treat status codes as an interface between operations and clients rather than as cosmetic detail, whether you can articulate the shed-versus-collapse tradeoff, and whether you recognise that client behaviour must be centralised for any of it to take effect.

  • Why is a per-request retry cap of three attempts insufficient protection during a partial outage?
    Because it bounds each caller individually while allowing the aggregate to triple exactly when the system is weakest. A service losing half its capacity sees offered load rise as every client retries, which pushes it fully over. A retry budget expressed as a fraction of successful traffic caps total retry volume across the fleet, so retries stop automatically once the error rate is high.
  • Service A depends on service B, and B times out. Should A return 504 to its caller?
    No. 504 asserts that the responder was acting as a gateway for the request, which A is not from its caller's perspective. A should decide its own availability story: shed with 503 if it genuinely cannot serve without B, degrade to a partial success if the contract permits, or return 500 if the failure is a defect in A. Passing B's code upward leaks A's internal topology and makes the caller's retry decision wrong.
  • Should deliberate load shedding count against a service's availability SLO?
    Usually yes. Shed requests are failed requests from the user's perspective, and excluding them creates an incentive to hide capacity shortfalls behind a shedding threshold. The nuance is that shedding is still preferable to collapse, so the SLO should count it while dashboards separate shed 503s from defect 500s, keeping the remediation — add capacity versus fix a bug — distinguishable.

An airline that distinguishes "flight cancelled, rebooked for 18:00" from "aircraft caught fire" can staff and message the two very differently; one that announces both as "technical issue" learns nothing from its own data.

saying these in an interview costs you the question

  • Letting each team choose its own 5xx conventions and then alerting on an aggregate 5xx rate.
  • Returning 500 for overload, which destroys 500 as a defect signal.
  • Emitting 502 or 504 from application code that is not acting as a gateway.
  • Relying on per-request retry caps with no fleet-wide retry budget, so retries amplify outages.
  • Preferring a slow 504 under saturation over an explicit fast 503, on the grounds that "the request might still succeed".

context