Your service is overloaded and starts shedding traffic. How do you decide between HTTP status 429, 503 and 500 for those responses, and what do the codes tell a client to do?
answer
- 429 = this caller too fast; Retry-After
- 503 = service down/shedding; must be cheap to reject
- 500 = bug, don't retry, page someone
- 502/504 = upstream problems
- Backoff + jitter + circuit breaker; retries need idempotency
basics
~20 s429 means this caller exceeded its rate or quota — retry after backing off. 503 means the service as a whole is unavailable or overloaded — retry later, ideally per Retry-After. 500 means an unexpected bug; retrying probably will not help and it should page someone.
solid answer
~50 sThe codes encode *whose problem it is and whether to retry*. - **429 Too Many Requests** — the caller exceeded a limit that applies to it. Include `Retry-After` and, ideally, `RateLimit` headers so clients can pace themselves. Fully retryable after waiting. - **503 Service Unavailable** — the service cannot handle the request right now: dependency down, load shedding, queue full, shutting down, in maintenance. Include `Retry-After`. Retryable, but the fleet is already hurting, so clients must use exponential backoff with jitter and circuit breakers or they will make it worse. - **500 Internal Server Error** — an unhandled defect. Not attributable to the client, not expected to resolve on retry, and the one that should wake someone. The operational reason to keep them apart: 429 and 503 are *expected* control signals during load management; a 500 spike is a bug. Reporting overload as 500 destroys that signal, and reporting bugs as 503 hides them.
go deeper
Know 429 = you sent too many requests, 503 = service temporarily unavailable, 500 = something broke on the server.
Add Retry-After, that 4xx is the client's to fix, and that 500 usually should not be retried.
Discuss cheap load shedding, retry amplification, backoff with jitter, circuit breakers, and keeping 5xx alerting meaningful.
Frame it as a system-wide reliability contract — retry budgets, quota policy per tenant, degradation strategy and what the platform guarantees clients on overload.
## Codes as instructions to the client Every failure status implicitly answers two questions: is this the client's doing, and will retrying help? Under load those answers matter more than usual because clients multiply whatever you tell them by their retry policy. ## 429 Too Many Requests "You, specifically, sent too much." The limit may be per API key, per user, per IP, per endpoint. The response should carry: ``` HTTP/1.1 429 Too Many Requests Retry-After: 30 RateLimit: limit=1000, remaining=0, reset=30 Content-Type: application/problem+json {"code":"RATE_LIMITED","scope":"api-key","limitPerMinute":1000} ``` `Retry-After` (seconds or an HTTP date) is what makes 429 cooperative — without it clients guess, and guessing badly is how a limiter turns into a thundering herd at the top of every minute. Publishing remaining-quota headers on *successful* responses too lets well-behaved clients slow down before hitting the wall. Be explicit in the body about *which* limit fired (per-key, per-endpoint, burst versus sustained), otherwise the integrator cannot fix anything. ## 503 Service Unavailable "Not you — us, right now." Legitimate triggers: - a critical dependency is down and the request cannot be served meaningfully, - deliberate load shedding when a queue or thread pool is saturated, - rolling deploy / graceful shutdown draining connections, - planned maintenance. 503 also carries `Retry-After`. The important design point is that 503 must be **cheap**. Load shedding only helps if rejecting costs far less than serving — reject at the edge, before the expensive work, ideally before acquiring a database connection. A 503 produced after a 30-second dependency timeout has already consumed the resource you were trying to protect; that is what turns overload into collapse. Related: `504 Gateway Timeout` (an upstream did not answer in time) and `502 Bad Gateway` (an upstream sent something invalid) are the proxy-flavoured cousins. Both are effectively "try again, it is not your fault", and both should be distinguishable in your dashboards from application 500s. ## 500 Internal Server Error The catch-all for defects: an unhandled exception, a null dereference, a serialization failure. Properties that matter: - **Not the client's fault** and not fixable by changing the request. If the client *can* fix it, the status was wrong — that request should have been a 4xx. - **Not usefully retryable.** The same input generally reproduces it. - **Alertable.** A 500 rate above baseline is a bug, and that should be the loudest signal in the dashboard. So never use 500 for load shedding, quota, validation, or a known-down dependency. Each of those has a code that tells the client something true. The body must stay opaque: a stable error code and a correlation/trace id, never a stack trace, SQL fragment, or internal hostname — 5xx bodies are attacker-visible and stack traces are a reconnaissance gift. ## Retry behaviour and the feedback loop The client contract should read roughly: | Status | Retry? | How | |---|---|---| | 429 | yes | honour `Retry-After`, then exponential backoff + jitter | | 503 / 504 | yes | exponential backoff + jitter, circuit breaker after N failures | | 500 | no (or once, cautiously) | surface to the caller, log with the correlation id | | 4xx (other) | no | fix the request | Two failure modes to name in an interview: **retry amplification** — three layers each retrying three times is 27 requests to a service already failing — and **synchronized retries**, which is why jitter is not optional. Budgeted retries (a cap on the fraction of traffic that may be retries) are the modern answer, along with circuit breakers that stop sending entirely for a while. ## Idempotency is the precondition for any of this Telling clients to retry is only safe if retries are safe. GET/PUT/DELETE are idempotent by definition; POST is not, so retry advice on POST requires an idempotency-key mechanism. An API that returns 503 on a payment POST and expects clients to retry, without deduplication, is asking for double charges. ## What good operations look like Separate dashboards for 4xx (client behaviour), 429 (limits firing), and 5xx (your health). Alert on 5xx rate and on 503 volume; treat sustained 429 as a capacity or customer-integration conversation rather than an outage. And make sure health checks and load balancers understand your 503s — a node returning 503 during graceful shutdown should be pulled from rotation, not restarted in a loop.
- Which headers should accompany a 429 response?Retry-After tells the client how long to wait, in seconds or as an HTTP date, and is what prevents guess-and-hammer behaviour. RateLimit-style headers (limit, remaining, reset) let clients pace themselves before they hit the wall, and are most useful when also sent on successful responses. The body should name which limit fired so an integrator can act on it.
- Why is returning 500 for an overloaded service harmful?It merges an expected control signal with a defect signal, so your 500 alert stops meaning "there is a bug" and on-call loses the ability to triage quickly. It also misinforms clients: 500 says retrying will not help, while an overloaded service actually wants clients to back off and retry later, which is what 503 with Retry-After expresses.
saying these in an interview costs you the question
- Returning 500 for rate limiting or for a known-down dependency
- Sending 429 or 503 with no Retry-After
- Retrying immediately without exponential backoff or jitter
- Producing 503 only after the expensive work has already been done, so shedding does not relieve load
- Including stack traces or internal hostnames in 5xx response bodies