How do you decide which errors in an HTTP API contract are worth retrying and which are permanent, and how do you communicate that distinction to client developers?
answer
- would the identical request succeed later?
- transient: 429/503/504/502 + transport
- permanent: 400/401/403/404/422
- ambiguous: timeout on a write → needs idempotency
- annotate retryable in the spec + at runtime
basics
~20 sRetryable means the same request could succeed later without changing: 429, 503, 504, connection failures. Permanent means the request itself is wrong: 400, 401, 403, 404, 422. Mark retryability explicitly per error code in the contract, and require idempotency before retrying writes.
solid answer
~50 sThe test is: *would this identical request plausibly succeed later with nothing changed?* **Transient** — the failure is about capacity or availability: `429`, `503`, `504`, connection resets and timeouts. Retry with exponential backoff and jitter, honouring `Retry-After`. **Permanent** — the request is wrong: `400`, `401`, `403`, `404`, `409` in most cases, `422`. Retrying wastes both sides' capacity and never succeeds; the client must change something or surface the error. **Ambiguous** — the important third bucket. A timeout or a dropped connection on a write means you don't know whether the server applied it. Retrying is only safe if the operation is idempotent, so the contract must say which operations are safe to repeat. Don't leave clients to infer this from status codes alone. Mark each documented error code `retryable: true/false` in the API contract, so retry behaviour is generated from the spec instead of guessed — and be internally consistent, because one endpoint returning 503 for a permanent misconfiguration teaches clients to retry forever.
code
json · 7 lines{
"type": "https://api.example.com/problems/upstream-timeout",
"title": "Upstream timed out",
"status": 504,
"retryable": true,
"requestId": "01J8YQ2R4K7ZC3M9"
}go deeper
Give the core split — capacity and availability failures are worth retrying, invalid requests are not — with examples on each side.
Add backoff with jitter, why 401 needs a refresh rather than a resend, and why 500 is not automatically retryable.
Lead with the ambiguous bucket and idempotency, and with publishing retryability per error code in the spec and at runtime.
Treat it as controlling the system's own load profile under failure: consistent classification across services, client obligations, and circuit breaking to prevent retry-amplified outages.
## The single question Classification comes down to: **would this exact request, unchanged, plausibly succeed at a later time?** If yes, it is transient and retrying is rational. If no, retrying is pure waste — it consumes the client's budget, adds load to a system that may already be struggling, and delays the moment a human learns something is broken. ## Transient The failure describes the *system's* current state, not the request: - **429 Too Many Requests** — quota, resets by definition. - **503 Service Unavailable** — overloaded, restarting, in maintenance. - **504 Gateway Timeout** and **502 Bad Gateway** — an upstream hop failed or was slow. - **Transport-level failures** — connection refused, connection reset, DNS failure, TLS handshake failure, client-side timeout. These warrant exponential backoff with jitter, honouring `Retry-After` when present, bounded by an attempt count *and* an overall deadline. ## Permanent The failure describes the *request*: - **400 / 422** — malformed or semantically invalid; identical resend gives an identical answer. - **401** — retryable only *after* obtaining a fresh credential, which is a different request, not a retry. Blind resending of the same expired token is a loop. - **403** — a policy decision; time doesn't change it. - **404** — the resource isn't there. (Read-after-write against an eventually consistent store is the notable exception, and one you must document explicitly if it applies.) - **405, 415** — the client used the wrong method or content type. **500** is the awkward one. Formally it says something went wrong server-side, which *might* be transient. But most 500s are deterministic bugs that reproduce on every attempt. A single cautious retry is defensible; aggressive retrying of 500s is a well-known way to amplify an incident, because a bug triggered by particular input gets hammered by every client at once. **409 Conflict** is contextual: an optimistic-concurrency conflict is retryable *after re-reading and rebuilding the request*, which again is not a blind retry. A business-rule conflict ("order already shipped") is permanent. ## The ambiguous bucket This is where real systems get hurt. A timeout or dropped connection on a `POST` leaves the client not knowing whether the server processed it. Retrying may duplicate a payment; not retrying may lose one. Resolution is a contract property, not a client guess: - **Naturally idempotent operations** (`GET`, `PUT`, `DELETE` with the same target state) can be retried freely. - **Non-idempotent operations** need an explicit mechanism the contract defines, so a repeated attempt is recognised as the same logical operation rather than a new one. The contract must state, per operation, whether repeating it is safe. Without that, cautious clients under-retry and lose work, and incautious ones double-charge customers. ## Communicating it Status codes alone under-specify. Concretely: 1. **Annotate every documented error code** with an explicit retryable flag and, where relevant, a suggested strategy. Put it in the machine-readable spec so generated clients can implement it. 2. **Emit the signal at runtime too** — a boolean or hint in the error body, alongside `Retry-After`, so a client can act correctly even on an error code it has never seen. 3. **Be internally consistent.** The most damaging failure here is a service returning 503 for a permanent condition (a missing configuration value, an upstream that will never come back). Clients retry forever against something that can never succeed. Equally bad is returning 400 for a transient dependency failure, which makes clients give up on work that would have succeeded. 4. **Document the client obligations**: backoff with jitter, an overall deadline, and a circuit breaker — not just an attempt count. ## Why classification is a server-side responsibility Only the server knows whether a failure was capacity or logic. If you leave clients to infer it, they infer differently, and your traffic during an incident becomes a function of a dozen third-party retry implementations. Publishing the classification is how you keep control of your own load profile under failure.
- Why is retrying HTTP 500 responses aggressively risky?Most 500s are deterministic bugs that reproduce on every attempt, so retries add load without any chance of success. Worse, when a bug is triggered by a common input, every client retrying simultaneously multiplies traffic against an already-failing path and can turn a partial failure into an outage. At most take one cautious retry, with backoff, a circuit breaker, and alerting so a human sees it.
- A POST times out and the client never sees a response. How should it decide whether to retry?The outcome is genuinely unknown, so safety must come from the contract rather than the client's judgement. If the operation is defined as safe to repeat — because it is naturally idempotent or because the contract provides a way to identify a repeated attempt as the same logical operation — the client retries. Otherwise it must not blindly retry; it should surface the ambiguity or reconcile by querying the resulting resource state.
- Is HTTP 401 retryable?Not as a blind resend. The same expired or invalid credential will be rejected identically, producing a loop. It becomes actionable only after obtaining a fresh credential, which makes the follow-up a different request rather than a retry — and that refresh attempt itself needs a limit, because a permanently invalid client secret will otherwise refresh forever.
saying these in an interview costs you the question
- Classifying purely by status class — retry all 5xx, never retry 4xx — without considering 500 bugs or 401 refresh
- Blindly retrying non-idempotent writes after a timeout, risking duplicate side effects
- Returning 503 for a permanent condition such as missing configuration, causing clients to retry forever
- Bounding retries only by attempt count with no overall deadline or circuit breaker
- Leaving retryability implicit and expecting every client team to infer it correctly