skip to content

A downstream HTTP service returns a 429 Too Many Requests response with a Retry-After: 30 header. How should a well-behaved client's retry logic use this header, and what should it do if the header is absent on a 429 or 503 response?

level: seniorimportance: should knowfreq 55%

answer

  1. Retry-After: seconds or HTTP-date
  2. 429 = rate limit reset; 503 = maintenance/overload estimate
  3. honor server value, don't override with shorter backoff
  4. still add jitter on top of Retry-After
  5. fallback to normal backoff when header absent

basics

~20 s

The Retry-After header tells the client exactly how long to wait, so the client should honor that value instead of using its own backoff timer. If the header isn't present, the client should fall back to its normal exponential-backoff-with-jitter schedule.

solid answer

~50 s

Retry-After is the server explicitly communicating its own recovery estimate or rate-limit reset time, and a well-behaved client should treat it as authoritative and wait at least that long before retrying — ignoring it and retrying sooner defeats the purpose of a server signaling backpressure, and can itself violate the API contract. The header can be either a number of seconds or an HTTP-date, so the client needs to parse both formats. Jitter can still be added on top of the server's value (e.g., wait Retry-After + a small random amount) to avoid many clients that all received the same Retry-After value retrying in lockstep. When the header is absent — which is common, since many services don't set it — the client falls back to its own exponential-backoff-with-jitter policy as the default. Some clients also cap how long they're willing to honor an extreme Retry-After value against their own deadline budget, rather than blindly waiting an arbitrarily long server-specified duration.

go deeper

for a junior

Should know that a 429/503 response can come with a hint about how long to wait, and that the client should generally respect it.

for a middle

Should know both Retry-After formats (seconds and HTTP-date) exist and describe the basic fallback to exponential backoff when it's missing.

for a senior

Should explain why honoring Retry-After takes precedence over client-computed backoff, that jitter can still layer on top of it, and connect it to retry-eligibility classification by status code.

for a principal

Should discuss defensive handling of adversarial/extreme Retry-After values against deadline budgets, and reference real provider behavior (e.g., stricter throttling for clients that ignore the header) as a systemic design consideration.

## What the header adds to blind retry logic Most retry logic operates blind: the client has no information from the server about why a request failed or how long a wait is likely to help, so it falls back to a generic policy — exponential backoff with jitter — calibrated to be reasonable across many possible causes of failure. The `Retry-After` HTTP response header, defined for responses like `429 Too Many Requests` and `503 Service Unavailable`, changes this by letting the server communicate an explicit, situation-specific recovery estimate directly to the client. - **For a 429** (rate limiting), the value usually reflects when the client's rate-limit window resets — the server knows exactly when it will start accepting requests from this client again, and it's telling the client rather than making it guess. - **For a 503** (service temporarily unavailable, e.g., during a deployment or planned maintenance window), the value reflects the server's own estimate of when it expects to be healthy again. The header's value can be expressed in one of two formats: an integer number of seconds to wait (`Retry-After: 30`), or an absolute HTTP-date (`Retry-After: Wed, 21 Oct 2026 07:28:00 GMT`) — a client parsing this header needs to handle both, since different servers and different response codes use each format. ## Treat the server's value as authoritative The correct client behavior is to treat the header as an authoritative lower bound on the wait time and honor it rather than override it with the client's own shorter backoff schedule. Retrying sooner than instructed doesn't just risk another 429/503 — for rate-limited APIs, some providers explicitly penalize clients that repeatedly ignore Retry-After by extending the rate-limit window or applying stricter throttling, effectively treating early retries as a signal of a non-compliant client. This makes Retry-After meaningfully different from ordinary exponential backoff: exponential backoff is the client's own best guess in the absence of server input, while Retry-After is a concrete signal from the party that actually knows the answer, and a well-designed client architecture treats server-provided timing as taking precedence over client-computed timing whenever both are available. ## Jitter and sanity checks still apply That said, honoring Retry-After doesn't mean abandoning jitter or sanity checks entirely. - If many clients hit the same rate limit at the same moment and all receive an identical Retry-After value, honoring it literally recreates the synchronized-retry problem jitter is meant to solve — so a robust client adds a small amount of jitter on top of the server-specified wait (e.g., `Retry-After + random(0, few seconds)`) rather than treating the header's value as a single precise instant to retry at. - There's also a defensive concern: a misbehaving or compromised server could return an extreme Retry-After value (hours or days), and a client that blindly waits that long risks violating its own caller's deadline or effectively hanging; well-behaved clients typically cap how long they're willing to wait based on the header, falling back to surfacing an error to the caller if the server-specified wait exceeds some sane ceiling or the caller's own deadline budget. ## When the header is absent When the header is absent — which is common, since not every API implements it even for 429/503 responses — the client has no choice but to fall back to its default exponential-backoff-with-jitter policy, since it has no server-provided signal to rely on. A robust retry implementation therefore checks for the header's presence first and only falls through to computed backoff when it's missing. It's also worth distinguishing Retry-After from status-code-based retry eligibility more broadly: a 429 or 503 is generally considered retryable, while a `400 Bad Request` or `404 Not Found` is not — retrying a malformed request or a genuinely nonexistent resource will simply fail again regardless of how long the client waits, so Retry-After handling and retry-eligibility classification (which status codes are worth retrying at all) are two separate but related pieces of a complete retry policy. ## Where it shows up A concrete real-world example is how most major cloud/API providers implement rate limiting: GitHub's REST API, for instance, documents returning a 403 or 429 with rate-limit headers when a client exceeds its request quota, and its own documentation instructs well-behaved clients to check these headers and back off accordingly rather than continuing to poll at the same rate — clients that ignore this guidance and keep hammering the API at their prior rate risk having their access token further throttled or temporarily blocked. The general pattern — server communicates its own recovery timing explicitly, client treats that as authoritative over its own guesswork — recurs across rate-limited public APIs, message brokers signaling backpressure, and load balancers returning 503 during rolling deployments.

  • Should a client ever retry sooner than a server-specified Retry-After value?
    Generally no — doing so works against the server's explicit signal and can trigger stricter throttling or be treated as abusive client behavior by APIs that track compliance with Retry-After. The one legitimate exception is when the specified wait would blow through the client's own hard deadline; in that case the sane response is to give up and surface an error to the caller, not to retry early.
  • How should a client's retry-eligibility logic treat a 429 versus a 400 response differently?
    A 429 signals a temporary, self-inflicted condition (rate limit) that is expected to clear, so it's retryable, ideally honoring any Retry-After header. A 400 signals the request itself was malformed in a way the server rejected outright, which retrying won't fix since the same malformed request will be rejected again — so 400s (and most other 4xx codes except 429 and sometimes 408) are generally excluded from a retry policy's list of retryable status codes.
  • If a client receives Retry-After: 30 on a 429 but its own internal exponential backoff would have suggested only a 2-second wait, which should it use?
    It should honor the server's 30-second value, since the server has direct knowledge of its own rate-limit reset time while the client's backoff schedule is just a generic estimate made without that information. Using the shorter client-computed delay would likely just produce another 429 and possibly escalate throttling.

Like a restaurant host who tells you 'come back in 20 minutes' instead of you guessing and hovering at the door every 2 minutes — ignoring the host's explicit estimate and showing up early anyway just annoys the host and doesn't get you seated any faster.

saying these in an interview costs you the question

  • Doesn't know Retry-After can be either seconds or an HTTP-date
  • Suggests ignoring Retry-After in favor of the client's own backoff schedule
  • Treats every 4xx status code as equally retryable
  • No fallback behavior described for when the header is absent
  • Doesn't consider capping an extreme Retry-After value against the caller's own deadline

context