skip to content

What does the HTTP Retry-After response header mean, what value formats does it accept, and on which responses should an API send it?

level: juniorimportance: must knowfreq 58%

answer

  1. seconds or HTTP-date
  2. 429 + 503 (also 3xx redirects)
  3. prefer seconds: clock skew
  4. floor, not a promise — add jitter
  5. clamp on the client; omit if unknown

basics

~20 s

Retry-After tells the client how long to wait before trying again. It takes either delay seconds (Retry-After: 120) or an HTTP-date. Send it on 429 and 503, and on 3xx redirects that ask the client to wait.

solid answer

~50 s

`Retry-After` is a response header carrying the server's instruction about *when* a retry is worth making. Two formats are defined: a **non-negative integer of seconds** (`Retry-After: 120`) or an **HTTP-date** (`Retry-After: Wed, 12 Aug 2026 14:00:00 GMT`). Delay-seconds is the safer choice in practice because it is immune to client clock skew and needs no date parsing. It belongs on: - **429 Too Many Requests** — when the rate-limit window resets. - **503 Service Unavailable** — how long the outage or maintenance is expected to last. - **3xx redirects** where the server wants the client to delay before following. It is an instruction to *wait at least* that long, not a promise of success at that instant. Clients should honour it as a floor and add jitter on top, because everyone told "retry in 60" retries at the same second and rebuilds the thundering herd the limit was protecting against. If you cannot estimate the wait honestly, sending no header is better than an invented one.

code

http · 5 lines
http
HTTP/1.1 429 Too Many Requests
Content-Type: application/problem+json
Retry-After: 42

{"type":"https://api.example.com/problems/rate-limited","title":"Rate limit exceeded","status":429}

go deeper

for a junior

State both formats and that the header belongs on 429 and 503, and that the client waits at least that long.

for a middle

Add why delay-seconds avoids clock skew, and that the value is a floor requiring client-side jitter.

for a senior

Cover server-side honesty during incidents, client-side clamping of hostile values, and fallback backoff when the header is absent.

for a principal

Discuss it as one input to a system-wide load-shedding policy: what the server advertises, what clients are required to do, and how the two are kept consistent across the estate.

## What the header says `Retry-After` is defined by the HTTP specification as a way for a server to tell a client how long to wait before issuing a follow-up request. It converts a bare "no" into an actionable "no, try again after N" — the difference between a client guessing and a client cooperating. ## The two formats **Delay-seconds**: a non-negative integer, relative to when the response was received. ``` Retry-After: 120 ``` **HTTP-date**: an absolute timestamp in the standard HTTP date format. ``` Retry-After: Wed, 12 Aug 2026 14:00:00 GMT ``` Both are legal, so a client must handle either. In practice **prefer delay-seconds** when generating: - It is immune to **client clock skew** — a client whose clock is minutes off will compute a nonsense wait from an absolute date, potentially zero or negative. - It requires no date parsing, historically a source of bugs. - It is trivially correct behind caches and proxies that add transit delay only in the client's favour. Absolute dates make sense when the reset is a genuinely fixed wall-clock moment, such as a scheduled maintenance window ending, and you want that stated exactly. ## Where it belongs - **429 Too Many Requests.** The natural value is the time until the client's rate-limit window resets. This is the single most useful place for the header: without it the client picks an arbitrary backoff, which is either too aggressive (it keeps getting 429s) or too conservative (it stalls unnecessarily). - **503 Service Unavailable.** The expected duration of the outage or maintenance. A short value tells a client to hold on; a long value tells it to fail the operation and requeue. - **3xx redirects.** The specification also allows it here, to ask the client to delay before following the redirect. Rare in practice. It is **meaningless on most other errors**. On a `400` or `422` the request is wrong and time changes nothing; sending a delay there invites pointless retries of a request that will never succeed. ## What it does not promise It is a *lower bound on the wait*, not a guarantee that a request at that moment will succeed. If load is still high, the next attempt may return another 429 with a new `Retry-After` — and the client must honour the new value rather than assuming the first one was final. ## The jitter problem If a server returns `Retry-After: 60` to ten thousand clients at once, all ten thousand retry within the same second. The header has synchronised them into a coordinated burst — precisely the pattern the rate limiter existed to prevent. The correct client behaviour is to treat the value as a **floor** and add randomised jitter above it. A well-behaved server can help by varying the value slightly per client. ## Server-side honesty Only send a value you can justify. An invented `Retry-After: 30` on an outage that lasts an hour produces a client retrying twice a minute for an hour — load you did not want during an incident. If the duration is unknown, either omit the header (the client falls back to its own exponential backoff) or send a deliberately conservative value. ## Client-side sanity checks A client should clamp the value. A hostile or buggy server can send `Retry-After: 86400`, and a client that blindly sleeps a day has effectively been shut down by a header. Cap it at something appropriate for the operation, and treat an unreasonably large value as "fail now and let a human or a scheduler decide". Equally, ignore malformed values rather than crashing, and never treat a missing header as permission to retry immediately.

  • Why is delay-seconds usually preferable to an HTTP-date?
    An absolute date depends on the client's clock being correct; a skewed clock can compute a wait of zero or a negative value and retry immediately, or wait far too long. Delay-seconds is measured from receipt of the response, so it is unaffected by clock differences and needs no date parsing. Use an HTTP-date only when the reset really is a fixed wall-clock moment worth stating precisely.
  • Should a client retry immediately if a 503 arrives with no Retry-After header?
    No. The absence of the header means the server has no estimate, not that no wait is needed. The client should fall back to its own exponential backoff with jitter, bounded by an attempt limit or deadline. Retrying immediately on a 503 is the classic way a client turns a degraded service into a fully collapsed one.
  • Does Retry-After guarantee the request will succeed after that delay?
    No. It is a lower bound on how long to wait, not a promise. If load is still elevated the next attempt can return another 429 or 503 with a fresh value, and the client must honour the new value rather than treating the first as final. Clients should still enforce their own overall attempt or deadline budget.

saying these in an interview costs you the question

  • Thinking Retry-After only accepts seconds and failing to parse an HTTP-date
  • Treating the value as a guarantee that the retry will succeed
  • Retrying at exactly the stated time with no jitter, re-synchronising all clients into one burst
  • Sending Retry-After on 400 or 422, where waiting cannot help
  • Blindly sleeping for whatever value the server sends, with no client-side clamp

context