skip to content

What does the HTTP Retry-After header mean on a 503 Service Unavailable response, what value formats may it take, and how should a well-behaved client honour it?

level: middleimportance: should knowfreq 50%

answer

  1. Seconds delta or HTTP-date
  2. Delta form avoids clock skew
  3. Floor, not exact — add jitter
  4. Cap the value, bound attempts
  5. Also legal on 429 and 301

basics

~20 s

Retry-After tells the client how long to wait before retrying. Its value is either a number of seconds (delta) or an HTTP-date. On a 503 it is a hint about when service resumes; clients should wait at least that long and add jitter.

solid answer

~50 s

`Retry-After` accompanies a 503 to convert "unavailable" into an actionable instruction. Two value forms are allowed: a **delta in seconds** (`Retry-After: 120`) or an absolute **HTTP-date** (`Retry-After: Tue, 12 Aug 2026 10:20:00 GMT`). The delta form is safer in practice because it needs no clock agreement between the two sides. A well-behaved client treats it as a floor, not an exact schedule: wait at least the stated interval, then add **randomised jitter** so that thousands of shedding clients do not return in a synchronised thundering herd. It should also cap absurd values — a `Retry-After: 86400` should not park a request handler for a day — and keep its own retry budget and attempt limit. On the server side, emit it whenever you know something: the drain window, the maintenance end, the circuit-breaker cooldown. Emitting a value you do not actually mean is worse than omitting the header.

code

http · 5 lines
http
HTTP/1.1 503 Service Unavailable
Retry-After: 120

HTTP/1.1 503 Service Unavailable
Retry-After: Tue, 12 Aug 2026 10:20:00 GMT

go deeper

for a junior

Know that Retry-After appears on 503, that it means "wait this long", and that both a seconds count and a date are valid.

for a middle

Explain both value forms and why delta-seconds avoids clock-skew problems, and describe a client that waits, jitters, caps and bounds attempts.

for a senior

Discuss where the server gets a real value from — drain time, breaker cooldown, queue depth — and how honouring the header belongs in the shared HTTP client rather than each caller.

for a principal

Position it inside a fleet-wide congestion-control story: retry budgets, jitter, load shedding, and how client behaviour is standardised so a shedding service can actually recover.

## The problem the header solves A bare 503 tells a client the server is temporarily unavailable but says nothing about *temporarily how long*. Left to guess, clients pick their own retry loop — usually far too aggressive — and a server that is already overloaded gets hit again immediately by every client at once. `Retry-After` exists so the server, which is the only party that knows why it is unavailable, can control the return schedule. ## The two value forms The field accepts either of two syntaxes: - **delta-seconds** — a non-negative integer count of seconds, e.g. `Retry-After: 30`. Relative to receipt of the response. - **HTTP-date** — an absolute timestamp in the IMF-fixdate format, e.g. `Retry-After: Tue, 12 Aug 2026 10:20:00 GMT`. The delta form dominates in practice for one reason: it requires no agreement about clocks. A client with a skewed clock reading an absolute date can compute a negative or wildly long wait. If you must send a date — for a maintenance window with a genuinely fixed end — send it in GMT in the exact HTTP-date format, and expect some clients to parse it poorly. ## What a correct client does with it 1. **Treat it as a minimum.** The server is saying "not before this". Retrying earlier is a protocol-hostile move against a machine that is already struggling. 2. **Add jitter.** If ten thousand clients are told to come back in 30 seconds and all obey exactly, the server gets a spike at t+30 identical to the one that caused the shed. Randomise within a window — for example wait the stated delay plus a uniform random fraction of it. 3. **Cap it.** Guard against nonsense or hostile values. A synchronous request path cannot honour an hour-long delay; either fail the call and let the caller decide, or hand the work to a queue. 4. **Bound attempts.** `Retry-After` is not permission to retry forever. Keep a maximum attempt count and an overall deadline, and respect the caller's own timeout budget. 5. **Respect idempotency.** Retrying is only safe when the operation is idempotent or protected by an idempotency mechanism. `Retry-After` speaks to *when*, never to *whether it is safe*. ## What a correct server does Emit `Retry-After` only when you know something real. Good sources of a value: the remaining drain time during a graceful shutdown, the scheduled end of a maintenance window, the cooldown left on a circuit breaker, or a computed estimate from queue depth divided by drain rate. If you have no basis for a number, sending an invented one trains clients to distrust the header; omitting it is honest. Also think about who reads it. A CDN or reverse proxy in front of you may act on the header, and monitoring dashboards often chart it. And be careful about interaction with caching: a 503 is not cacheable by default, and you generally do not want an intermediary storing your unavailability. ## Beyond 503 The header is defined generically rather than as a 503-only field: it is also used on 429 Too Many Requests to communicate a rate-limit reset, and on 301 Moved Permanently to hint how long the resource will remain at the new location before the client should re-check. Within server errors, 503 is the natural home, because it is the one 5xx where the server has an informed opinion about recovery timing. A 500 by definition is an unanticipated defect with no schedule, and a 504 tells you upstream was slow without telling you when it will be fast. ## A common failure mode Teams often add `Retry-After` on the server and then discover that nothing changed, because their clients — SDKs, mobile apps, other services — ignore it entirely and run a fixed retry loop. Honouring `Retry-After` has to be implemented in the shared HTTP client library, not asked of every caller individually. That is the difference between a header that documents intent and one that actually shapes traffic.

  • Why is the delta-seconds form usually preferred over an HTTP-date?
    The delta form is interpreted relative to when the client receives the response, so it needs no agreement between client and server clocks. An absolute date is misread by any client whose clock is skewed, and can produce a negative wait — an immediate retry — or an absurdly long one. Dates are only worth using for a genuinely fixed window such as scheduled maintenance.
  • A client honours Retry-After exactly as given. What can still go wrong at scale?
    Synchronisation. Every client that received the same value returns at the same instant, recreating the load spike that caused the shed in the first place. The fix is randomised jitter around the stated delay, so the returning traffic is spread over a window rather than arriving as a wall.

A sign on a shop door reading "back in 20 minutes": it tells you not to keep rattling the handle, but if every customer returns at exactly minute 20 the queue is as bad as before.

saying these in an interview costs you the question

  • Believing the value can only be a number of seconds and rejecting the HTTP-date form.
  • Treating Retry-After as a promise that retrying is safe, ignoring idempotency.
  • Retrying at exactly the stated instant with no jitter across a fleet of clients.
  • Blocking a request thread for an arbitrarily long stated delay instead of capping it.
  • Emitting an invented value on every 503 when the server has no idea when it will recover.

context