How should an HTTP service signal rate limiting with status 429 Too Many Requests, and what are the two legal formats of the Retry-After header?
answer
- 429 = this client's quota; 503 = server-wide
- Retry-After: delta-seconds OR HTTP-date, nothing else
- delta-seconds beats dates under clock skew
- no-store so shared caches never replay it
- Jitter the retry; RateLimit-* headers pre-empt
basics
~20 sReturn 429 when a client exceeded its request quota, and include Retry-After telling it when to come back. Retry-After takes either delta-seconds (Retry-After: 120) or an HTTP-date (Retry-After: Wed, 21 Oct 2026 07:28:00 GMT). Clients should back off, not retry immediately.
solid answer
~50 s**429 Too Many Requests** (RFC 6585) says the client has sent too many requests in a given time window. The response should explain the limit in the body and, crucially, carry **`Retry-After`**, which has exactly two legal forms: - **delta-seconds**: `Retry-After: 120` — retry after 120 seconds; - **HTTP-date**: `Retry-After: Wed, 21 Oct 2026 07:28:00 GMT` — an absolute instant in IMF-fixdate form. Delta-seconds is safer in practice because it is immune to client clock skew; dates are useful for fixed window resets. Good 429 practice: make the response **`no-store`** so intermediaries never serve a cached rate-limit error; add quota headers (`RateLimit-Limit`, `RateLimit-Remaining`, `RateLimit-Reset`) so clients can self-pace before being rejected; keep 429 cheap to produce so a flood does not cost the same as real work. Distinguish it from **503 Service Unavailable**, which also accepts `Retry-After` but means the *server* is overloaded or down, not that this client exceeded a quota. And expect clients to add jitter — synchronised retries at exactly `Retry-After` re-create the spike.
code
http · 13 linesGET /api/search?q=shoes HTTP/1.1
Host: api.example.com
Authorization: Bearer ...
HTTP/1.1 429 Too Many Requests
Retry-After: 30
RateLimit-Limit: 1000
RateLimit-Remaining: 0
RateLimit-Reset: 30
Cache-Control: no-store
Content-Type: application/json
{"code":"RATE_LIMITED","scope":"per-api-key","limit":1000,"window":"60s"}go deeper
Know that 429 means too many requests and that Retry-After tells the client when to try again, in seconds or as a date.
Add the exact two legal formats, why seconds beat dates, and the 503 contrast; mention RateLimit-* headers.
Talk about jitter and thundering herds, no-store to protect shared caches, cheap rejection at the edge, limiter keys and algorithms, and per-client 429 alerting.
Frame quota policy across the platform: which identity dimension is limited, where enforcement lives, how clients and SDKs are expected to react, and how throttling interacts with circuit breakers and error budgets.
## What 429 means **429 Too Many Requests**, defined in RFC 6585, means the user agent has sent too many requests in a given amount of time — quota exhausted. The subject is *this client* (by API key, token, account, IP, or tenant), not the health of the server. RFC 6585 says the response *should* include an explanation of the condition and *may* include `Retry-After`. Contrast the neighbours: - **503 Service Unavailable** — the *server* is overloaded or in maintenance; every client is affected. Also carries `Retry-After`. - **403 Forbidden** — a permission decision, permanent for this credential. - **408 Request Timeout** — the client was too slow to send its request. Using 503 where a per-client quota was hit tells the client "the service is down", which is false and can trigger circuit breakers that stop *all* traffic instead of just slowing this caller. ## Retry-After: the two legal forms RFC 9110 defines `Retry-After` as either: 1. **delta-seconds** — a non-negative decimal integer number of seconds: `Retry-After: 30`. 2. **HTTP-date** — a date in the preferred IMF-fixdate format, always GMT: `Retry-After: Wed, 21 Oct 2026 07:28:00 GMT`. Only those two. `Retry-After: 30s`, an ISO-8601 timestamp, or a millisecond value are all invalid and will be ignored or mis-parsed by conformant clients and SDKs. **Which to choose?** delta-seconds is relative, so it is immune to client clock skew and to the delay the response spent in transit — the usual default. An HTTP-date is convenient when your limiter uses fixed windows that reset at a known wall-clock instant, and it survives a client that queues the response for a while, but it depends on the client's clock being roughly correct. `Retry-After` is also legal on 503 and on 3xx redirects; it is not exclusive to 429. ## Making 429 useful to clients A bare 429 forces clients into blind exponential backoff. Better responses include: - **Quota headers.** The widely adopted trio `RateLimit-Limit`, `RateLimit-Remaining`, `RateLimit-Reset` (or the older `X-RateLimit-*` spellings) lets a well-behaved client slow down *before* it gets rejected. Send them on successful responses too — that is where they change behaviour. - **A body with a stable code**, e.g. `{"code":"RATE_LIMITED","scope":"per-api-key","limit":1000,"window":"1m"}`, so client code and support can tell *which* limit fired when several exist (per-IP, per-key, per-endpoint). - **`Cache-Control: no-store`.** 429 is not in the heuristically cacheable set, but explicit headers keep CDNs and shared proxies from ever reusing a rate-limit response for another client — a nasty failure mode where one caller's throttle is served to everyone. ## Client-side behaviour that interviewers probe - **Honour `Retry-After` as a floor, then add jitter.** If a thousand clients are told "retry in 30", they all return at t+30 and re-create the spike. Randomised jitter across the window is the fix. - **Cap total attempts and total wait.** Infinite retries turn a throttle into a self-inflicted DDoS. - **Do not retry non-idempotent requests blindly.** A 429 is normally safe to retry because the request was rejected before processing, but only if you can be confident nothing happened — an idempotency mechanism removes the doubt. - **Treat 429 differently from 5xx in circuit breakers.** Tripping a breaker on 429 stops legitimate low-rate traffic; usually you want to throttle locally instead. ## Server-side design notes - **429 must be cheap.** If producing the rejection costs a database lookup and a template render, an abusive client still consumes real capacity. Enforce at the edge (gateway/CDN) where possible. - **Choose the identity key deliberately.** Per-IP limits collapse whole offices behind NAT and are trivially spread across a botnet; per-API-key or per-account limits are the meaningful ones. If you limit per IP behind a proxy, only trust a forwarded client address from a hop count you control — otherwise a spoofed header defeats the limiter. - **Distinguish burst from sustained.** Token bucket allows short bursts with a sustained refill rate; fixed windows produce boundary spikes (a client can send 2× the limit around the boundary); sliding windows smooth that at higher cost. - **Decide what happens to concurrency-based rejection.** Rejecting because too many requests are *in flight* is closer to load shedding — 503 with `Retry-After` is often the more honest answer there. ## Observability Count 429s by client key and by limiter rule, and alert on "a normally-quiet client is suddenly being throttled" — that usually means a client-side retry loop or a bad deploy, and it is far more actionable than a global 429 rate.
- Why is delta-seconds usually preferred over an HTTP-date in Retry-After?Delta-seconds is relative to when the client receives the response, so it is unaffected by client clock skew, timezone bugs, or the response sitting in a queue. An HTTP-date requires the client's clock to be roughly correct; if it is minutes behind, the client retries too early and gets throttled again, or waits far too long.
- A thousand clients all receive Retry-After: 60. What happens at t+60 and how do you avoid it?They retry simultaneously and reproduce the original spike — a thundering herd that immediately re-triggers throttling. Clients should treat Retry-After as a minimum and add randomised jitter across a window; servers can help by varying the advertised delay per client or returning the actual per-client bucket reset rather than one shared value.
- When is 503 with Retry-After a better answer than 429?When the rejection is about server health rather than a client quota — overload shedding, a dependency outage, or planned maintenance. 429 tells a client "you specifically are going too fast", which is misleading during a global incident and can stop clients from applying the broad backoff or failover behaviour an outage warrants.
saying these in an interview costs you the question
- Writing Retry-After: 30s or an ISO-8601 timestamp — only delta-seconds or an HTTP-date are legal
- Using 503 for per-client quota exhaustion (or 429 for a server-wide outage)
- Letting a shared cache or CDN store and replay a 429 to other clients
- Retrying immediately or without jitter, re-creating the spike
- Rate limiting purely by client IP and trusting an unvalidated forwarded-for header