In a Go rate-limiting HTTP middleware, what is the difference between an x/time/rate Limiter's Allow and Wait?
answer
- one answers, the other postpones
- does the caller wait, or get told no?
- shedding versus queueing
- a parked request still owns a goroutine
- one returns a bool, one takes a context
basics
~20 sAllow returns true or false immediately, so an over-limit request is shed with 429 straight away. Wait blocks the request's goroutine until a token frees up or the request context ends, queueing the caller instead of refusing it.
solid answer
~50 sThey are the two ways a middleware can answer an over-limit request. `Allow()` is non-blocking: it returns a bool, and if it is false the middleware writes `429 Too Many Requests` and returns without calling the next handler. `Wait(ctx)` blocks the goroutine serving that request until a token is available, returning an error if the context is canceled or its deadline would pass first; the caller sees latency rather than an error. In an HTTP server that difference is load shedding versus queueing. Shedding keeps latency honest and bounded, and pushes back-pressure onto the client. Queueing smooths a short burst but costs a live goroutine, its stack, the accepted connection and any downstream resources for every waiter, so under sustained overload it converts a clean error into growing memory and timeouts. Public request paths usually shed; queueing is for short, internally bounded callers.
code
go · 12 lines// shed: answer now, never block the caller
if !lim.Allow() {
w.Header().Set("Retry-After", "1")
http.Error(w, "rate limited", http.StatusTooManyRequests)
return
}
// queue: hold the request until a token frees or the caller gives up
if err := lim.Wait(r.Context()); err != nil {
return // context canceled, or the wait outlasts its deadline
}
next.ServeHTTP(w, r)go deeper
Be ready to say which of the two returns a bool and which blocks, and to write the four-line rejection branch: header, status, return. Knowing that the rejected request must not reach the next handler is the part interviewers check.
Explain what each choice costs the server: a shed request costs a comparison, a queued one costs a goroutine, a connection and memory for as long as it waits. Be able to say which context you pass to the blocking call and what its error means.
Show the operating judgment: default to shedding on a public path, queue only for bounded internal callers, and bound the waiters when you do. Expect to describe what you would watch to know the choice is wrong — goroutine count and tail latency climbing while throughput does not.
Own the posture as policy rather than a per-handler choice: which callers may ever be queued, what deadline a queued request gets, and how the rejection is documented so client teams implement back-off instead of a hot retry loop against your 429s.
## The two methods A rate limiter in the `golang.org/x/time/rate` package hands out permission to proceed. A middleware asks it once per request, and it has two ways to ask. `Allow()` asks and gets an answer immediately. It returns a `bool`: true means the request may proceed, false means it may not. It never blocks and never sleeps. If it returns false, the middleware's job is to answer the request itself and stop: set any headers first, write the status, and return without invoking the next handler in the chain. `Wait(ctx)` asks and is prepared to be told "not yet". It parks the calling goroutine until permission becomes available, and returns `error`. It returns a non-nil error when the context passed to it is canceled or its deadline would elapse before permission could be granted. On a nil return, permission has been granted and the request proceeds. ## What that means for an HTTP server `net/http` serves every accepted request on its own goroutine. So the choice is not a stylistic one about which API reads better; it decides what the server does with an excess request. With `Allow()`, an excess request is **shed**. The server spends almost nothing on it: a comparison, a status line, a closed request. The client learns instantly that it is over its limit and can back off. Latency for the requests that are admitted stays whatever the service normally does, because rejected work is not competing with them. With `Wait(ctx)`, an excess request is **queued**. The queue is not a data structure you can see or size; it is the set of goroutines parked inside `Wait`. Each one holds a goroutine and its stack, the accepted TCP connection and its read and write buffers, whatever the middleware already allocated for the request, and a slot in any upstream connection pool or proxy that is waiting on the response. If the arrival rate stays above the limit, that set grows without bound: memory climbs, the tail latency of admitted requests grows because the machine is doing more of everything, and eventually clients time out anyway, having spent your resources to learn nothing. Queueing pays off only when the overload is a short burst that the bucket's spare capacity can absorb, so the queue drains as fast as it fills. ## Using Wait safely If you queue, pass the request's own context, `r.Context()`, so a client that disconnects stops occupying a slot: `net/http` cancels that context when the connection goes away. Give the context a deadline, so a waiter cannot wait longer than the answer is useful for; without one, a waiting request has no natural end. Treat the error from `Wait` as terminal for the request: log it if you like, but do not call the next handler after it. And bound the number of concurrent waiters separately if the limit is per-client, because a thousand clients each entitled to wait is still a thousand goroutines. ## Writing the rejection When the middleware sheds, the response mechanics matter. Headers must be set on `w.Header()` **before** anything writes the status line; once the status is written, later header mutations are ignored, and a second write of the status logs a "superfluous WriteHeader call". `http.Error(w, msg, http.StatusTooManyRequests)` writes the status and a plain-text body in one call, so set `Retry-After` (or any diagnostic header) above it. Then `return` — falling through to the next handler after writing a 429 produces a response that is both rate-limited and served. ## Choosing A useful default for a public API is: shed. A 429 is a truthful, cheap, actionable answer, and it keeps your own resource usage proportional to the traffic you actually intend to serve rather than to the traffic that arrives. Queue where the caller set is known and bounded — an internal batch job, a fan-out you wrote yourself — and where a little added latency is genuinely better than an error the caller would only convert into a retry. Some services do both: `Wait` with a short deadline, falling back to a 429 when the deadline would be exceeded, which is really just shedding with a small absorbing buffer in front of it.
- If the middleware uses Wait, which context should it pass and why?Pass `r.Context()`, which `net/http` cancels when the client disconnects, so a waiter that nobody is listening for stops holding resources. Give it a deadline as well, so a request cannot wait longer than the response is useful; `Wait` returns an error rather than parking past a deadline it cannot meet. Treat that error as terminal and do not call the next handler.
- What must the middleware get right about the response when it sheds?Set every header on `w.Header()` before the status is written, then write the status once — `http.Error(w, "rate limited", http.StatusTooManyRequests)` does the status and body together. Then return, so the inner handler never runs. Writing the status and continuing produces a double response and a superfluous WriteHeader warning in the server's error log.
- Why can queueing turn a small overload into an outage?Each waiting request keeps a goroutine, its stack, the accepted connection, and anything the middleware allocated for it. Nothing bounds that set except the arrival rate, so a sustained excess grows memory and the tail latency of admitted requests, and the waiters eventually time out anyway. Shedding costs one comparison and one short response per excess request.
Allow is a doorman who says yes or no and closes the door. Wait is a doorman who lets you stand in the queue outside — cheap for him, but every person in that queue is still yours to feed.
saying these in an interview costs you the question
- Thinks Allow blocks until a token becomes available
- Thinks Wait discards the request when the limit is hit
- Queues every over-limit request with no deadline or cap
- Calls the next handler after writing the 429
- Sets a response header after the status has been written
- Assumes waiting requests are free because they are idle