In golang.org/x/time/rate, what is the difference between Limiter.Allow and Limiter.Wait?
answer
- one asks, one queues
- which one can return false?
- the blocking form takes a context
- shed the work, or pace the work
- blocking with no deadline is an unbounded queue
basics
~20 sAllow never blocks: it takes a token if one is free and returns true, otherwise false, so the caller must shed or retry. Wait blocks until a token is free, the context is cancelled, or its deadline passes.
solid answer
~40 s`Allow` is the non-blocking probe. It consumes one token and returns true, or returns false immediately when the bucket is empty, consuming nothing. Reach for it when you would rather drop or degrade the work than delay it. `Wait(ctx)` blocks until a token is granted and returns nil, and returns a non-nil error if the context is cancelled, if its deadline would pass before a token could arrive, or if it asks for more tokens than the limiter's burst. Wait suits a client that should be paced rather than shed — a background exporter calling a quota-limited API — but only with a context carrying a deadline, otherwise overload becomes a growing queue of blocked goroutines instead of a clean failure. `Reserve` never blocks and reports how long the wait would be.
go deeper
Be ready to say which of the two blocks and which does not, and to give one concrete situation for each: Allow when the work can be dropped, Wait when the caller should simply be slowed down.
Explain that both draw on the same token bucket, that Wait takes a context and returns its error on cancellation or an impending deadline, and that Reserve is the non-blocking third form that reports the delay.
Show that Wait turns overload into latency and goroutine growth, so a deadline is mandatory, and that you keep allowed-versus-denied counters so you can tell whether the limit is the binding constraint.
Own the policy: whether exceeding the limit queues callers or fails them fast is a product and cost decision, and it has to be stated once and applied the same way by every client sharing the quota.
## What the limiter is `golang.org/x/time/rate` (one of the `golang.org/x` modules, not the standard library) provides a **token bucket**. A bucket holds at most `burst` tokens and refills continuously at `rate` tokens per second. Every operation you want to pace asks the bucket for one token; if a token is there, the operation proceeds and the token is gone. The bucket is what makes short spikes acceptable while the long-run average stays at the configured rate. One `*rate.Limiter` is safe for concurrent use by many goroutines, so a single limiter guarding an outbound client is the normal shape: build it once, share the pointer, and have every call site ask it for permission. ## Allow — ask and move on ``` if !limiter.Allow() { // no token now; skip, degrade, or return a busy error } ``` `Allow` returns a plain `bool` and never sleeps. True means a token was taken. False means the bucket was empty **and nothing was consumed** — a rejected call does not spend a future token, so a tight retry loop on `Allow` is a pure spin, not a queue. Because `Allow` gives control straight back, the caller must decide what happens to the rejected work: drop it, record it as denied, fall back to a cheaper path, or hand it to something that will retry later. That decision being explicit is exactly why `Allow` is the right call when losing the work is acceptable — sampled logging, metrics pushes, opportunistic refreshes. There is also `AllowN(t time.Time, n int)` when one operation should cost several tokens (a batch of `n` items, for instance) or when you need to evaluate against a clock you control in a test. ## Wait — be paced ``` if err := limiter.Wait(ctx); err != nil { return err // ctx done, or the request can never be satisfied } ``` `Wait(ctx)` blocks the calling goroutine until a token is available, then returns nil. It returns a non-nil error in three situations: - the context is cancelled or its deadline passes while waiting; - the deadline is already close enough that the required delay would overrun it, in which case it fails immediately rather than sleeping pointlessly; - the requested number of tokens exceeds the limiter's burst (`WaitN` with `n > burst`), which no amount of waiting can fix. When `Wait` fails it releases the tokens it had tentatively reserved, so a cancelled caller does not silently consume capacity that nobody used. The operational point about `Wait` is that it **converts overload into latency**. Every waiting caller is a live goroutine holding whatever it was holding. If arrivals exceed the configured rate for a sustained period, the number of blocked goroutines grows without limit and the symptom is climbing latency and memory, not errors. That is why `Wait(context.Background())` on a hot path is a smell: give the context a deadline so the queue has a length, or use `Allow` and shed deliberately. ## Reserve — decide for yourself `Reserve()` never blocks. It returns a `*rate.Reservation`: `OK()` reports whether the request can ever be satisfied, `Delay()` says how long you would have to wait, and `Cancel()` returns the tokens if you decide not to go ahead. It is the escape hatch for code that wants the wait time as data — to log it, to choose a different backend, or to give up when the delay exceeds what the caller can afford. The obligation that comes with it is real: a reservation you abandon without cancelling keeps its tokens spent, so a limiter can appear far busier than the traffic justifies. ## Choosing between them Ask what should happen to work that arrives above the limit. - If it should disappear, or be answered with a cheap fallback: `Allow`. - If it should simply go slower, and the caller can afford to wait a bounded time: `Wait` with a deadline. - If the caller needs to see the cost of waiting before committing: `Reserve`. Instrument the choice. A counter of allowed versus denied calls, or a histogram of the delay `Wait` actually imposed, is what later tells you whether the configured rate is the binding constraint or whether something upstream is. One misconception worth killing early: a rate limiter bounds **operations per unit of time**, not how many operations are in flight simultaneously. A limiter set to 100 per second in front of calls that each take ten seconds will happily leave a thousand of them outstanding. Rate and concurrency are separate properties and need separate controls.
- What does Limiter.Reserve give you that Allow and Wait do not?Reserve never blocks and hands back a `Reservation`. `OK` says whether the request can ever be satisfied, `Delay` reports how long you must wait, and `Cancel` gives the tokens back if you choose not to proceed. It fits code that wants to weigh the wait before committing — and abandoning a reservation without cancelling it leaves those tokens spent.
- What happens if you call limiter.WaitN(ctx, n) with n larger than the limiter's burst?It fails immediately with an error instead of blocking forever. A token bucket never holds more than burst tokens, so the request is unsatisfiable by construction. `AllowN` returns false for the same reason, and `ReserveN` returns a reservation whose `OK` is false.
- Why is calling Wait with a context that has no deadline dangerous in a busy client?Every caller blocked in Wait is a goroutine plus whatever it holds. If the arrival rate stays above the configured rate, that queue grows without bound and overload appears as rising latency and goroutine count rather than as errors you can alert on. Give the context a deadline, or use Allow and decide explicitly what to do with rejected work.
saying these in an interview costs you the question
- Says Allow blocks until a token becomes available
- Thinks Wait ignores the context passed to it
- Calls Wait with context.Background on every outbound call
- Believes a rate limiter caps how many calls are in flight
- Assumes a false from Allow has already spent a token