skip to content

How do you set a free-tier API's per-client rate limit, and decide whether over-limit requests queue or get 429?

level: principalimportance: should knowfreq 35%

answer

  1. two decisions, two owners
  2. the ceiling comes from a load test
  3. count first, enforce later
  4. who can change it at 3am?
  5. in-process buckets multiply by replica count

basics

~20 s

The number is a product decision bounded by measured capacity, not an engineering guess. Ship it in shadow mode first. Default to shedding with 429 on a public tier, and keep the value changeable without a deploy.

solid answer

~50 s

Two people own the two halves. The quota is a product call — what a free tier is worth and which usage it should support — bounded by a capacity number engineering measures rather than guesses. The over-limit behaviour is an engineering and operational call: shedding with 429 keeps the server's cost proportional to the traffic it intends to serve, while blocking converts an error into latency, goroutines and memory, and under sustained overload fails worse than the error it avoided. So a public free tier sheds; queueing is defensible only for a caller set you control, with a hard deadline. Run the number in shadow mode first, counting per client what would have been rejected, then enforce from configuration that changes without a deploy, because on-call must move it during an incident. Remember an in-process limiter is per replica.

go deeper

for a junior

You are not expected to set the number. Know that a limit is configuration rather than a constant in the code, and that rejecting with 429 is the normal answer for a public API.

for a middle

Be able to explain what rejecting costs versus what waiting costs the server, and why a burst allowance and a sustained rate are separate settings that serve different traffic shapes.

for a senior

Show how you would derive the ceiling from a load test, roll the limit out in shadow mode, and instrument shed counts per client. Be ready to say what you would change during an incident and what you would leave alone.

for a principal

Own the split: the quota belongs to whoever owns the tier's economics, the over-limit behaviour and the global cap belong to engineering, and both must be movable without a deploy. Be able to defend refusing a quota the measured capacity cannot honour.

## Separate the number from the behaviour Two decisions are usually collapsed into one and should not be. **What the limit is** is a product decision: the free tier exists to let people evaluate the service and to convert, so the limit expresses how much evaluation is free. **What happens past it** is an engineering and operational decision about how the service behaves under more load than it intends to serve. Confusing them produces the classic outcome where a number nobody owns is hard-coded in a middleware, and the only person who can change it during an incident is whoever can ship a deploy. ## Bounding the number with a measurement The product number is not free to pick, because it is bounded above by what the service can actually serve. That ceiling comes from a load test, not from arithmetic on a CPU count: the usable throughput is where latency turns up sharply, not where the machine saturates. Divide that by the concurrent client population you expect and you have the region the product number must live inside. If the product's desired quota is above that ceiling, the honest answer is that the tier needs more capacity or a smaller promise — not a limiter tuned to a number the service cannot honour. Burst and sustained rate are separate knobs and should be argued separately. Real clients are bursty in a benign way: a page load fires eight parallel calls, a batch job wakes on a cron minute. A limit that permits a reasonable burst while holding the sustained rate down rejects the abuse without rejecting the legitimate shape of the traffic. Setting burst to one is the single most common way to make a correct sustained limit feel broken to honest users. ## Choosing shed over queue Shedding is the default for anything public. A rejection costs a comparison and a small response, so the server's resource use stays proportional to the work it intends to do, and the caller learns immediately and can back off. Queueing — parking the request inside a blocking limiter until a token frees — trades an error for latency, and the resources it spends are real: a goroutine and stack per waiter, the accepted connection, whatever the middleware already allocated, and a slot in every upstream pool between the client and you. Under a burst that drains, that trade is good. Under sustained excess, the queue is unbounded and you have converted a clean 429 into rising memory, rising tail latency for the requests you did admit, and eventually timeouts anyway. So: queue only where the caller set is known, bounded and yours, and always with a deadline on the wait so a request cannot outlive its usefulness. One complication you must state rather than discover: if the limiter lives in the process, every replica has its own buckets, so the effective per-client limit is your configured rate times the number of replicas — and it moves every time the deployment scales. Either divide the configured number by the replica count and accept the imprecision, or move the counter out of the process, which is a different and much more expensive design. Decide this deliberately; do not let an autoscaler silently redefine your product's quota. ## Rolling it out Enforcing a new limit is a change to your product's behaviour for existing customers, so ship it in two steps. First, shadow mode: evaluate the limiter, do not reject, and emit a counter per client of what *would* have been shed. A week of that data tells you which real customers the number bites, and it is the artefact that lets the product owner move the number knowingly rather than after the escalation. Then enforce, with the value read from configuration that changes without a deploy, so on-call can raise it when the limiter is hurting more than it protects and lower it under abuse. Instrument admitted and shed counts, and the number of distinct limiter keys — the last one is your early warning that the limiter's own memory is becoming the problem. Document the limit and the rejection publicly, because a limit clients cannot discover produces retry loops, and a retry loop against your 429s costs you more than the traffic you rejected. ## What you own when it goes wrong The posture to defend in the room is: the per-client limit is fairness, and it will not save the service when the load is not one client's fault. That is what a separate global admission cap is for, and it is the one on-call reaches for at 3am. The per-client number is the product's; the global cap is engineering's, and the argument for having both is that they fail for different reasons.

  • How do you introduce a limit on a tier that has never had one?
    Shadow mode first: evaluate the limiter and record, per client, what would have been rejected, without rejecting anything. A week of that shows which real customers the number bites and turns the conversation with the product owner into data. Then enforce, with the value in configuration that changes without a deploy, and watch the shed counter per client for the first days.
  • The limiter lives in the process and the service runs on many replicas. What does that do to the quota?
    Each replica keeps its own buckets, so a client can get up to the configured rate from every replica it reaches, and the effective limit changes whenever the deployment scales. Either divide the configured value by the replica count and accept the imprecision, or move the counter out of the process — an expensive design change. The one thing you must not do is let autoscaling silently redefine the product's quota.
  • When is queueing over-limit requests actually the right call?
    When the caller set is known, bounded and yours — an internal batch job or a fan-out you wrote — and the overload is a short burst that drains, so added latency really is better than an error the caller would just retry. Always with a deadline on the wait and a cap on concurrent waiters, so the queue cannot grow with the arrival rate.
  • On call at 3am, the limiter is rejecting a major customer. What do you change and what do you not?
    Raise or disable that client's limit through configuration, because the per-client number is fairness and fairness is not worth an outage for a paying customer. Do not remove the global admission cap in the same move — that is what is keeping the service alive. Then hand the incident's data to whoever owns the quota, so the number is revisited deliberately rather than left at the value you typed at 3am.

saying these in an interview costs you the question

  • Picks the limit from a CPU count instead of a load test
  • Hard-codes the quota so changing it needs a deploy
  • Queues public over-limit traffic to avoid returning errors
  • Ignores that in-process limiters multiply by replica count
  • Enforces a new limit on existing customers with no shadow period
  • Treats a per-client limit as protection against total overload