A paid vendor caps your account at 100 requests per second. How do you set your client's rate, burst and coalescing policy?
answer
- the number is a budget, not a constant
- who else is spending the same allowance?
- replicas multiply a local limit
- duplicate calls are duplicate invoices
- the ceiling belongs to whoever pays it
basics
~20 sTreat the ceiling as a budget to allocate, not a constant to hard-code. Target below it, divide it explicitly between callers and replicas, decide whether overflow waits or is shed, and coalesce duplicate in-flight calls because each one is billable.
solid answer
~50 sThe contractual number is a ceiling, not a target. I start by asking who is spending it: how many replicas, which services, and whether batch work shares the account with interactive traffic. Local limiters compose additively, so six replicas each set to the full 100 send 600; the quota must be divided explicitly, and autoscaling makes a static division fragile, so either the share is derived from replica count or one component owns the outbound path. I aim below the ceiling, because their counter and mine disagree over retries and in-flight calls, and tripping their limiter costs more than leaving capacity unused. Burst is a latency choice constrained by their measurement window. Coalescing duplicate in-flight calls is the largest cost lever, and its price is coupling. The split between interactive and batch, and whether to buy a larger tier, belongs to whoever owns the vendor budget.
go deeper
Take away one idea: the vendor's published limit applies to the whole account, not to your one process, so several copies of your program each holding that limit will exceed it together.
Be able to compute the fleet-wide rate from a per-process configuration and explain why the target should sit below the contractual ceiling rather than on it.
Show the operational instrumentation — your allow and deny counts against the vendor's rejections — and the concrete mechanics for dividing one allowance across replicas and workloads.
Own the allocation as a budget with a named owner: the fraction of the paid ceiling you target, the split between interactive and batch, the coupling coalescing introduces, and when buying a larger tier beats engineering around the limit.
## Why this is a judgment call and not a constant A contractual limit looks like a number you can paste into a constructor. It is really an allowance that several parties are spending simultaneously, measured by someone else's clock, priced by someone else's contract, and owned by someone who may not be an engineer. The interesting decisions are about allocation, overflow behaviour and cost — and each of them can legitimately be overruled by the person who pays the invoice. ## Decide the target, not the ceiling Run below the contractual number. Two reasons, both practical. First, the two counters disagree: retries you consider one logical call are several to them, their window boundaries are not yours, in-flight calls are counted at different instants, and other consumers of the same account exist that you may not have inventoried. Second, the penalties are asymmetric. Leaving five percent of a paid quota unused costs five percent of a line item; tripping their limiter costs failed user-visible work, and with some vendors a rejected call still bills or still counts. How far below is a real decision — a comfortable margin for a hard contractual cap that triggers penalties, a thin one for a soft limit that merely rejects. ## Allocate the budget across the fleet and across workloads This is the part teams get wrong most often. A per-process limiter bounds one process. N replicas each holding the full allowance send N times the allowance, and the failure appears exactly when you scale up to handle load — the worst possible moment. The options each have a cost: - **Divide by replica count.** Simple and local, but wrong the moment the replica count changes, and wasteful when replicas are unevenly loaded: a quiet replica's share is unusable by a busy one. - **Funnel the traffic through one component** that owns the vendor relationship. Accurate and easy to reason about, at the price of a new dependency in the path and a new thing to make highly available. - **Coordinate through shared state.** Precise but adds latency and a failure mode to every call; you then have to decide what happens when the coordination store is unreachable — fail open and risk the contract, or fail closed and take an outage for a dependency the vendor never sees. Separately, allocate across *workloads*. Interactive requests and an overnight backfill drawing on one quota is a policy question: does the backfill get the leftovers, a fixed slice, or the whole thing outside business hours? Writing that split down, per caller, is the artefact that makes the decision reviewable later. ## Burst is a latency knob with a contractual edge A larger burst improves tail latency for spiky interactive traffic by letting queued work leave at once. It does not change the average. The constraint is the vendor's measurement window: if they enforce per-second and you burst a hundred in one instant, you are non-compliant on their meter even though your minute is well under. Absent documentation, a burst near their stated window's allowance is the defensible choice, and the smallest burst that meets your latency goal is the safest. ## Coalescing is the cost lever, and coupling is its price When the same answer is requested by many callers at once, collapsing those into one billable call is often the single largest saving available, and it costs nothing in correctness *if* the key is exact. The price is coupling: every waiter now shares one execution's latency and one execution's failure. That is a tradeoff the product owner should see stated — fewer vendor calls and lower spend, against a failure that lands on many callers at once instead of one. Whether the savings justify it is measurable: the share of calls that were coalesced tells you directly what you are buying. ## Choose the overflow behaviour deliberately Above the limit, work either waits or is rejected. Waiting preserves the work and spends latency and memory; rejecting preserves latency and loses work. Different callers deserve different answers, and "whatever the library did by default" is not one of them. For interactive traffic, a bounded wait then a clear failure is usually right; for a backfill, waiting indefinitely is fine because nobody is watching. ## Instrument so the policy can be argued about Three signals make the whole thing legible: allowed versus denied counts at your limiter, the vendor's own rejection rate, and the coalesced share of calls. The combinations are diagnostic. Denials at your limiter while the vendor never rejects means you are self-limiting below the capacity you already pay for — raise the target or ask why. Vendor rejections while your limiter is happy means your model of their accounting is wrong, most often because a caller you did not know about shares the account. Both are cheaper to discover on a graph than in an invoice. ## Name the owner The engineering choices — where the limiter lives, how the fleet divides the allowance, which calls coalesce — are yours. The target as a fraction of the paid ceiling, the split between paying and batch workloads, and whether the right move is to buy a larger tier rather than engineer around the limit belong to the person who owns the vendor budget. Present the arithmetic and the options; expect to be overruled on the number, and design so that being overruled is a configuration change rather than a redeployment.
- The service autoscales, so dividing the quota by replica count keeps going stale. What do you do?Either derive each replica's share from a count it can observe, and accept a conservative margin while that count is in flux, or move the vendor call behind one component that owns the allowance. I would pick based on how expensive a breach is: a hard contractual cap justifies the extra dependency, a soft limit that merely rejects usually does not.
- How would you argue for buying a larger vendor tier rather than engineering around the limit?By pricing both sides. The engineering cost is the work plus the permanent complexity of coalescing, coordination and their failure modes; the tier cost is a known monthly number. If projected traffic makes the limit binding within a couple of quarters, buying capacity is usually cheaper than owning a distributed limiter forever. That comparison is exactly the one the budget owner is equipped to decide.
- Your limiter denies a meaningful share of calls while the vendor never rejects anything. What does that tell you?That you are self-limiting below the capacity you already pay for. Either the target was set conservatively and can be raised, or the traffic is spikier than the burst allows and the fix is burst rather than rate. Both are cheap wins, and neither is visible without charting your own denials against their rejections.
- What would make you refuse to coalesce duplicate calls even though it would cut spend?When callers must not share a fate. If one execution's failure or staleness landing on every waiter is unacceptable — an authorisation decision, or a call whose result differs subtly per caller so the key cannot be exact — the coupling is not worth the saving. Coalescing is safe only when the key genuinely identifies identical work.
saying these in an interview costs you the question
- Hard-codes the contractual ceiling as the client rate
- Gives every replica the full account allowance
- Never asks who else spends the same quota
- Treats the limit as purely an engineering decision
- Ignores that duplicate concurrent calls are billed twice