On a shared ad-screening inference tier, what does a per-tenant concurrency quota protect that a per-tenant request-rate quota does not?
answer
- door versus room
- arrivals are not occupancy
- concurrency equals rate times service time
- three dimensions, not one
- resident memory is also a quota
basics
~20 sOccupancy. A concurrency quota caps how much of the tier a tenant holds at once, so it still binds when requests get slower. A request-rate quota counts arrivals, and arrivals stay flat while in-flight work multiplies.
solid answer
~50 sA rate quota is a statement about the door; a concurrency quota is a statement about the room. Little's Law shows the gap: at 100 requests per second and 50 ms per request a tenant holds about 5 requests in flight, but if its policy model or a downstream lookup slows to 500 ms, the same 100 requests per second holds about 50 in flight - ten times the occupancy, with the rate quota never tripping. A multi-tenant serving tier therefore needs three per-tenant dimensions: request rate, concurrency (counted in in-flight requests, or better in accelerator-milliseconds), and resident model memory, which bounds how many pinned versions a tenant may keep loaded. All three are checked at admission; a tenant below its reserved floor is never shed, and traffic above its ceiling is rejected with a retry-after hint rather than queued indefinitely.
code
json · 12 lines{
"tenant": "tenant-47",
"pinnedModel": { "name": "ad-policy-screen", "version": "2026-08-11.3" },
"quotas": {
"requestsPerSecond": { "reserved": 40, "ceiling": 120 },
"concurrentRequests": { "reserved": 4, "ceiling": 10 },
"residentModelBytes": 3221225472
},
"burst": { "creditSeconds": 30, "refillPerMinute": 10 },
"onExceed": "shed",
"retryAfterSeconds": 2
}go deeper
Know that a quota limits what one tenant may consume, and that requests per second is only one of several things worth limiting on a shared tier.
Explain the three dimensions - rate, concurrency and resident model memory - and use arrival rate times service time to show why a rate cap stops binding when latency drifts up.
Describe enforcement in the real path: where admission runs, what a reserved floor guarantees, what the shed response tells the caller, and how the counter is released on error paths.
Own the overcommitment ratio and the degradation order between tenants, plus the cost of the headroom that makes a reserved floor a guarantee rather than a number in a config file.
## Three things a tenant consumes A tenant on a shared ad-screening tier consumes three different scarce resources, and a quota on one of them says nothing about the other two. | dimension | unit | what it protects | where it goes blind | |---|---|---|---| | request rate | requests per second | arrival fairness and downstream call volume | per-request cost varies, so equal rates are not equal load | | concurrency | in-flight requests, or accelerator-milliseconds in flight | occupancy of the serving tier itself | needs a per-tenant ceiling well below the whole pool | | resident memory | bytes, or resident model slots | other tenants' models staying loaded | says nothing at all about traffic volume | Most platforms ship the first one, because it is easy to count at the edge, and discover the other two during an incident. ## Why a request-rate quota stops binding **Little's Law** states that concurrency equals arrival rate times service time. Hold the arrival rate fixed and let service time move, and occupancy moves with it: - 100 requests per second at 50 ms each is about **5 requests in flight**; - the same 100 requests per second at 500 ms each is about **50 requests in flight**. Nothing the rate quota measures has changed. Service time drifts upward for reasons that have nothing to do with the tenant's behaviour - a heavier pinned model version, a slower feature lookup, a colder cache, more contention from other tenants - and each of those multiplies the tenant's footprint while its rate stays legal. A concurrency quota binds immediately in all of those cases, because it measures the thing that is actually scarce. The rate quota is still worth having. It is the cheapest control for downstream call volume and for billing, and it is the only one that limits a tenant's demand *before* work starts. It simply does not bound occupancy. ## Where the check runs 1. The tenant is known (identity resolution happens upstream of the serving tier). 2. The tenant's **pinned policy-model version** is resolved from the serving policy. 3. All three quota dimensions are evaluated together against the tenant's current usage. 4. The request is admitted, queued in the tenant's own bounded lane, or shed with a retryable response carrying a retry-after hint. 5. On completion - including on error and on timeout - the concurrency count is released; a release path that only runs on success leaks the counter until the tier appears permanently full. Two details decide whether this works in production. **Retries must be charged to the same bucket**, or a throttled client amplifies its own overload. And **the shed response must be distinguishable from a failure**, so a caller backs off rather than treating it as a reason to retry immediately. ## Floors, ceilings and borrowing - A **reserved floor** is the capacity a tenant is guaranteed. It is only real if the sum of floors fits within provisioned capacity; otherwise it is a number in a config file. - A **ceiling** bounds blast radius: the most one tenant can take even when the tier is idle. - **Work-conserving borrowing** lets a tenant use idle capacity above its floor, which is what makes sharing cheaper than dedicated capacity. - **Reclaim is not instant.** When the lender's traffic returns, the borrowed capacity comes back only as in-flight requests finish, so reclaim latency is bounded by the longest request the tier permits. That is a second reason to cap per-request execution time. ## Failure modes seen in practice - Quotas denominated in requests while per-request cost across tenants varies by an order of magnitude. - A single tier-wide limit instead of per-tenant limits: it stops the tier collapsing but lets one tenant hold the whole pool. - Counting only at the front door while retries and internal fan-out multiply behind it. - No quota on resident model memory, so one tenant pinning many versions evicts everybody else's models. - Throttling without a retry-after signal, turning a shed into a retry storm. - Floors that sum to more than capacity, which means the guarantee fails precisely when it is needed.
- Should the concurrency quota count requests or accelerator time?Accelerator-milliseconds is the honest unit when tenants run models of very different per-request cost, because counting requests charges a 400 ms inference the same as a 20 ms one. Counting in-flight requests is acceptable when every tenant's model sits in roughly the same cost band, and it is far cheaper to measure and explain.
- What if every tenant's ceiling is honoured but the ceilings sum to more than the tier can serve?The tier is overcommitted, which is usually deliberate because tenants peak at different times. It is safe only if the reserved floors sum to less than capacity and the shed rule takes first from tenants above their floor. Otherwise the first simultaneous peak degrades everyone equally, including tenants that never exceeded their contract.
- Where does a quota on resident model memory actually bite?At load time. When a tenant's next pinned version would push its resident footprint past the quota, the platform evicts one of that tenant's own versions rather than a neighbour's. Without it, eviction is global and the tenant loading the most versions quietly pushes everyone else's models out of accelerator memory.
saying these in an interview costs you the question
- A request-rate quota is enough to stop noisy neighbours
- One tier-wide concurrency limit gives each tenant a fair share
- Counting admitted requests is the same as counting capacity
- Model memory is the platform's problem, not a quota dimension
- Throttled requests can be retried immediately without backoff