skip to content

Many advertiser tenants share one ad-screening inference tier; one tenant triples its traffic and all tenants slow down - why?

level: juniorimportance: must knowfreq 60%

answer

  1. one queue, many tenants
  2. finite accelerator capacity
  3. arrival rate times service time
  4. waiting explodes near saturation
  5. no errors, only latency

basics

~20 s

Every tenant draws on the same finite pool of workers, accelerator time and queue slots. The extra requests wait in the same line, so queueing delay rises for everyone as occupancy climbs. Nothing has to fail for this to happen.

solid answer

~40 s

The tier is one shared queue in front of a fixed amount of accelerator capacity, and by default that queue is tenant-blind: requests are served roughly in arrival order, so one tenant's surge sits in front of everybody else's work. Little's Law sizes the surge - required concurrency equals arrival rate times service time - so tripling arrivals triples the concurrency needed to hold waiting flat. Past that point waiting grows far faster than traffic: at about 30% occupancy a request waits under half a service time, at 90% it waits roughly nine. No request errors, so dashboards look healthy while every tenant's end-to-end p99 degrades together. The fix is per-tenant admission control with a reserved floor, not a longer shared queue.

go deeper

for a junior

Recall that a shared tier means one queue and one pool of capacity, and that contention shows up as waiting rather than as errors.

for a middle

Explain why waiting rises non-linearly as occupancy approaches saturation, and use arrival rate times service time to size the concurrency the surge actually demands.

for a senior

Show how you would prove it in production: per-tenant occupancy and p99 rather than aggregate latency, then admission control with a reserved floor while capacity is being added.

for a principal

Decide the contract: whether the sum of tenants' ceilings may exceed provisioned capacity, and who is allowed to degrade when it does. Overcommitment is a pricing decision as much as an engineering one.

## What the tenants actually share A multi-tenant ad-screening tier looks like one service, but underneath it is a handful of **finite, shared resources**, and every tenant's request competes for the same ones: - the **inbound request queue** and the worker pool that drains it; - **accelerator compute** - one device executes one batch at a time, so requests serialise there; - **accelerator memory**, which decides how many tenant policy models stay resident; - **host memory and artifact-loader bandwidth**, used whenever a model has to be pulled in; - the **batcher**, which groups requests before a forward pass and therefore mixes tenants together. None of these is per tenant unless somebody made it so. A tripling of one tenant's traffic therefore arrives as a tripling of *total* arrivals at a pool sized for the old total. ## Why waiting grows much faster than traffic A serving tier is a queueing system, and queueing delay is not linear in load. Using the standard single-queue approximation, the time a request spends *waiting* before it is served is roughly `rho / (1 - rho)` multiplied by one service time, where `rho` is occupancy - the fraction of capacity in use. | occupancy | waiting time, in service times | |---|---| | 30% | about 0.4 | | 50% | about 1 | | 80% | about 4 | | 90% | about 9 | | 95% | about 19 | So a tier that was comfortable at 30% occupancy and is pushed to 90% by one tenant's surge does not get three times slower. Waiting goes from under half a service time to about nine - a twenty-fold change in the part of latency the tenant did not pay for. That is why the complaint arrives from tenants whose own traffic never moved. Little's Law gives the other half of the picture: **concurrency = arrival rate x service time**. If the screening call takes 40 ms and total arrivals go from 500 to 1500 requests per second, the concurrency the tier must sustain goes from 20 to 60 in-flight requests. If the pool cannot hold 60, the excess becomes queue, and queue becomes latency. ## Why nothing looks broken This failure is quiet by construction: - **no errors** until requests start exceeding client timeouts, so error-rate alerts stay silent; - **accelerator utilisation near 100%** reads as good news on a cost dashboard; - **aggregate p99** shows the tier is slow but not which tenant caused it, because the aggregate mixes all tenants together; - the surging tenant may be **inside its contracted request rate**, so there is no policy violation to point at. The measurements that do resolve it are per tenant: requests per second, in-flight concurrency (or accelerator-milliseconds in flight), and p99 - each broken out by tenant, so the share of occupancy is visible next to the share of pain. ## What actually fixes it 1. **Per-tenant admission control.** Give each tenant a reserved floor it is never shed below and a ceiling it may not exceed, then evaluate both before the request reaches the accelerator. 2. **Stop serving one tenant-blind FIFO.** Separate queues per tenant or per cost class, drained by a fair-share rule, so a surge lengthens the surging tenant's own line first. 3. **Bound the queue and shed early.** A request that will wait longer than its caller's timeout is wasted work that still occupies capacity; rejecting it with a retry-after hint is strictly better than serving it late. 4. **Scale on occupancy, and expect it to be slow.** Model-serving replicas start empty and must load artifacts before they help, so autoscaling is the recovery, not the defence. ## What looks like a fix and is not - **Adding replicas on its own** - they arrive cold, and until quotas exist the surging tenant simply consumes the new capacity too. - **A tier-wide rate limit** - it caps the total but not the offender, and under contention it can reject the quiet tenants' requests just as readily. - **Raising client timeouts** - this converts a fast failure into a long wait and deepens the queue, making the tail worse for everyone. The underlying point is that sharing is the reason the tier is affordable. Isolation is bought back deliberately, in quotas and scheduling, rather than assumed.

  • The surging tenant is still inside its contracted request rate. Is anything actually wrong?
    Yes. A contract that caps arrivals does not cap occupancy, and the tier was sized for the aggregate, not for every tenant hitting its ceiling at once. Either the contracted ceilings must fit the provisioned capacity, or the tier needs reserved floors plus a shed rule that takes first from tenants above their floor.
  • Why does adding replicas not stop the slowdown immediately?
    New replicas start empty. They must be placed, pull tenant model artifacts into accelerator memory and warm up before they take useful traffic, which on a model-serving tier is seconds to minutes. Shedding and fair queueing hold the line during that window; autoscaling arrives after the damage.

One shared checkout lane: a customer with fifty items does not break the till, but everyone behind them waits. Opening another till takes minutes, while a lane reserved for small baskets protects them right now.

saying these in an interview costs you the question

  • No errors in the logs, so the serving tier is healthy
  • Adding replicas fixes a multi-tenant slowdown immediately
  • A traffic spike only hurts the tenant that caused it
  • A tier-wide rate limit protects the quiet tenants
  • Raising client timeouts makes the slowdown go away