A screening tier serves 400 tenant policy models but keeps only 48 resident in accelerator memory; why is a rarely-called tenant's end-to-end p99 measured in seconds?
answer
- resident slots are a cache
- the tail is always the miss
- eviction follows recent traffic
- fetch, deserialise, warm up
- pinned slots cost shared capacity
basics
~20 sResident slots are a cache, and the rare tenant is the miss. Called once an hour, its model is evicted between calls, so nearly every request first pays a cold load - fetch, deserialise and warm a multi-gigabyte artifact - before inference starts.
solid answer
~40 sWith 400 pinned versions competing for 48 resident slots and least-recently-used eviction, residency follows traffic: busy tenants keep touching their models, and the long tail is always the victim. For a low-traffic tenant the hit rate is near zero, so its typical request pays the full cold path - roughly 4 s to pull a 2 GB artifact from object storage at about 500 MB/s, plus deserialisation and a first-call warmup, so about 5 s against a 40 ms warm inference. The same effect poisons the whole tier when misses cluster: concurrent loads contend for host memory and loader bandwidth, so warm tenants slow down too. The fixes are reserved residency slots for tenants with a latency commitment, a bounded single-flight loader, scheduled pre-warm, and a separate published SLO for the cold pool.
code
pseudocode · 22 linesfunction handle(request):
version = pinnedVersion(request.tenantId)
slot = residentSlots.get(version)
if slot != null:
return slot.infer(request.creative)
if loadsInFlight >= MAX_CONCURRENT_LOADS:
return shed("cold-load capacity exhausted", retryAfter = 2)
slot = singleFlight(version, function():
victim = chooseEvictable(residentSlots, excluding = pinnedForSlo)
if victim == null:
return null
drainThenEvict(victim)
return load(version) // about 5 s: fetch, deserialise, warm
)
if slot == null:
return shed("no evictable slot", retryAfter = 2)
return slot.infer(request.creative)go deeper
Recall that a model must be loaded into accelerator memory before it can score anything, and that loading a multi-gigabyte artifact takes seconds.
Explain residency as a cache with eviction, and work out why a small overall miss rate lands entirely in the tail percentiles rather than in the mean.
Show the operational defences: single-flight and bounded loading, reserved slots for latency-committed tenants, pre-warm, and shedding instead of queueing behind a cold load.
Decide whether the platform sells one latency contract or two, who pays for pinned residency, and how much utilisation the organisation gives up to make the tail tenant's SLO true.
## Residency is a cache, and the tail is the miss Accelerator memory holds a fixed number of models. With 8 devices holding 6 models each, 48 of 400 pinned versions are resident at any moment - about 12%. Under least-recently-used eviction, residency is allocated in proportion to recent traffic, which is exactly the wrong allocation for a tenant that cares about latency but sends little traffic. A tenant calling once an hour is guaranteed to have been evicted by dozens of other tenants' requests in the meantime, so its hit rate is effectively zero and its *median* request is a cold one. ## What a cold load actually costs Break the cold path into its parts rather than quoting one number: 1. **Fetch the artifact** - 2 GB from object storage at about 500 MB/s of effective throughput: **4.0 s**. 2. **Deserialise and allocate device memory** - about **0.6 s**. 3. **First-call warmup** - graph or kernel initialisation and allocator growth on the first inference: about **0.4 s**. Total **5.0 s**, against a warm inference of 40 ms. The cold path is roughly **125 times** the warm one, which is why it dominates any percentile it reaches. ## Why a small miss rate ruins p99 but barely moves the mean Suppose 5% of all requests across the tier are cold: | statistic | value | why | |---|---|---| | mean latency | about 288 ms | 0.95 x 40 ms + 0.05 x 5000 ms | | p95 | about 40 ms to 5 s | the boundary sits exactly at the cold fraction | | p99 | about 5 s | the slowest 5% of requests are all cold loads | The mean looks tolerable and the p99 is catastrophic, which is why a tier can pass an average-latency dashboard while a handful of tenants are unusable. For the rare tenant itself the arithmetic is worse: its own miss rate is near 100%, so its p50 is the cold path. ## Eviction policy is a design decision, not a default Least-recently-used is traffic-proportional, so it always sacrifices the tail. Alternatives, each with a price: - **Reserved (pinned) residency slots** for tenants with a latency commitment. This works, and it costs real capacity: every pinned slot is one fewer slot the shared pool can use, so the remaining tenants contend harder. - **Cost-aware eviction** - evict the model whose reload cost multiplied by its hit probability is lowest, rather than simply the least recently used. - **Two pools with two contracts** - a hot pool with pinned residency and a tight latency SLO, and a cold pool where a multi-second first response is the *published* behaviour rather than an incident. - **Shrink the footprint** - where tenants share a common base model with small per-tenant deltas, only the delta needs loading, and the resident count rises by an order of magnitude. ## Guard the loader, or a miss storm becomes an outage Cold loads are not merely slow; they are expensive for everybody: - **Single-flight per version** - ten simultaneous requests for one evicted model must trigger one load, not ten. - **Bounded loader concurrency** - loads consume host memory and artifact bandwidth, and unbounded loading starves warm inference on the same device. - **Shed rather than queue** behind a load that will outlive the caller's timeout; a request that returns nothing is cheaper than one that holds a slot for 5 s and is then discarded. - **A local disk cache** between object storage and the device: 2 GB from local disk at about 2 GB/s is roughly 1 s instead of 4 s, turning the worst case into something a generous timeout can absorb. - **Never evict a version with requests in flight**; drain first, or the in-flight request dies with the slot. ## What this changes in the SLO conversation The honest outcome is usually two contracts rather than one: tenants who need a tight p99 buy pinned residency and pay for the slot, while everyone else gets a published cold-start behaviour and a warm-path SLO that applies once the model is loaded. Pretending one number covers both is how a tier ends up with a p99 that nobody can explain and no tenant recognises as their own.
- Ten requests for the same evicted model arrive at once. What should the tier do?Collapse them into one load with a single-flight guard and let the rest wait on that load or shed with a retry-after hint. Starting ten loads multiplies bandwidth and host memory use, can exhaust the device, and slows the warm tenants sharing it - turning one tenant's miss into everyone's incident.
- Does routing traffic for a version to fewer replicas help or hurt residency?It helps residency and costs headroom. Concentrating a version on a small replica subset means fewer copies to keep loaded, so more distinct models fit overall. The price is that the tenant's burst capacity is capped by that subset, which is why the subset size is tuned per tenant rather than globally.
saying these in an interview costs you the question
- Average latency is fine, so the tier is healthy
- Least-recently-used eviction is fair to every tenant
- A longer client timeout solves cold-start latency
- Pinning residency for a tenant costs nothing
- Loading a model on demand has no effect on other tenants