skip to content

In a multi-tenant screening tier where each tenant pins one policy-model version, how should a request reach a replica that already holds it?

level: seniorimportance: should knowfreq 38%

answer

  1. key is tenant plus version
  2. resolve once, at the router
  3. never the word latest
  4. sticky subset keeps models resident
  5. pre-warm before the table flips

basics

~20 s

Resolve the pinned version once at the router from a versioned routing table, then balance only across replicas that already hold that version, keeping the routing key sticky. Cold-loading elsewhere is the fallback, never a different version.

solid answer

~50 s

The routing key is the pair (tenant, pinned version), never the word latest. Resolve it once, at the router, so a retry of the same submission cannot land on a different version. Hash that key onto a small subset of replicas - a handful of candidates rather than the whole fleet - and pick the least loaded candidate that already holds the version; concentrating a version on few replicas is what keeps it resident at all. If no candidate holds it, a bounded cold load on the least loaded one is acceptable, and shedding with a retry-after is acceptable, but silently serving a neighbouring version is not: in policy enforcement that means the same creative gets two different verdicts. Pin changes are a control-plane update with a revision number, applied after the new version is pre-warmed on the target subset.

code

pseudocode · 18 lines
pseudocode
function route(request):
    pin        = routingTable.lookup(request.tenantId)   // { version, revision }
    key        = hash(request.tenantId + ":" + pin.version)
    candidates = ring.subset(key, k = subsetSizeFor(request.tenantId))
    holders    = filter(candidates, r -> r.holds(pin.version))

    if holders is empty:
        target = leastLoaded(candidates)
        if target.coldLoadsInFlight >= 1:
            return shed(retryAfter = 2)       // never fall back to another version
        target.requestLoad(pin.version)
    else:
        target = leastLoaded(holders)

    response = target.infer(request.creative, pin.version)
    response.servedVersion    = pin.version
    response.routingRevision  = pin.revision
    return response

go deeper

for a junior

Recall that a tenant is tied to one specific model version, and that the request has to reach a replica that has that exact version loaded.

for a middle

Explain why the routing key pairs tenant with pinned version, and why concentrating a version on a few replicas is what makes residency affordable.

for a senior

Describe the whole path under change: pre-warm, a revisioned table flip, draining the old version, and a fallback order that never silently substitutes a different version.

for a principal

Set the policy: what the tier owes a tenant about which version judged their content, and what the platform is allowed to do when the pinned version cannot be served in budget.

## The routing key, and where it is resolved Every request carries a tenant; the serving policy turns that tenant into exactly one pinned model version. Two properties matter more than the lookup itself: - **Resolution happens once, at the router.** If each replica resolves the version for itself, two replicas with different views of the routing table will answer the same submission under different versions. - **Resolution is stamped onto the response.** The version that actually scored the request travels back with the verdict, so an appeal weeks later is answerable rather than a guess. The key that drives placement is the pair `(tenantId, pinnedVersion)` - not the tenant alone, because a pin change must move traffic, and not the version alone, because tenants sharing a version can still be scheduled apart. ## Sticky placement without hot spots Spreading every version across every replica would be perfect for load balancing and fatal for residency: each replica would have to keep all of them in accelerator memory. Concentrating each version on a subset does the opposite. Consistent hashing of the routing key onto a subset of `k` replicas, then choosing the least loaded member of that subset, gives both properties: - small `k` means **few copies** of each version, so more distinct models fit in the fleet; - larger `k` means **more headroom** for one tenant's burst, and more copies to keep warm. The sizing is arithmetic, not taste. With 400 pinned versions at about 2 GB each, the fleet must hold 800 GB of models per full copy. Sixty replicas with 24 GB of usable accelerator memory each provide 1440 GB, so the average replication factor cannot exceed about 1.8. Setting `k = 3` for everyone would demand 2400 GB and simply will not fit; the workable answer is a larger subset for high-traffic tenants and `k` of one or two for the long tail. ## What breaks when routing is loose 1. **A replica resolving latest itself.** Mid-rollout, two replicas disagree, and one advertiser's creative is approved by one version and rejected by the other. 2. **Re-resolving on retry.** A client retry that resolves the pin again can hit a newer version, so the same submission produces a different verdict without anything being logged as a change. 3. **Per-replica caching of the routing table with different expiry.** A pin rollback then takes effect unevenly, and the tier serves a mixture for as long as the longest cache lives. 4. **Not returning the served version.** The verdict becomes unattributable; nobody can say which version produced it. ## Fallbacks that do not make things worse When no candidate replica holds the pinned version: - **cold-load it** on the least loaded candidate, with the loader bounded so a routing miss cannot become a load storm; - **shed with a retry-after hint** if the cold load would exceed the caller's budget; - **never substitute another version**. In an enforcement path a wrong-version verdict is worse than no verdict, because it is acted on and looks legitimate. ## Pin changes and draining A version change is a routing-table update, and it is safe only when the target subset is prepared: - **pre-warm** the new version on the replicas the new key hashes to, before the table flips; - **flip the table with a revision number**, so every component can report which revision it is serving; - **drain the old version** rather than evicting it - in-flight requests must finish on the replica that started them; - **keep the old artifact available** for long enough that a pin rollback is a table revert and not a re-download; a rollback whose artifact has been garbage-collected is not a rollback. The theme throughout is that the pin is a promise to the tenant about which policy their creatives were judged under. Routing is the mechanism that keeps that promise on every request, including retries and during changes.

  • Why not spread every version evenly across the whole replica fleet?
    Because every replica would then need every model resident, and accelerator memory does not hold them. Concentrating each version on a subset is what makes residency possible at all; the cost is that a tenant's burst capacity is bounded by its subset, which is why the subset is sized per tenant rather than globally.
  • What has to be true for a pin rollback to actually work?
    The previous artifact must still be retrievable and ideally still resident somewhere, the routing table must be revertible as a single revision, and in-flight requests on the new version must drain rather than be killed. A rollback whose artifact has been garbage-collected is a re-deployment with a long cold start, not a rollback.

saying these in an interview costs you the question

  • Replicas can resolve the latest version for themselves
  • Serving a near-identical version is fine under load
  • A retry may safely re-resolve which version to use
  • Spreading every model across every replica balances load best
  • A pin change takes effect the moment the table is written