skip to content

How do you route LoRA-adapter requests across a fleet of LLM replicas?

level: seniorimportance: should knowfreq 35%

answer

  1. base weights once, adapters many
  2. the adapter must already be resident
  3. a cap on distinct adapters per batch
  4. hash by adapter id, not round-robin
  5. hot tenants earn their own replicas

basics

~20 s

Route by adapter identity, not round-robin. A replica can only serve an adapter it currently has resident, and engines cap how many distinct adapters may appear in one batch, so hash adapters to replicas, keep each replica's hot set stable, and give a heavy tenant its own group.

solid answer

~60 s

Multi-LoRA serving loads the base weights once and layers small per-tenant adapters on top, so one replica can serve many tenants — vLLM does this behind `--enable-lora`, with `--max-loras` capping how many distinct adapters may appear in a single batch. That cap is the routing constraint: if requests for fifty adapters arrive uniformly at every replica, each replica thrashes between adapter sets and pays repeated loads, while a replica that serves a stable handful runs efficiently. So make adapter id the routing key. Consistent-hash adapters onto replicas so each holds a stable working set and a membership change moves only about 1/N of them, and keep a general pool for the cold long tail. Watch for skew: one tenant with real traffic needs its own group of replicas rather than a slot in someone else's, and a tenant with a strict latency SLO may deserve a dedicated replica or a merged checkpoint. Autoscale the pool on the usual saturation signals, not per adapter — adapters are megabytes, replicas are the expensive unit.

go deeper

for a junior

Know that many small adapters can share one loaded base model, which is far cheaper than running a separate server per customer.

for a middle

Explain the residency constraint and the per-batch adapter cap, and why routing by adapter id keeps each replica's working set small instead of making every replica juggle every adapter.

for a senior

Show the production design: consistent hashing with a replication factor, a shared pool for the cold tail, promotion of hot tenants to dedicated replicas, and readiness that includes adapter fetch.

for a principal

Own the tenancy policy — which tenants get isolation, what per-tenant SLOs you are willing to sign given shared KV-cache and batch slots, and when merging an adapter into a dedicated checkpoint is worth the extra artifact and pipeline.

## What multi-LoRA serving buys, and what it constrains A LoRA adapter is a small pair of low-rank matrices per targeted layer — typically tens to a few hundred megabytes, against tens of gigabytes of base weights. Serving many adapters over one shared base is therefore enormously cheaper than one replica per tenant: you pay for the base once and add a rounding error per tenant. Engines implement this with batched adapter kernels that let sequences using different adapters share a single forward pass over the base weights. The constraints that shape routing are: - **Residency.** An adapter must be loaded before it can serve. Loading is fast relative to base weights but not free, and it competes for host and device memory. - **A cap on distinct adapters per batch.** vLLM exposes this as `--max-loras`, alongside `--max-lora-rank` for the largest rank it will accept and a separate cap on how many adapters it will keep cached in host memory. Requests for adapters beyond the cap wait for a slot rather than joining the current batch. - **A per-token cost that rises with adapter diversity.** A batch spanning many adapters does more work than a batch spanning one, because the adapter kernels must be applied per group. All three push the same direction: a replica performs best when it serves a small, stable set of adapters. ## Why round-robin is the wrong default here Spray requests for fifty adapters uniformly across eight replicas and every replica sees all fifty. Each one continually evicts and reloads adapters, exceeds its per-batch cap, and forms queues behind adapter-slot contention — while the fleet as a whole is nowhere near its compute limit. The symptom is puzzling: TTFT is poor, GPU utilization is high, and adding replicas barely helps, because each new replica also sees all fifty adapters. ## Adapter-affine routing Make the adapter id the routing key and hash it onto the fleet: - **Consistent hashing** so a scale-out or a rolling restart relocates about 1/N of adapters rather than all of them. Each replica then has a stable working set well inside its per-batch cap. - **Replication factor.** Map each adapter to two replicas rather than one, so a single pod failure or restart does not take a tenant offline and so the router has a choice when one of the pair is busy. This is the direct analogue of choosing between affinity and balance in prefix routing. - **A general pool** for the long tail. Adapters with a handful of requests a day do not deserve a residency guarantee; route them to a shared group that accepts the load cost, and promote an adapter into the affine tier when its rate crosses a threshold. ## Tiering by traffic and by SLO Adapter populations are almost always power-law shaped: a few tenants generate most traffic, and hundreds are nearly idle. Treat the tiers differently. - **Hot tenants** get a dedicated group of replicas. At that point the multiplexing benefit is gone anyway — the replica is serving one adapter almost exclusively — and you gain isolation and predictable latency. If the adapter is stable, merging it into the base weights and serving it as a plain checkpoint removes the adapter kernels entirely, at the cost of losing per-request switching and doubling the storage for that model. - **Warm tenants** live in the affine tier, hashed to a small replica set. - **Cold tenants** live in the shared pool and pay a load on first use after eviction. A per-tenant latency SLO is the usual reason to promote a tenant out of the shared tier, and noisy-neighbour risk is the usual reason to keep the tiers separate at all: on a shared replica, one tenant's flood of long generations consumes KV-cache blocks and batch slots that every co-resident tenant needs. ## Autoscaling this fleet Scale the replica pool on the same saturation signals as any other inference fleet — queued requests, KV-cache occupancy, TTFT — not on adapter count. Adding a replica to a hash ring is what redistributes adapters; there is no meaningful notion of scaling one adapter. Two wrinkles matter. First, a new replica must fetch its assigned adapters as part of becoming ready, so the readiness gate should include them; happily they are small compared with base weights. Second, because the ring moves keys on membership change, aggressive scale-in and scale-out churn the adapter assignments, which argues for the same conservative scale-in policy that cold start already forces on you. ## The anti-pattern to name The reflex design is one deployment per tenant. It is operationally simple and financially indefensible: each tenant gets a full copy of the base weights in VRAM, so a hundred tenants on a 16 GB model need a hundred GPUs' worth of weight storage to serve traffic that would fit on a handful of shared replicas. Multi-LoRA exists precisely to collapse that, and adapter-affine routing is what makes the collapsed version behave.

  • When should a hot tenant's adapter be merged into the base weights instead?
    When its traffic justifies dedicated replicas anyway and the adapter is stable. Merging produces a plain checkpoint, removing per-token adapter kernel overhead and the per-batch adapter cap. The costs are a second full-size checkpoint to store and deploy, and the loss of per-request switching — you can no longer serve that tenant and others from one process.
  • How do you keep one tenant on a shared replica from starving its neighbours?
    Cap per-tenant concurrency and output length at the router, so no single adapter can occupy the whole running batch or hoard KV-cache blocks. Track queue depth and cache occupancy per adapter, and promote a tenant to its own replica group once its share of a shared replica stops being incidental. Fair queueing at admission is more effective than anything the engine offers internally.
  • What must a new replica do before it is ready to join an adapter hash ring?
    Load the base weights and warm up as usual, then fetch the adapters the ring has assigned it. Adapters are small — megabytes against tens of gigabytes — so this adds seconds, not minutes. The important part is gating readiness on it, so the router does not start sending adapter traffic to a replica that would have to load on the critical path of a user request.
  • Does adapter-affine routing conflict with prefix-affine routing?
    They compose, because they are keys at different granularities. Route on adapter id first to satisfy residency, then apply prefix or session affinity within the replica group that holds that adapter. If a tenant maps to only one replica the second key is moot; with a replication factor of two or three there is still room for prefix-aware choice inside the group.

saying these in an interview costs you the question

  • Runs one deployment per tenant adapter
  • Round-robins adapter traffic across every replica
  • Ignores the per-batch cap on distinct adapters
  • Assumes adapters are free to load on the request path
  • Tries to autoscale per adapter rather than the pool

context