skip to content

Serving 200 per-tenant LoRA adapters on one base model: merge them or swap them?

level: principalimportance: should knowfreq 34%

answer

  1. one base resident, many small files
  2. count the checkpoints, not the adapters
  3. batching across tenants is the win
  4. merging is a graduation, not an optimization
  5. 200 adapters means 200 eval sets

basics

~20 s

Swap by default. Keeping adapters separate means one resident base serves every tenant, with small per-tenant files loaded on demand; merging bakes an adapter into the weights and produces a full-size model per tenant, which does not scale past a handful of dedicated deployments.

solid answer

~60 s

With adapters kept separate, one copy of the base occupies GPU memory and each tenant contributes only a small adapter — tens to hundreds of megabytes — that is applied as a parallel branch at inference. Serving stacks can batch requests for different tenants together while applying each request's own adapter, so 200 tenants genuinely share hardware. The costs are a modest per-layer compute overhead, a memory budget for the resident adapter cache, and cold-start latency when an adapter must be fetched and loaded. Merging removes that overhead entirely by folding the scaled adapter product into the base weights, but the merged artifact is a full-size model: 200 merged 7B models is 200 times the storage and cannot share a GPU. So merge only for the one or two tenants with dedicated capacity and a hard latency budget, and swap for the long tail. The larger question a principal should raise is whether 200 separate adapters are the right design at all, given 200 training runs, 200 eval sets and 200 rollback stories to own.

go deeper

for a junior

Know the difference: a swappable adapter stays a separate small file applied on top of the base, while a merged one is folded into the weights and produces a complete standalone model.

for a middle

Be able to do the storage arithmetic for many tenants and explain that one shared base plus small adapters is what allows tenants to share a GPU at all.

for a senior

Reason about the serving path end to end: batching across adapters, the resident adapter cache and its eviction policy, cold-start tail latency, and the mechanics of merging correctly including the scaling factor.

for a principal

Own the level above the question — whether per-tenant adapters are the right unit, given the training, evaluation and versioning burden they create, and what evidence would move you to segment adapters or retrieval instead.

## The two deployment shapes A trained adapter can live in one of two forms. **Separate (swappable).** The adapter stays as its own artifact. At inference the layer computes the frozen weight's output plus the scaled low-rank branch. The base is loaded once; adapters are attached per request. **Merged.** The scaled product of the adapter matrices is added into the base weight, producing a new full weight matrix. The result is architecturally indistinguishable from an ordinary fine-tuned model, and there is no adapter left to swap. For a single-tenant deployment the choice barely matters. For a multi-tenant product with a per-customer adapter over one shared base, it determines the entire cost structure. ## Why swapping is the default at 200 tenants The arithmetic is decisive. A 7B base is roughly 14 GB in 16-bit. An adapter at a moderate rank is tens to a few hundred megabytes. Keeping adapters separate means one 14 GB residency plus a pool of small files; merging means 200 separate 14 GB checkpoints, roughly three terabytes of storage, and — fatally — no sharing of GPU memory, since each merged model is a distinct set of weights that must be resident to serve. Modern serving stacks support attaching different adapters to different requests in the same batch, so tenants share not only the device but the batch. That is what makes the shared-base economics work: utilization does not fragment across tenants. Swapping also buys operational properties that matter more than they look. A tenant's adapter can be retrained and redeployed without touching the base or any other tenant. A bad adapter has a blast radius of exactly one customer. Rollback is replacing a small file. A/B testing a tenant's new adapter against their old one, or against the raw base, is a routing decision rather than a deployment. ## What swapping costs **Inference overhead.** Each adapted layer runs an extra pair of skinny matmuls. It is usually a small single-digit percentage of latency, but it is not zero and it grows with how many modules were adapted. **Adapter cache management.** The resident adapter pool is bounded by memory. With 200 tenants and a long-tailed traffic distribution you need an eviction policy, and an evicted tenant's next request pays a load. Cold-start latency is the thing that actually shows up in tail latency reports for low-traffic tenants, so pinning the top tenants and letting the tail churn is the usual compromise. **Scheduling complexity.** Batching across adapters is a solved problem in mainstream serving stacks but is not free — it constrains which requests can batch together and interacts with your queueing. ## When merging earns its cost Merge for a tenant that has all three of: enough traffic to justify dedicated capacity, a latency budget tight enough that the adapter branch matters, and a stable adapter that is not retrained frequently. That is typically a small number of accounts, and they are precisely the ones you would give dedicated infrastructure anyway. In that case merging is not really a serving optimization — it is the acknowledgment that this tenant has graduated to their own deployment. Two mechanical cautions. Merging must apply the same scaling factor the adapter was trained with, or you deploy a model nobody evaluated. And merging into a quantized base compounds error: merge into the higher-precision weights and quantize afterwards if needed, then re-run evaluation on the merged artifact rather than trusting the unmerged numbers. ## Stacking is not a free composition A tempting third option is composing adapters — a base behaviour adapter plus a tenant adapter. Adding two low-rank updates is arithmetically trivial but behaviourally unreliable: they were trained independently against the same base and their effects do not compose in any guaranteed way. Treat any stacked configuration as a distinct model requiring its own evaluation, and be sceptical of designs that assume compositionality holds. ## The question above the question The framing a principal should push back on is whether per-tenant adapters are the right unit at all. Two hundred adapters is 200 training runs to schedule, 200 evaluation sets to curate and keep honest, 200 artifacts to version, and 200 chances for a quiet quality regression that only one customer notices. The alternatives deserve a hearing: one multi-tenant adapter trained across all customers with tenant context supplied at inference; a small number of segment adapters covering clusters of similar customers; or retrieval over per-tenant data on a single shared model, which moves the per-customer state out of weights entirely and makes updates instantaneous. The honest answer is that this is contested and workload-dependent. Per-tenant adapters win when tenants genuinely differ in *form* — their taxonomy, their conventions, their output structure — and when the volume per tenant justifies the pipeline. They lose when the difference is mostly *facts*, which retrieval handles better, or when the tenant tail is long and low-volume, where the per-tenant training and evaluation overhead dwarfs the quality gain. Deciding which regime you are in, with evidence, is the actual senior judgment; merge-versus-swap is the easier question underneath it.

  • What is the practical downside of keeping adapters unmerged for a high-traffic tenant?
    Two things: the extra low-rank matmuls in every adapted layer add a small but persistent latency cost, and the adapter must be resident or paid for on load. For a tenant with a strict latency budget on dedicated capacity, folding the adapter into the weights removes both. For everyone else the overhead is far cheaper than the storage and utilization loss of a dedicated merged model.
  • Can you stack a shared behaviour adapter and a per-tenant adapter on the same base?
    Arithmetically yes — low-rank updates add. Behaviourally there is no guarantee they compose, since each was trained independently against the unmodified base and neither anticipated the other. Any stacked configuration is effectively a new model and needs its own evaluation. Designs that assume compositionality holds across many adapters tend to discover otherwise in production, on the combinations nobody tested.
  • What would make you argue against per-tenant adapters entirely?
    A long tail of low-volume tenants, or differences that are factual rather than formal. Two hundred adapters means 200 training runs, evaluation sets, versions and regression risks to own. If tenants differ mainly in the data they reference, retrieval on one shared model handles it with instant updates and no per-tenant pipeline. Reserve adapters for tenants whose required output form genuinely differs and whose volume pays for the machinery.

Swappable adapters are like one press with interchangeable plates; merging is casting a dedicated press for a single customer — worth it only for the customer whose volume keeps that press busy.

saying these in an interview costs you the question

  • Merging per tenant without counting checkpoint storage
  • Assuming merged models can share GPU memory
  • Forgetting the scaling factor when merging an adapter
  • Expecting independently trained adapters to compose reliably
  • Ignoring cold-start latency for the low-traffic tenant tail

context