skip to content

How do you serve several LoRA adapters from one vLLM server and select one per request?

level: seniorimportance: should knowfreq 45%

answer

  1. Base weights loaded once, deltas per request
  2. Adapters look like models to clients
  3. Selected through the ordinary model field
  4. Rank ceiling is preallocated memory
  5. Runtime load is gated by an env var

basics

~20 s

Start with --enable-lora and register adapters via --lora-modules name=path. Each name appears in /v1/models, and a request selects one by putting that name in its model field. Adapters must all target the base checkpoint the server loaded.

solid answer

~40 s

`vllm serve <base-model> --enable-lora --lora-modules sql=/models/sql-lora support=/models/support-lora` loads the base weights once and keeps the adapters alongside them. The adapter names are advertised at `/v1/models` and are addressed through the ordinary `model` field, so a client picks an adapter exactly the way it picks a model — the base model stays addressable under its own name. Two sizing flags matter: `--max-lora-rank` must be at least the largest adapter's rank and is preallocated, so raising it costs memory whether or not you use it; `--max-loras` caps how many distinct adapters may appear in a single batched step. Adapters can also be added and removed at runtime by posting to `/v1/load_lora_adapter` and `/v1/unload_lora_adapter`, which are only enabled when `VLLM_ALLOW_RUNTIME_LORA_UPDATING` is set. Every adapter must have been trained against the exact base checkpoint being served.

code

bash · 5 lines
bash
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-lora \
  --lora-modules sql=/models/sql-lora support=/models/support-lora \
  --max-lora-rank 16 \
  --max-loras 4

go deeper

for a junior

Know that vLLM can serve LoRA adapters on top of one base model, that --enable-lora turns this on, and that a client picks an adapter by name in the same model field it already uses.

for a middle

Explain the launch shape: --lora-modules name=path, adapters listed in /v1/models, and the two sizing flags — the rank ceiling that is preallocated and the cap on distinct adapters per batched step.

for a senior

Cover operations: runtime load and unload behind VLLM_ALLOW_RUNTIME_LORA_UPDATING as a control-plane surface, the hard adapter-to-base pairing across revisions, and the real per-token overhead a LoRA-enabled server carries.

for a principal

Own the naming and artifact policy: adapter names are a public contract clients hardcode, adapter and base form one versioned pair, and the boundary between the inference audience and whoever may load weights into a running GPU process must be explicit.

## What the server is doing A LoRA adapter is a small set of low-rank weight deltas trained against a specific base checkpoint. Serving one does not require materializing a whole new model: vLLM loads the base weights once and applies the selected adapter's deltas during the forward pass, per request. That is what makes many adapters on one replica economically interesting — the expensive part, the base weights in GPU memory, is shared, and each adapter adds only megabytes. From the API's point of view the trick is that adapters are exposed as *models*. `--enable-lora` turns the feature on; `--lora-modules` registers adapters as `name=path` pairs (a JSON form is also accepted when you need to be explicit about the base model name). Each registered name then shows up in `GET /v1/models` next to the base model, and a request routes to it by setting `"model": "sql"`. No new endpoint, no custom header, no separate port. Clients that already know how to pick a model already know how to pick an adapter. ## The two capacity flags `--max-lora-rank` declares the largest adapter rank the server will accept. Rank is the width of the low-rank decomposition and is fixed at training time; the server preallocates buffers sized for this value, so it is not free. Set it below an adapter's actual rank and that adapter is rejected; set it far above (say 64 when every adapter is rank 8) and you have paid memory for nothing. Take the maximum over the adapters you intend to serve, and remember that admitting a higher-rank adapter later is a restart. `--max-loras` is different: it caps how many *distinct* adapters may participate in one batched step. It does not cap how many are registered. With a low value, requests for adapters beyond the cap wait for a step in which a slot frees, which shows up as latency for the unlucky adapters rather than as an error. Raising it lets more adapters share a batch at some throughput cost, because the batched matrix work becomes more heterogeneous. There is also a CPU-side cache (`--max-cpu-loras`) that holds adapters not currently resident on the GPU, so swapping one in does not mean re-reading it from disk. ## Runtime add and remove Registering everything at launch is fine when the adapter set is static, but a per-tenant fine-tuning product ships new adapters continuously and cannot restart a GPU process for each one. vLLM exposes `POST /v1/load_lora_adapter` with a `lora_name` and `lora_path`, and `POST /v1/unload_lora_adapter` to drop one. These routes are gated behind the `VLLM_ALLOW_RUNTIME_LORA_UPDATING` environment variable and are off by default, for good reason: a route that makes the server read a filesystem path and load weights into a running process is an obvious abuse target, so it must never be reachable from untrusted callers. Treat it as a control-plane operation, on a private network path, behind your own authorization. ## The constraints people trip over **Same base, exactly.** An adapter's deltas are shaped for a specific checkpoint. Point a server at a different base — even a different revision or a differently quantized build of the "same" model — and the adapter is at best rejected, at worst semantically wrong. The pairing of adapter and base is a versioned artifact. **One base per process.** All adapters on a server share the one base model. Serving adapters for two different bases means two servers. **LoRA is not free at inference.** Applying deltas adds per-token work and heterogeneity to the batch, so a LoRA-enabled server is measurably slower than the same server without it. That cost is what you pay for not standing up one replica per adapter, and it is usually an easy trade at any meaningful adapter count — but do measure it rather than assuming the overhead is zero. **Name collisions matter.** Adapter names live in the same namespace as the served model name, since both are addressed by `model`. Choose names deliberately, because they are now a public API contract that clients hardcode. ## Operating it Make `/v1/models` your source of truth in health checks — it tells you which adapters this replica currently has, which is the question that matters after a rolling restart or a runtime load. Log the requested `model` value so you can attribute traffic per adapter. And keep adapter artifacts wherever your model weights live, with the base checkpoint id recorded next to them, so a restart re-registers exactly the pairs that were serving before.

  • What is the difference between --max-loras and --max-lora-rank?
    `--max-lora-rank` is a shape ceiling: the largest adapter rank the server will accept, with buffers preallocated for it, so setting it high costs memory even if unused. `--max-loras` is a concurrency ceiling: how many distinct adapters may appear in a single batched step. Registering more adapters than that is fine; the excess simply waits for a slot.
  • Why are the runtime adapter load and unload endpoints disabled by default?
    Because they make the server read a caller-supplied path and load weights into a live process. That is a remote code and data ingestion surface, so vLLM requires VLLM_ALLOW_RUNTIME_LORA_UPDATING to be set before the routes exist. Treat them as control plane: private network path, your own authorization in front, never exposed to the same audience as the inference routes.
  • A request names an adapter that this replica does not have. What does the client see?
    The same failure as naming a model that is not served — a rejected request, because the `model` field is validated against what the process advertises. That is why `/v1/models` matters operationally: after a rolling restart or a runtime load, replicas can legitimately disagree about which adapters they carry, and a request must reach one that has the adapter.

saying these in an interview costs you the question

  • Thinks each adapter needs its own server process
  • Expects a custom header or body field to pick an adapter
  • Assumes adapters work against any base checkpoint
  • Believes LoRA serving adds no inference overhead
  • Confuses the rank ceiling with the concurrency cap

context