skip to content

When should you merge a LoRA adapter into Llama's base weights?

level: middleimportance: should knowfreq 38%

answer

  1. fold the low-rank product into the weights
  2. cheap at runtime, expensive on disk
  3. many behaviours on one resident base
  4. never fold bf16 into four bits
  5. re-evaluate the artifact you actually ship

basics

~20 s

Merge when one adapter will serve all traffic and you want zero adapter overhead — merging folds the low-rank update into the weights, producing a plain checkpoint. Keep it unmerged when you need to hot-swap several adapters over one shared base.

solid answer

~50 s

Merging computes `W + (alpha/r) * B*A` once and writes the result into the weight matrices, giving you an ordinary Llama checkpoint with no PEFT wrapper — in PEFT that is `PeftModel.merge_and_unload()` followed by `save_pretrained()`. The upside is a serving path that needs no adapter support and no per-layer extra matmuls. The downsides are that you now store a full model copy per task instead of a small adapter file, and you lose the ability to serve many task adapters over one shared base. **Merge in a float dtype, not into a 4-bit base**: if you trained with QLoRA, reload the base in bf16, attach the adapter, merge there, and quantize the merged model afterwards if you want it quantized — folding a bf16 adapter into 4-bit storage loses precision that shows up as quality drift. Also merge exactly one adapter per checkpoint unless you have deliberately validated the combination.

code

python · 12 lines
python
import torch
from transformers import AutoModelForCausalLM
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B-Instruct",
    dtype=torch.bfloat16,   # unquantized, even if you trained with QLoRA
)

model = PeftModel.from_pretrained(base, "out/llama-lora")
merged = model.merge_and_unload()
merged.save_pretrained("out/llama-merged")

go deeper

for a junior

Know that a LoRA adapter can either be kept as a small separate file loaded on top of Llama, or folded permanently into the weights to produce one ordinary checkpoint.

for a middle

Explain that merging computes the low-rank product once and adds it to the weights, name merge_and_unload, and state the storage-versus-runtime-overhead tradeoff.

for a senior

Demonstrate the operational care: merge in a float dtype rather than into a 4-bit base, quantize afterwards, re-evaluate the merged artifact, and keep the base and adapter for reproducibility.

for a principal

Own the serving architecture — one resident base with per-request adapters against a fleet of merged checkpoints, and what each choice implies for GPU memory, storage cost, rollout speed and rollback.

## What merging actually does During training and unmerged inference, a LoRA-adapted layer computes `y = Wx + (alpha/r) * B(Ax)` — two extra small matmuls per adapted layer. Merging evaluates the constant part `(alpha/r) * B*A` once, adds it into `W`, and deletes the adapter modules. The arithmetic is exact in principle; in practice the result depends on the dtype you merge in, because you are adding a small matrix to a large one and rounding the sum. After merging you hold a plain Llama checkpoint. Anything that loads Llama loads it — no PEFT dependency, no adapter file, no runtime that must understand adapters. ## The case for merging - **One model, one job.** If the deployment serves a single fine-tuned behaviour, an adapter buys nothing at runtime and costs a little latency and complexity. - **Runtime compatibility.** Some inference paths and conversion pipelines accept only ordinary checkpoints. Quantizing a fine-tuned model for local inference generally starts from merged weights. - **Latency.** The extra per-layer matmuls are small but not free, and they interact badly with some fused-kernel implementations. ## The case against merging - **Storage.** An adapter for an 8B Llama is typically tens to a few hundred megabytes; a merged checkpoint is the full 16 GB in bf16. Ten task adapters versus ten full copies is a 100x difference in what you store and ship. - **Multi-tenancy.** Serving runtimes that support adapters can hold one base model resident and apply a different adapter per request, which is how you serve many tuned behaviours on one GPU. Merging forfeits that entirely. - **Iteration.** Adapters are cheap to version, diff and roll back; full checkpoints are not. ## The QLoRA merge trap This is the part that bites people. If you trained with QLoRA, the base in memory is 4-bit NF4 while the adapter is bf16. Adding a bf16 update into 4-bit storage cannot be represented faithfully — either the operation is refused or the result is quantized again and loses fidelity. The correct sequence is: 1. Load the **original** base checkpoint in bf16 or fp16, unquantized. 2. Attach the trained adapter with `PeftModel.from_pretrained`. 3. Call `merge_and_unload()`. 4. Save the merged model. 5. Quantize the merged model separately if the serving target needs it. A common symptom of getting this wrong is a merged model that is subtly worse than the adapter-attached model was during evaluation — always re-evaluate the merged artifact rather than trusting the pre-merge numbers. ## Stacking several adapters PEFT can load multiple adapters and combine them, weighted, into a single set of weights. This sometimes works and sometimes produces a model that is worse at both tasks, because the two low-rank updates were fitted independently against the same frozen base and nothing constrained them to be compatible. Treat any combination as an experiment requiring its own evaluation, not as composition that is guaranteed to hold. ## Reversibility Merging is destructive to the loaded model object but not to your artifacts: you still have the base checkpoint and the adapter file, so you can always redo the merge differently. PEFT also offers `unmerge_adapter()` to back the update out of a still-attached adapter, which subtracts the same product — but this is a within-session convenience, and the durable safety net is simply keeping the base and the adapter on disk. Never delete the adapter after merging. ## The decision in one line Merge for a dedicated single-purpose deployment or a conversion pipeline; keep adapters separate when you serve several behaviours, iterate frequently, or care about storage. And whichever you choose, benchmark the exact artifact you will deploy, because a merged model is not automatically identical in quality to the adapter-attached model you evaluated.

  • You trained with QLoRA on a 4-bit base. Why not merge straight into that loaded model?
    Because the adapter update is bf16 and the base is 4-bit NF4 — the sum cannot be stored faithfully at 4-bit resolution, so you either hit a refusal or silently lose precision. Reload the original base unquantized, attach the adapter, merge there, and quantize the merged checkpoint afterwards if the serving target needs it.
  • What do you give up by merging when you run several fine-tuned behaviours?
    Multi-tenancy. Unmerged, one base model stays resident and a different adapter is applied per request, so many tuned behaviours share one GPU's weights. Merged, each behaviour is a full checkpoint that must be loaded separately — far more memory and far more storage for the same set of tasks.
  • Is a merged model guaranteed to score identically to the adapter-attached model you evaluated?
    No. The merge adds a small matrix into a large one and rounds the result in the merge dtype, so tiny numerical differences are possible, and a merge done at the wrong precision can differ noticeably. Always re-run your evaluation against the exact merged artifact you intend to deploy rather than carrying over the pre-merge numbers.
  • Can you merge two adapters trained separately on the same base?
    Mechanically yes — PEFT can combine loaded adapters with weights — but there is no guarantee of quality. The two low-rank updates were fitted independently against the same frozen weights and nothing made them compatible, so the combination is often worse at both tasks. Treat it as an experiment that needs its own evaluation.

saying these in an interview costs you the question

  • Merges a bf16 adapter directly into a 4-bit quantized base
  • Deletes the adapter file once the merge is done
  • Assumes merging reduces the model's inference memory
  • Expects merged quality to match pre-merge numbers without re-testing
  • Thinks stacking two adapters composes their skills reliably

context