skip to content

What does QLoRA change about a LoRA run, and why does it fit on one GPU?

level: middleimportance: must knowfreq 62%

answer

  1. the base weights are the leftover term
  2. four bits, frozen, dequantized per matmul
  3. adapters stay at training precision
  4. error is introduced once, not compounded
  5. you pay in step time

basics

~20 s

QLoRA stores the frozen base model in 4-bit instead of 16-bit and trains 16-bit adapters on top of it. Since the base never updates, its low precision is static; cutting weight memory roughly fourfold is what brings a multi-GPU run onto a single card.

solid answer

~60 s

Ordinary LoRA already removes gradients and optimizer state for the frozen parameters, but the base weights themselves still sit in memory at 16 bits — about 14 GB for a 7B model, 140 GB for a 70B one. QLoRA attacks that remaining term by holding the frozen base in a 4-bit representation, dequantizing blocks back to higher precision on the fly as each matmul needs them, and keeping the trainable adapters in bf16. Because the base is frozen, its quantization error is fixed at load time rather than compounding through updates — nothing in the base is ever written back. The result is roughly a fourfold cut in weight memory, which is what turns a 7B run into something a single consumer-class card can hold and a 70B run into a single high-memory card job. Two supporting tricks travel with it: quantizing the quantization constants themselves to reclaim a little more, and paging optimizer state to host memory so transient spikes do not abort the run. You pay for it in step time, since every forward and backward pass dequantizes on the fly.

code

python · 4 lines
python
for name, params in [("7B", 7e9), ("13B", 13e9), ("70B", 70e9)]:
    bf16 = params * 2 / 1e9
    four_bit = params * 0.5 / 1e9
    print(f"{name:>4}: {bf16:6.1f} GB bf16 -> {four_bit:5.1f} GB at 4-bit")

go deeper

for a junior

Know that QLoRA means the frozen base model is stored at 4 bits while the trained adapters stay at normal precision, and that the point is fitting the run into less GPU memory.

for a middle

Be able to do the arithmetic — roughly 2 bytes per parameter down to roughly half a byte — and explain that dequantization happens per block during the forward and backward passes.

for a senior

Explain why freezing makes low precision tolerable, name the step-time cost you accept, and describe the merge path that avoids compounding quantization error on deployment.

for a principal

Own the decision framing: this is a capacity unlock, not a default. Weigh slower steps and extra operational coupling against the alternative of claiming more accelerators, and record the quantization setup as part of the adapter's contract.

## The memory term LoRA does not remove Standard LoRA removes gradients, optimizer moments and master copies for the frozen parameters. What it cannot remove is the frozen model itself: every weight must be resident to run the forward pass. At 16 bits per parameter that is roughly 2 GB per billion parameters — about 14 GB for a 7B model, 26 GB for a 13B, 140 GB for a 70B. Once optimizer memory is gone, the base weights become the dominant term, and for larger models they alone exceed a single accelerator's memory. QLoRA is the direct answer: hold the frozen base at 4 bits instead of 16. ## The mechanism Three pieces do the work. **A 4-bit frozen base.** The pretrained weights are quantized once, at load time, into a 4-bit representation and never updated. During training, each weight block is dequantized back to a compute precision immediately before its matrix multiplication and discarded afterwards, so the high-precision copy exists only transiently for the block being used. Weight memory drops by roughly a factor of four. **Higher-precision adapters.** The LoRA matrices stay in bf16 and are the only trainable tensors. Gradients flow backward *through* the quantized frozen layers to reach them — the frozen weights participate in the backward computation but receive no update. **Memory hygiene around the edges.** Quantization stores per-block scaling constants, which themselves consume memory; quantizing those constants in turn reclaims a further slice. Separately, optimizer state can be paged out to host memory so that transient allocation spikes — long sequences, an unlucky batch — do not abort a run that would otherwise fit. ## Why freezing makes low precision tolerable This is the conceptual point worth making in an interview. Quantization error in a *trainable* weight compounds: you quantize, update, requantize, and the rounding interacts with the optimization trajectory. In QLoRA the base is frozen, so its quantization error is introduced exactly once, at load, and stays fixed for the whole run. The trainable path — the adapters — is at full training precision throughout. The optimization problem is therefore a clean bf16 problem over a small parameter set, applied on top of a slightly perturbed but static function. ## What it costs **Step time.** Dequantizing blocks on every forward and backward pass is real work that a bf16 run does not do. Expect meaningfully slower steps than plain LoRA on the same hardware — you are trading throughput for the ability to run at all. **Merge friction.** Adding the scaled adapter product back into a 4-bit base is awkward: you would be merging a high-precision delta into a low-precision weight and then requantizing, which introduces a second round of error. The clean path is to merge into the original 16-bit weights and, if a quantized deployment artifact is wanted, quantize afterwards. **Operational coupling.** The adapter was trained against a specific quantized view of the base. Serving it against a differently quantized base is a configuration mismatch waiting to surprise you, so the quantization setup belongs in the adapter's recorded metadata alongside rank and scaling. ## When to reach for it QLoRA is a memory-constrained choice, not a default. If the run already fits in bf16 with room for your sequence length and batch, plain LoRA is faster and simpler. Reach for the 4-bit base when the base weights are what is blocking you: a large model on a single card, a shared cluster where you can claim only one device, or a development loop on a workstation. The decision is a capacity decision first — does the run fit at all — and everything else follows from that. ## The economics, concretely At roughly 2 bytes per parameter in bf16 and roughly 0.5 bytes plus small per-block overhead in 4-bit, a 7B base moves from about 14 GB to under 4 GB, a 13B from about 26 GB to around 7 GB, and a 70B from about 140 GB to roughly 35-40 GB. Add adapter parameters, their optimizer state, activations and workspace, and the qualitative shift is that models which needed several accelerators for adapter training become single-device jobs. That accessibility shift, more than any quality argument, is why the technique became the standard fallback whenever memory is the binding constraint.

  • If the base is quantized, why does its low precision not degrade the optimization itself?
    Because the base never receives an update. Its quantization error is introduced once at load time and stays constant, so it perturbs the function being adapted but does not interact with the optimizer's trajectory. The trainable path — the adapter matrices — is held at normal training precision throughout, so gradients and optimizer moments are computed and stored without extra rounding.
  • How would you merge a QLoRA adapter back into the base weights for deployment?
    Merge into the original 16-bit weights, not the 4-bit ones. Adding a high-precision delta into a quantized weight and requantizing introduces a second round of error on top of the first. Load the unquantized base, add the scaled product of the adapter matrices, then quantize the merged result if a low-precision artifact is what you want to serve. Re-evaluate afterwards, because the merged-and-requantized model is not the model you trained against.
  • When is plain LoRA the better choice over QLoRA?
    Whenever the run already fits in higher precision. The 4-bit base buys memory at the cost of dequantization work on every forward and backward pass, so steps are slower for no benefit if memory was never the constraint. Treat it as the fallback that unblocks a run you otherwise could not launch, rather than as a default configuration to apply everywhere.

saying these in an interview costs you the question

  • Thinking the adapters are what gets quantized
  • Assuming quantization error compounds through training updates
  • Expecting QLoRA to speed up training as well as save memory
  • Merging an adapter directly into the 4-bit base weights
  • Treating a 4-bit base as a default rather than a memory fallback

context