What does QLoRA change compared with plain LoRA when fine-tuning Llama?
answer
- frozen base shrinks, trainable part does not
- a 4-bit type shaped for a bell curve
- the scaling constants get quantized too
- dequantize per block, compute in bf16
- optimizer state that can spill to host RAM
basics
~20 sQLoRA loads the frozen Llama base in 4-bit NF4 instead of 16-bit, then trains ordinary LoRA adapters in bfloat16 on top. Weights are dequantized block by block during the forward and backward passes, cutting base-model memory roughly fourfold.
solid answer
~50 sPlain LoRA keeps the frozen Llama weights in bf16 or fp16, which for a 70B model is still ~140 GB before you train anything. QLoRA quantizes those frozen weights to **4-bit NormalFloat (NF4)** — a data type shaped for normally-distributed weights — and dequantizes each block back to bf16 on the fly inside the matmul. The LoRA adapters themselves stay in bf16 and are what actually receives gradients, so the trainable part is unquantized and full-precision. Two extras come with it: **double quantization**, which quantizes the per-block quantization constants for a further ~0.4 bits per parameter, and **paged optimizers**, which spill optimizer state to host memory during gradient-checkpointing spikes instead of OOM-ing. The cost is throughput — dequantizing on every forward pass is slower than bf16 — and slightly noisier gradients. In `transformers` you enable it with `BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4")`.
code
python · 14 linesimport torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
bnb = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B",
quantization_config=bnb,
)go deeper
Know that QLoRA means loading the frozen Llama weights in 4 bits so a big model fits on one GPU, while the small LoRA adapters you train stay in normal precision.
Explain NF4 and per-block scaling, what double quantization recovers, what bnb_4bit_compute_dtype does, and why the base is only dequantized transiently inside the matmul.
Show that you budget a real run: weight memory versus activation memory, gradient checkpointing and paged optimizers for long sequences, the throughput penalty you accept, and how you validate quality against a bf16 baseline.
Own the economics — when buying more accelerator memory beats absorbing QLoRA's slower steps across many tuning jobs, and how the 4-bit training choice constrains the downstream merge-and-serve path for the whole fleet.
## Why QLoRA exists LoRA removes the optimizer and gradient memory for the base model, but it does not remove the base model itself. Holding a Llama 70B checkpoint in bf16 costs about 140 GB of accelerator memory before a single training step runs, which puts it out of reach of a single-node setup. QLoRA attacks the remaining term: the frozen weights. Since those weights are never updated, they do not need to be stored at training precision — only read accurately enough for the forward and backward passes. ## NF4: 4-bit NormalFloat NF4 is a 4-bit data type with 16 levels placed so that they are information-theoretically optimal for values drawn from a zero-centred normal distribution — which is a good description of pretrained transformer weights. Quantization is done per block (a small group of contiguous weights, commonly 64), each block carrying its own scaling constant, so outliers in one block do not wreck the resolution of its neighbours. At runtime, the kernel dequantizes a block back to the compute dtype (`bnb_4bit_compute_dtype`, normally bfloat16), does the matmul, and discards the dequantized copy. Memory is paid in 4-bit; arithmetic happens in 16-bit. ## Double quantization Per-block scales are themselves parameters: with block size 64 and a 32-bit scale, that is 0.5 extra bits per weight. Double quantization quantizes those constants too (to 8-bit, with a second-level scale over groups of them), recovering roughly 0.37 bits per parameter. On a 70B model that is several gigabytes — the difference between fitting and not fitting. Enabled with `bnb_4bit_use_double_quant=True`. ## Paged optimizers Gradient checkpointing recomputes activations on the backward pass, which produces sharp, transient memory spikes. Paged optimizers allocate optimizer state in NVIDIA unified memory so that pages can migrate to host RAM under pressure and back when needed, turning what would be a hard out-of-memory crash into a slowdown. This is what makes long-sequence QLoRA runs survivable on a single GPU. ## What stays in full precision This is the part interviewers probe. The **adapters are not quantized**. `A` and `B` are ordinary bf16 tensors with ordinary AdamW state; gradients flow through the dequantized base weights to reach them, but the base weights receive no update and are never written back. So QLoRA is 4-bit *storage* of a frozen model plus 16-bit *training* of a small module — not 4-bit training. ## What it costs Three costs, in rough order of how often they bite: 1. **Throughput.** Dequantizing on every forward pass adds work. QLoRA runs are commonly noticeably slower per step than bf16 LoRA on the same hardware; you are trading time for the ability to run at all. 2. **Quality.** The QLoRA authors reported that 4-bit NF4 plus LoRA can match 16-bit fine-tuning quality on their benchmarks, but that result is not a universal guarantee. On a demanding task with a small model, a bf16 LoRA baseline is worth measuring if you can afford it. 3. **Merging.** You cannot cleanly fold a bf16 adapter into a 4-bit base and keep the arithmetic exact. The normal path is to reload the base in bf16, attach the adapter, merge, and then quantize the merged model separately for serving if you want it quantized. ## Sizing intuition As a rough guide for LoRA-style training on Llama: an 8B model in 4-bit is about 5-6 GB of frozen weights, leaving room for adapters, activations and optimizer state inside a 24 GB card with gradient checkpointing and a short sequence length. A 70B model in 4-bit is roughly 35-40 GB of weights, which is why the QLoRA paper's headline was fine-tuning a 65B model on a single 48 GB GPU. Sequence length is the other dominant term — activation memory grows with it, and long-context runs blow the budget faster than parameter count does. ## Where it fits in a pipeline QLoRA is the default for supervised fine-tuning of Llama on constrained hardware, and the same 4-bit trick is routinely used for preference tuning afterwards. Tooling wraps it: Unsloth exposes it through `FastLanguageModel.from_pretrained(..., load_in_4bit=True)` with fused kernels, and Axolotl selects it from YAML with `adapter: qlora` and `load_in_4bit: true`. Underneath, all of them are calling bitsandbytes through the same `BitsAndBytesConfig`.
- Are the LoRA adapters themselves stored in 4 bits under QLoRA?No. Only the frozen base weights are quantized. The adapter matrices train in bfloat16 with normal optimizer state, and gradients reach them through the dequantized base weights. Calling QLoRA "4-bit training" is a mischaracterisation — it is 4-bit storage of a frozen model plus 16-bit training of a small trainable module.
- What does bnb_4bit_compute_dtype actually control?The precision the quantized weights are dequantized into for the matmul. Setting it to bfloat16 means blocks are expanded to bf16 before the multiply, then discarded. Leaving it at the default float32 wastes bandwidth and slows training with no quality benefit on hardware that supports bf16 well.
- Why is QLoRA slower per step than bf16 LoRA on the same GPU?Because every forward and backward pass dequantizes the 4-bit blocks back to the compute dtype before multiplying, which is extra work that bf16 weights never pay. QLoRA buys memory headroom at the price of throughput — you accept slower steps in exchange for fitting a model that otherwise would not fit at all.
- How does sequence length change your QLoRA memory budget?Sharply. Quantization only shrinks the weight term; activation memory scales with sequence length and batch size and is untouched by NF4. A run that fits comfortably at 1k tokens can OOM at 8k on the same card, which is why gradient checkpointing and paged optimizers matter most on long-context fine-tuning.
saying these in an interview costs you the question
- Claims QLoRA trains the adapters in 4-bit precision
- Says QLoRA updates the quantized base weights
- Thinks NF4 is just standard int4 with a different name
- Expects QLoRA to be faster than bf16 LoRA
- Plans to merge a bf16 adapter directly into the 4-bit base