In LoRA fine-tuning, what does the low-rank update train while the base stays frozen?
answer
- the pretrained weights never move
- two skinny matrices beside each layer
- one of the pair starts at zero
- r times the two dimensions, summed
- optimizer state is the real saving
basics
~20 sLoRA leaves every pretrained weight frozen and trains two small matrices per targeted layer, A and B. Their product BA is added to the frozen weight as a rank-r update, so only those matrices hold gradients and optimizer state.
solid answer
~50 sLoRA assumes the *change* you need to make to a pretrained weight is low-rank, so instead of learning a full delta it learns a factorization. For a weight matrix W of shape d_out x d_in, LoRA freezes W and injects two trainable matrices: A of shape r x d_in and B of shape d_out x r, with r far smaller than either dimension. The layer computes W x + (alpha/r) * B (A x). A is initialized randomly and B to zeros, so at step zero the adapter contributes nothing and the model behaves exactly like the base. Trainable parameters drop from d_in * d_out to r * (d_in + d_out) per adapted matrix — often well under one percent of the model. The base weights still occupy memory, but gradients, optimizer moments and fp32 master copies now exist only for the adapter, which is what turns a multi-GPU job into a single-card one.
code
python · 13 linesimport numpy as np
d_in, d_out, r, alpha = 4096, 4096, 8, 16
W = np.random.randn(d_out, d_in) * 0.01 # frozen pretrained weight
A = np.random.randn(r, d_in) * 0.01 # trainable
B = np.zeros((d_out, r)) # trainable, zero-init
x = np.random.randn(d_in)
base = W @ x
adapted = base + (alpha / r) * (B @ (A @ x))
print("identical at step 0:", np.allclose(base, adapted))
print("trainable:", A.size + B.size, "frozen:", W.size)go deeper
Be able to say plainly that the pretrained weights are frozen and two small matrices are trained beside them, and that the result is a small adapter file rather than a new full model.
Explain the factorization itself: shapes r x d_in and d_out x r, the parameter count r times the sum of the dimensions, the zero-initialized side, and why the model is unchanged at step zero.
Show where the memory actually goes. An interviewer expects you to separate weights from gradients and optimizer state, note that activations are unaffected, and predict that time savings are far smaller than memory savings.
Own the framing that rank is a capacity budget, not a quality setting. Be ready to argue when the low-rank assumption holds for your workload and what evidence would tell you it has been exceeded.
## Why full fine-tuning is expensive Updating every weight of a pretrained transformer costs far more memory than simply holding the model. With a standard adaptive optimizer you pay for the weights themselves, a gradient for every weight, and roughly two optimizer moments per weight — usually kept in 32-bit, often alongside a 32-bit master copy of the weights. For a 7-billion-parameter model held in bf16 that is about 14 GB of weights and several times that again in gradients and optimizer state, before activations. A model that serves comfortably on one accelerator becomes a multi-GPU training job. Parameter-efficient fine-tuning (PEFT) is the family of methods that avoid this bill by training a small number of new parameters instead of all the old ones; LoRA — low-rank adaptation — is the dominant member of that family. ## The decomposition The starting observation is that the difference between a pretrained model and a task-adapted one appears to be a *low-rank* change, not an arbitrary full-rank one. A rank-r matrix of shape d_out x d_in can be written as the product of a d_out x r matrix and an r x d_in matrix. LoRA exploits this directly: pick the weight matrices you want to adapt, freeze each one, and attach a pair of new matrices beside it. For a frozen weight W, LoRA learns A (shape r x d_in) and B (shape d_out x r), and the layer's output becomes W x + s * B (A x), where s is a fixed scalar. Nothing about W changes; the adapter runs as a parallel branch whose output is summed into the original. Because the branch is applied to the same input x, it can be folded back into W after training by simply adding s * B A to it. The parameter arithmetic is the whole point. A 4096 x 4096 projection holds about 16.8 million numbers. Its rank-8 adapter holds 8 * (4096 + 4096) = 65,536 — about 0.4% of the original matrix. Applied to the query and value projections of a 7B model, that is roughly four million trainable parameters against seven billion frozen ones: well under a tenth of a percent. ## Initialization A is drawn from a small random distribution and B is initialized to zeros. That asymmetry matters twice over. First, B A = 0 at initialization, so the adapted model is numerically identical to the base model on step one; training can only ever move away from a known-good starting point, and a run that diverges cannot be blamed on a random perturbation of the pretrained function. Second, the pairing is what lets learning start at all: with B zero but A random, B still receives a nonzero gradient (it sees the nonzero activations A x), and once B moves off zero, A begins receiving gradient too. Initialize *both* to zero and the gradients of both are identically zero forever — the adapter never leaves the origin. ## Where the memory savings actually come from A common misreading is that LoRA shrinks the model. It does not. The full base model is resident in memory during training and during inference; LoRA changes only which tensors are *trainable*. The savings are: - **No gradient buffer** for the billions of frozen parameters. - **No optimizer state** for them — usually the largest single term, since adaptive optimizers keep two moments per trainable parameter. - **Small checkpoints.** A trained adapter is tens or hundreds of megabytes rather than tens of gigabytes, which is what makes per-task and per-tenant adapters practical to store and ship. Activations are *not* saved by LoRA in general: gradients must still flow backward through the whole network to reach adapters in the early layers, so activation memory and the backward pass through frozen layers remain. This is why LoRA reduces training memory dramatically but reduces training *time* only modestly. ## The low-rank hypothesis and where it breaks The adapter's capacity is bounded by r. Teaching a model a house response format, a tone, a fixed output schema, or a routing behaviour is a small change and a small rank carries it. Absorbing a large body of new domain material is a bigger change, and if the information in the dataset exceeds what r * (d_in + d_out) parameters per matrix can represent, the run becomes capacity-bound: the loss plateaus above what full fine-tuning would reach, and the gap widens as you add data. That is the honest boundary of the hypothesis — not that LoRA is a quality compromise in general, but that a *given rank* is a capacity budget. ## What it does not change LoRA does not alter the architecture, the tokenizer, or the base model's knowledge. It does not make inference cheaper — if anything an unmerged adapter adds one extra small matmul per adapted layer. And because the base is frozen, whatever the base gets wrong is still there unless the adapter can steer around it.
- Why is B initialized to zeros while A is random, rather than both being random?Zero-initializing B makes the adapter's contribution exactly zero at step one, so the adapted model starts identical to the pretrained model and every change is deliberate. Random-initializing both would perturb a known-good function before any learning happens. Zero-initializing both would be worse still: the gradients of A and B would each be zero, and the adapter would never leave the origin.
- Does an unmerged LoRA adapter add inference latency?Yes, a small amount: each adapted layer runs an extra pair of skinny matmuls plus an addition. It is usually a few percent, but it is not free, and it scales with how many modules you adapted. Adding the scaled product of the adapter matrices back into the base weight removes the overhead entirely, at the cost of no longer being able to swap the adapter out.
- Does LoRA reduce the memory needed to hold the base model during training?No. The full base is resident, in whatever precision you loaded it. LoRA removes the gradient buffer and optimizer state for the frozen parameters, which is typically the largest term in a full fine-tuning run. Activation memory for the backward pass also remains, since gradients still flow through the whole network to reach adapters in the early layers.
Think of it as a transparent overlay on a printed map rather than reprinting the map: the original is untouched, the overlay is tiny, and you can peel it off, swap it, or press it permanently into a reprint.
saying these in an interview costs you the question
- Saying LoRA compresses or shrinks the base model
- Claiming LoRA trains a selected subset of the original weights
- Initializing both adapter matrices randomly, or both to zero
- Assuming LoRA speeds up training as much as it cuts memory
- Believing the adapter can be dropped in without the base model