In LoRA fine-tuning of Llama, what do rank r and lora_alpha control?
answer
- two small matrices beside frozen weights
- one knob is capacity, one is strength
- the scale divides by the rank
- B starts at zero, so training starts neutral
- q_proj, k_proj, v_proj, o_proj and the MLP trio
basics
~20 sRank r sets the width of LoRA's low-rank update matrices, fixing adapter capacity and the trainable-parameter count. lora_alpha scales that update — PEFT applies it as alpha/r — so raising alpha strengthens the adapter's effect without adding parameters.
solid answer
~50 sLoRA freezes the Llama checkpoint and injects a trainable pair of matrices **A (d x r)** and **B (r x d)** beside chosen linear layers, so the effective weight becomes `W + (alpha/r) * B*A`. **r** is the bottleneck rank: it decides how much the adapter can express and directly sets the trainable parameter count — typical values are 8 to 64, and doubling r roughly doubles adapter size. **lora_alpha** is a fixed scaling constant, not learned; because the update is scaled by `alpha/r`, people usually set alpha to about 2x r so that changing r does not silently change the effective update magnitude. In PEFT you also pick `target_modules` — for Llama, the attention projections `q_proj`, `k_proj`, `v_proj`, `o_proj` and optionally the MLP projections `gate_proj`, `up_proj`, `down_proj`. Adapting more modules usually helps more than pushing r very high.
code
python · 13 linesfrom peft import LoraConfig, get_peft_model
config = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
task_type="CAUSAL_LM",
)
model = get_peft_model(base_model, config)
model.print_trainable_parameters()go deeper
Be able to say that LoRA freezes the original Llama weights and trains two small matrices next to them, so you train under one percent of the parameters and ship a small adapter file.
Explain the W + (alpha/r) * B*A form, why B starts at zero, how r sets capacity and trainable-parameter count, and name the Llama projection layers you would put in target_modules.
Show judgment on configuration: pick rank against dataset size, prefer broader module coverage over extreme rank, justify the higher learning rate LoRA tolerates, and explain how you detect overfitting on a held-out split.
Own the tradeoff at fleet level — one base checkpoint with many per-task adapters versus divergent full copies, the versioning burden when the base model is upgraded, and how adapter rank interacts with training cost across dozens of tuning jobs.
## The problem LoRA solves A Llama checkpoint is a stack of transformer blocks whose parameters are large dense matrices. Updating every one of them (full fine-tuning) means storing a gradient and optimizer state for each parameter, which is why an 8B model needs well over a hundred gigabytes of accelerator memory to train conventionally. LoRA (Low-Rank Adaptation) sidesteps this: it leaves every original weight frozen and learns a small additive correction instead. ## The math, in plain terms For a frozen linear layer with weight matrix `W` of shape `d_out x d_in`, LoRA adds two small matrices: `A` of shape `r x d_in` and `B` of shape `d_out x r`, where `r` (the rank) is small — 8, 16, 32, 64. The layer now computes `y = Wx + (alpha/r) * B(Ax)`. `A` is initialised randomly and `B` is initialised to zeros, so at step zero `B*A` is exactly zero and the adapted model is bit-identical to the base model. Training only updates `A` and `B`. The parameter saving is dramatic. A 4096 x 4096 projection has ~16.8M parameters; its rank-16 LoRA pair has `16*4096 + 4096*16` = ~131K, under 1%. Across a whole Llama model, LoRA typically trains 0.1%-1% of the parameters. ## What rank r actually buys Rank is a capacity knob. The product `B*A` can only ever have rank r, so r caps how many independent directions the adaptation can move the weights in. Low ranks (4-8) suit narrow stylistic or format adaptation on a few thousand examples. Higher ranks (32-128) give the adapter more room for a genuinely different task or a large dataset, at the cost of more trainable parameters, more optimizer memory, and a greater chance of overfitting a small dataset. The empirical finding that has held up well is that **breadth beats depth**: adapting more module types at r=16 usually outperforms adapting only the query and value projections at r=128. ## What lora_alpha actually buys `lora_alpha` is not a learning rate and not a trainable parameter — it is a constant multiplier on the adapter's output. PEFT divides it by r, so the effective scale is `alpha/r`. That division exists so that r and the update magnitude are decoupled: if you tune at r=8, alpha=16 (scale 2.0) and then raise r to 32, keeping alpha at 2x r (alpha=64) keeps the scale at 2.0 and you only change capacity, not strength. If you instead hold alpha fixed while raising r, the effective update shrinks and your carefully tuned learning rate is now wrong. A variant, rank-stabilised LoRA (`use_rslora=True` in PEFT), scales by `alpha/sqrt(r)` instead, which behaves better at high ranks. ## Choosing target_modules on Llama Llama's decoder blocks expose these linear layers by name: the attention projections `q_proj`, `k_proj`, `v_proj`, `o_proj`, and the SwiGLU MLP projections `gate_proj`, `up_proj`, `down_proj`. A conservative configuration targets the four attention projections; a stronger one targets all seven. Adapting the embedding and output head is possible but expensive on Llama because the vocabulary is large, and it is only necessary if you added new special tokens. PEFT also accepts `target_modules="all-linear"` as a shorthand for every linear layer outside the head. ## Practical defaults and their consequences A workable starting point for instruction-style adaptation of an 8B Llama is r=16, lora_alpha=32, lora_dropout=0.05, all seven projections, learning rate around 1e-4 to 2e-4 — an order of magnitude higher than full fine-tuning uses, because you are training a small randomly-initialised module rather than nudging a pretrained one. Train for one to three epochs and watch a held-out split; LoRA overfits small datasets quickly at high rank. ## What LoRA does not change The base weights never move, so you keep the original checkpoint intact and ship a few tens of megabytes of adapter per task instead of a full model copy. Inference memory for the base model is unchanged — LoRA is a training-time saving plus a distribution convenience, not a way to make a 70B model fit where it did not fit before. And because the adapter is additive on top of a frozen model, catastrophic forgetting is far milder than with full fine-tuning, though it is not zero.
- If you double r from 16 to 32, what should you do with lora_alpha and why?Double it too, to 64, so the `alpha/r` scale stays at 2.0. Otherwise the effective update magnitude halves and your tuned learning rate is silently mismatched to the new configuration. Keeping alpha at roughly 2x r means changing r changes only capacity, not update strength — which is the whole point of the alpha/r parameterisation.
- Why is B initialised to zeros while A is initialised randomly?So the adapted model starts exactly equal to the base model: `B*A` is the zero matrix at step zero, so training begins from the pretrained behaviour rather than from a random perturbation. A must be random to break symmetry — if both were zero, the gradients would stay zero and nothing would learn.
- You have a fixed parameter budget. Do you raise r or adapt more module types?Adapt more module types. Empirically, covering the attention projections plus the MLP projections at a modest rank beats a high rank applied only to q_proj and v_proj. The MLP layers hold a large share of a Llama block's parameters, so leaving them untouched limits what the adapter can express regardless of rank.
- What does rank-stabilised LoRA change?With `use_rslora=True`, PEFT scales the adapter by `alpha/sqrt(r)` rather than `alpha/r`. The plain `alpha/r` scaling shrinks the update too aggressively as rank grows, so high-rank runs underperform; the square-root form keeps gradients better conditioned and makes large ranks actually pay off.
Think of the frozen Llama weights as a printed textbook you cannot edit. LoRA is a thin transparency sheet laid over each page: r is how much ink the sheet may hold, and alpha is how darkly that ink prints over the text underneath.
saying these in an interview costs you the question
- Says LoRA updates the base Llama weights during training
- Claims lora_alpha is the adapter's learning rate
- Assumes higher rank is always better regardless of data size
- Thinks LoRA reduces the base model's inference memory
- Targets only q_proj and v_proj and calls it complete coverage