skip to content

Fine-Tuning

Adapting a base model to your data with LoRA or QLoRA adapters rather than full fine-tuning, and shaping behaviour with DPO or RLHF. The interview angle is when fine-tuning beats prompting plus retrieval — usually format, tone and task shape, rarely new facts.

on this pageshow

questions

6

In LoRA fine-tuning of Llama, what do rank r and lora_alpha control?

level: middleimportance: must knowfreq 72%

answer

  1. two small matrices beside frozen weights
  2. one knob is capacity, one is strength
  3. the scale divides by the rank
  4. B starts at zero, so training starts neutral
  5. q_proj, k_proj, v_proj, o_proj and the MLP trio

basics

~20 s

Rank r sets the width of LoRA's low-rank update matrices, fixing adapter capacity and the trainable-parameter count. lora_alpha scales that update — PEFT applies it as alpha/r — so raising alpha strengthens the adapter's effect without adding parameters.

solid answer

~50 s

LoRA freezes the Llama checkpoint and injects a trainable pair of matrices **A (d x r)** and **B (r x d)** beside chosen linear layers, so the effective weight becomes `W + (alpha/r) * B*A`. **r** is the bottleneck rank: it decides how much the adapter can express and directly sets the trainable parameter count — typical values are 8 to 64, and doubling r roughly doubles adapter size. **lora_alpha** is a fixed scaling constant, not learned; because the update is scaled by `alpha/r`, people usually set alpha to about 2x r so that changing r does not silently change the effective update magnitude. In PEFT you also pick `target_modules` — for Llama, the attention projections `q_proj`, `k_proj`, `v_proj`, `o_proj` and optionally the MLP projections `gate_proj`, `up_proj`, `down_proj`. Adapting more modules usually helps more than pushing r very high.

code

python · 13 lines
python
from peft import LoraConfig, get_peft_model

config = LoraConfig(
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
    task_type="CAUSAL_LM",
)

model = get_peft_model(base_model, config)
model.print_trainable_parameters()

go deeper

for a junior

Be able to say that LoRA freezes the original Llama weights and trains two small matrices next to them, so you train under one percent of the parameters and ship a small adapter file.

for a middle

Explain the W + (alpha/r) * B*A form, why B starts at zero, how r sets capacity and trainable-parameter count, and name the Llama projection layers you would put in target_modules.

for a senior

Show judgment on configuration: pick rank against dataset size, prefer broader module coverage over extreme rank, justify the higher learning rate LoRA tolerates, and explain how you detect overfitting on a held-out split.

for a principal

Own the tradeoff at fleet level — one base checkpoint with many per-task adapters versus divergent full copies, the versioning burden when the base model is upgraded, and how adapter rank interacts with training cost across dozens of tuning jobs.

## The problem LoRA solves A Llama checkpoint is a stack of transformer blocks whose parameters are large dense matrices. Updating every one of them (full fine-tuning) means storing a gradient and optimizer state for each parameter, which is why an 8B model needs well over a hundred gigabytes of accelerator memory to train conventionally. LoRA (Low-Rank Adaptation) sidesteps this: it leaves every original weight frozen and learns a small additive correction instead. ## The math, in plain terms For a frozen linear layer with weight matrix `W` of shape `d_out x d_in`, LoRA adds two small matrices: `A` of shape `r x d_in` and `B` of shape `d_out x r`, where `r` (the rank) is small — 8, 16, 32, 64. The layer now computes `y = Wx + (alpha/r) * B(Ax)`. `A` is initialised randomly and `B` is initialised to zeros, so at step zero `B*A` is exactly zero and the adapted model is bit-identical to the base model. Training only updates `A` and `B`. The parameter saving is dramatic. A 4096 x 4096 projection has ~16.8M parameters; its rank-16 LoRA pair has `16*4096 + 4096*16` = ~131K, under 1%. Across a whole Llama model, LoRA typically trains 0.1%-1% of the parameters. ## What rank r actually buys Rank is a capacity knob. The product `B*A` can only ever have rank r, so r caps how many independent directions the adaptation can move the weights in. Low ranks (4-8) suit narrow stylistic or format adaptation on a few thousand examples. Higher ranks (32-128) give the adapter more room for a genuinely different task or a large dataset, at the cost of more trainable parameters, more optimizer memory, and a greater chance of overfitting a small dataset. The empirical finding that has held up well is that **breadth beats depth**: adapting more module types at r=16 usually outperforms adapting only the query and value projections at r=128. ## What lora_alpha actually buys `lora_alpha` is not a learning rate and not a trainable parameter — it is a constant multiplier on the adapter's output. PEFT divides it by r, so the effective scale is `alpha/r`. That division exists so that r and the update magnitude are decoupled: if you tune at r=8, alpha=16 (scale 2.0) and then raise r to 32, keeping alpha at 2x r (alpha=64) keeps the scale at 2.0 and you only change capacity, not strength. If you instead hold alpha fixed while raising r, the effective update shrinks and your carefully tuned learning rate is now wrong. A variant, rank-stabilised LoRA (`use_rslora=True` in PEFT), scales by `alpha/sqrt(r)` instead, which behaves better at high ranks. ## Choosing target_modules on Llama Llama's decoder blocks expose these linear layers by name: the attention projections `q_proj`, `k_proj`, `v_proj`, `o_proj`, and the SwiGLU MLP projections `gate_proj`, `up_proj`, `down_proj`. A conservative configuration targets the four attention projections; a stronger one targets all seven. Adapting the embedding and output head is possible but expensive on Llama because the vocabulary is large, and it is only necessary if you added new special tokens. PEFT also accepts `target_modules="all-linear"` as a shorthand for every linear layer outside the head. ## Practical defaults and their consequences A workable starting point for instruction-style adaptation of an 8B Llama is r=16, lora_alpha=32, lora_dropout=0.05, all seven projections, learning rate around 1e-4 to 2e-4 — an order of magnitude higher than full fine-tuning uses, because you are training a small randomly-initialised module rather than nudging a pretrained one. Train for one to three epochs and watch a held-out split; LoRA overfits small datasets quickly at high rank. ## What LoRA does not change The base weights never move, so you keep the original checkpoint intact and ship a few tens of megabytes of adapter per task instead of a full model copy. Inference memory for the base model is unchanged — LoRA is a training-time saving plus a distribution convenience, not a way to make a 70B model fit where it did not fit before. And because the adapter is additive on top of a frozen model, catastrophic forgetting is far milder than with full fine-tuning, though it is not zero.

  • If you double r from 16 to 32, what should you do with lora_alpha and why?
    Double it too, to 64, so the `alpha/r` scale stays at 2.0. Otherwise the effective update magnitude halves and your tuned learning rate is silently mismatched to the new configuration. Keeping alpha at roughly 2x r means changing r changes only capacity, not update strength — which is the whole point of the alpha/r parameterisation.
  • Why is B initialised to zeros while A is initialised randomly?
    So the adapted model starts exactly equal to the base model: `B*A` is the zero matrix at step zero, so training begins from the pretrained behaviour rather than from a random perturbation. A must be random to break symmetry — if both were zero, the gradients would stay zero and nothing would learn.
  • You have a fixed parameter budget. Do you raise r or adapt more module types?
    Adapt more module types. Empirically, covering the attention projections plus the MLP projections at a modest rank beats a high rank applied only to q_proj and v_proj. The MLP layers hold a large share of a Llama block's parameters, so leaving them untouched limits what the adapter can express regardless of rank.
  • What does rank-stabilised LoRA change?
    With `use_rslora=True`, PEFT scales the adapter by `alpha/sqrt(r)` rather than `alpha/r`. The plain `alpha/r` scaling shrinks the update too aggressively as rank grows, so high-rank runs underperform; the square-root form keeps gradients better conditioned and makes large ranks actually pay off.

Think of the frozen Llama weights as a printed textbook you cannot edit. LoRA is a thin transparency sheet laid over each page: r is how much ink the sheet may hold, and alpha is how darkly that ink prints over the text underneath.

saying these in an interview costs you the question

  • Says LoRA updates the base Llama weights during training
  • Claims lora_alpha is the adapter's learning rate
  • Assumes higher rank is always better regardless of data size
  • Thinks LoRA reduces the base model's inference memory
  • Targets only q_proj and v_proj and calls it complete coverage

context

open as a page

What does QLoRA change compared with plain LoRA when fine-tuning Llama?

level: middleimportance: must knowfreq 60%

basics

~20 s

QLoRA loads the frozen Llama base in 4-bit NF4 instead of 16-bit, then trains ordinary LoRA adapters in bfloat16 on top. Weights are dequantized block by block during the forward and backward passes, cutting base-model memory roughly fourfold.

open as a page

When is fine-tuning Llama the wrong tool compared with prompting plus RAG?

level: principalimportance: must knowfreq 50%

basics

~20 s

Fine-tuning teaches form — output shape, tone, task convention, refusal behaviour. It is a poor way to install facts that change, because updating them means retraining. Retrieval handles changing knowledge; fine-tuning handles how the answer is produced.

open as a page

When should you merge a LoRA adapter into Llama's base weights?

level: middleimportance: should knowfreq 38%

basics

~20 s

Merge when one adapter will serve all traffic and you want zero adapter overhead — merging folds the low-rank update into the weights, producing a plain checkpoint. Keep it unmerged when you need to hot-swap several adapters over one shared base.

open as a page

When aligning a fine-tuned Llama, how does DPO differ from PPO-based RLHF?

level: seniorimportance: should knowfreq 44%

basics

~20 s

DPO trains directly on chosen/rejected preference pairs with a classification-style loss against a frozen reference model. PPO-based RLHF first trains a separate reward model, then optimizes the policy by sampling completions and scoring them — far more moving parts.

open as a page

Why does full fine-tuning of an 8B Llama need far more VRAM than its weights?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Training stores far more than weights: a gradient per parameter, two AdamW moment tensors, usually an fp32 master copy, plus activations. That is roughly 16 bytes per parameter, so an 8B model needs well over 100 GB before activation memory is counted.

open as a page