skip to content

How do you choose the LoRA rank for a fine-tuning run on a 7B base model?

level: seniorimportance: should knowfreq 58%

answer

  1. capacity budget, not a quality dial
  2. count the trainable parameters it buys
  3. format is cheap, taxonomy is not
  4. plateau above full fine-tuning's loss
  5. the gap widens as data grows

basics

~20 s

Match rank to how much the model must absorb. A small rank around 8 carries a fixed response format or tone; tens to low hundreds are needed to absorb a domain's taxonomy or a post-training-scale dataset. Undersizing plateaus the loss; oversizing mainly costs memory.

solid answer

~50 s

Rank is a capacity budget, not a quality dial. Each adapted matrix gets `r * (d_in + d_out)` trainable parameters, so rank determines how much information the adapter can hold. For narrow behavioural changes — a house response format, a consistent tone, a fixed output schema learned from a few hundred examples — a rank around 8 is usually plenty. For absorbing a customer's domain taxonomy or a large instruction set, tens to low hundreds is the working range, and post-training-scale supervised runs are commonly done at rank in the low hundreds. Reinforcement-style fine-tuning sits at the other extreme: each episode carries very little information, so single-digit to low-double-digit ranks generally suffice. The diagnostic that matters is whether you are capacity-bound: sweep rank with everything else held fixed, and if loss keeps improving monotonically as rank rises, the adapter was too small. Oversizing costs memory, step time and checkpoint size, but rarely costs quality outright, so err upward when unsure.

code

python · 7 lines
python
def lora_params(r, layers=32, d=4096, ffn=11008):
    attn = 4 * r * (d + d)               # q, k, v, o
    mlp = 2 * r * (d + ffn) + r * (ffn + d)  # gate, up, down
    return layers * (attn + mlp)

for r in (8, 64, 256):
    print(f"rank {r:>3}: {lora_params(r)/1e6:6.1f}M trainable")

go deeper

for a junior

Know that rank sets how many trainable parameters the adapter gets, and that small ranks handle simple style or format changes while bigger jobs need more.

for a middle

Be able to compute the parameter count from rank and layer dimensions, and name concrete starting ranks for a format task versus a domain-absorption task.

for a senior

Demonstrate the diagnostic: a fixed-alpha rank sweep, a loss plateau sitting above full fine-tuning's, and a gap that widens with more data. Say which direction you err in and why.

for a principal

Own the framing that rank is a capacity budget measured against the information content of the dataset, and that oversizing is the cheap error while undersizing silently caps the ceiling of every downstream evaluation.

## What rank actually buys A LoRA adapter on a weight of shape d_out x d_in holds `r * (d_in + d_out)` trainable numbers. Rank is therefore a direct, countable capacity budget: it is the number of independent directions in which the adapter can push that layer's transformation. Everything about rank selection follows from treating it as capacity rather than as a quality setting. On a 7B-shaped model — hidden size 4096, feed-forward width around 11000, 32 layers — adapting every linear projection gives roughly 20 million trainable parameters at rank 8, around 160 million at rank 64, and around 640 million at rank 256. Those numbers are worth memorizing as anchors, because they turn an abstract knob into a comparison against the size of what you are teaching. ## Sizing by what the model must learn **Form and style: small rank.** Teaching a model to always answer in a fixed structure, adopt a house voice, refuse in a particular way, or emit a specific schema is a small, systematic change to the output distribution. A rank of 8 is typically sufficient, and pushing higher tends to buy nothing measurable. These runs also need very little data — a few hundred well-curated examples can carry a format. **Domain absorption: mid-to-high rank.** Teaching a model a customer's ticket taxonomy, an internal product vocabulary, or a large body of task-specific instruction behaviour is a substantially bigger change. Ranks in the tens work for a narrow taxonomy; a genuine post-training-scale supervised run over tens of thousands of examples is commonly done at rank in the low hundreds. The current practical guidance for supervised fine-tuning at that scale is a rank around 256, chosen so that adapter capacity is not the binding constraint. **Reinforcement-style fine-tuning: low rank.** When the training signal is a reward on sampled trajectories rather than a token-level target, each episode carries dramatically less information than a labeled example. The total information the run can transfer is small, so a rank in the single digits to low tens is usually enough, and large ranks mostly burn memory. ## The failure mode: capacity-bound training Undersizing has a recognizable signature. Training loss descends normally and then plateaus at a level noticeably above what full fine-tuning reaches on the same data, and — the tell that distinguishes it from ordinary underfitting — the gap *widens* as you add more training data. A capacity-bound adapter cannot use the extra data; a merely under-trained one can. The clean diagnostic is a rank sweep with every other setting frozen. If final loss improves monotonically as rank rises and only flattens at the top of your sweep, the earlier ranks were capacity-bound. If loss is flat across ranks from 8 upward, you found the plateau and the smallest rank on the plateau is the right one. Run the sweep with a fixed alpha rather than an alpha tied to rank, so the scaling factor is not confounding the comparison. ## The cost of oversizing Oversizing is the cheaper mistake. Its costs are real but bounded: more optimizer state and gradient memory, a slower step, a larger adapter checkpoint to store and ship, and a longer load time when adapters are swapped at serving. What it does *not* reliably do is degrade quality — a large adapter trained on a small dataset does not automatically overfit worse than a small one, because the base model's frozen weights still dominate the function. So when the dataset size is uncertain, biasing upward and confirming with a downward sweep is a defensible order of operations. ## What rank does not fix Two things are worth stating explicitly because they come up as follow-ups. First, rank does not compensate for adapting the wrong modules. An attention-only adapter at high rank underperforms an all-linear adapter at matched parameter count; where you put the capacity matters as much as how much you allocate. Second, rank does not fix LoRA's sensitivity to very large effective batches. Adapter training tolerates large batches less gracefully than full-parameter training does, and raising rank does not recover the loss — the two are independent axes, and the batch axis belongs to your optimizer settings, not your adapter configuration. ## A workable default procedure Start from the shape of the change: format-only, start at 8; domain behaviour on a few thousand examples, start at 32 or 64; a large supervised post-training set, start in the low hundreds; reinforcement-style, start low. Then run one downward or upward rank sweep at fixed alpha, watch whether the loss curve is still moving, and settle on the smallest rank sitting on the plateau. Report the rank alongside the dataset size — a rank quoted without a dataset size is uninterpretable.

  • How would you tell a capacity-bound adapter apart from one that is simply under-trained?
    Sweep rank with everything else fixed. A capacity-bound run improves monotonically as rank rises and its loss gap to full fine-tuning widens when you add more data, because it cannot absorb the extra signal. An under-trained run improves with more steps at the same rank instead, and its rank sweep is flat. The two diagnoses call for different fixes, so run the sweep before touching anything else.
  • Would raising the rank help if the run degrades at a large effective batch size?
    No. Adapter training tolerates very large effective batches worse than full-parameter training, and that sensitivity is independent of rank — raising rank does not recover it. The fix lies on the batching side, by keeping the effective batch modest rather than by spending adapter capacity. Treat them as separate axes and diagnose them separately, or you will spend a rank sweep chasing a batching problem.
  • Why does reinforcement-style fine-tuning need so much less rank than supervised fine-tuning?
    The amount of information transferred per training example differs by orders of magnitude. A supervised example supplies a target for every token; a reinforcement episode supplies roughly a single scalar reward for an entire trajectory. Total information transferred over the run is what the adapter must store, so a low rank is generally sufficient, and spending capacity there buys memory cost without buying learning.

saying these in an interview costs you the question

  • Treating higher rank as strictly higher quality
  • Picking rank without reference to dataset size
  • Confusing capacity-bound plateaus with overfitting
  • Assuming rank compensates for adapting the wrong modules
  • Sweeping rank while alpha moves with it

context