Why does a LoRA fine-tune need a higher learning rate than full fine-tuning?
answer
- depends on what is trainable
- adapter versus full-weight run
- fewer parameters, bigger steps
- about an order of magnitude apart
- 1e-5 full, 1e-4 to 2e-4 adapter
basics
~20 sA LoRA run trains only a small adapter, so each step must move far fewer parameters and needs a bigger step size. The working rule is roughly ten times the full fine-tuning rate — about 1e-4 versus 1e-5.
solid answer
~50 sFull fine-tuning updates every weight in the model, so even a small step changes a lot of capacity; typical rates for a 7–8B model are **1e-5 to 2e-5**. LoRA freezes the base and trains a small low-rank adapter, so that same nominal rate barely moves it — the run underfits and the loss curve looks almost flat. The rule of thumb, backed by the 2025 *LoRA Without Regret* results and now repeated in mainstream training documentation, is about **10x the full fine-tuning rate** (closer to 15x for very short runs), which lands at **1e-4 to 2e-4** in practice. Under the standard alpha-over-rank scaling used by most implementations, that optimum is roughly independent of adapter size, which is why a single number transfers across configurations. The mistake runs both ways: carrying a 2e-4 adapter rate over to a full-parameter run usually degrades the model into repetitive or incoherent output.
go deeper
Know that fine-tuning has a learning rate and that adapter-based runs use a much larger one than full-weight runs. Quoting the ballpark — 1e-5 for full fine-tuning, 1e-4 for LoRA — is enough at this level.
Be ready to explain the mechanism: only a small adapter is trainable, so each update must move it further, hence roughly ten times the full fine-tuning rate. State real numbers and note that the schedule's peak is what those numbers refer to.
Expect to diagnose from evidence — a flat loss curve pointing at an adapter trained at a full-weight rate, degenerate samples pointing at the reverse — and to describe a cheap three-point sweep that finds the rate without spending the training budget.
Own the defaults: publish a per-training-mode rate so teams stop re-deriving folklore, and decide how much compute a sweep deserves against simply shipping a known-good recipe. Be clear about when a new base model invalidates the default.
## What the learning rate is doing here The learning rate is the size of the step the optimizer takes along the gradient at each update. It is the single hyperparameter that most often decides whether a fine-tune works, and it is not a property of the data or the model alone — it is a property of *what you are training*. Two runs on identical data, with identical batch sizes and identical schedules, need learning rates an order of magnitude apart if one updates all the weights and the other updates only an adapter. ## Full fine-tuning: every weight is in play In full-parameter supervised fine-tuning, every parameter of the base model receives an update. The model already sits at a good point in weight space after pretraining, and a fine-tuning dataset is tiny compared with the pretraining corpus. Large steps therefore overwrite pretrained behaviour rather than nudging it: the model loses fluency, starts repeating tokens, or collapses into a single response style. The conventional range for post-training a mid-size model is **1e-5 to 2e-5**, sometimes as low as 5e-6 for larger models or noisier data. ## Adapters: a small number of trainable parameters LoRA freezes the pretrained weights entirely and inserts small trainable matrices whose product is added to the frozen weight. The trainable parameter count drops by two or three orders of magnitude. Those new parameters start from an initialization where the adapter contributes nothing, so the run has to move them a meaningful distance before the model's behaviour changes at all. At 1e-5, that journey takes far more steps than a short fine-tuning run has — the loss curve descends slightly and then flattens, and the tuned model is nearly indistinguishable from the base. Practitioners routinely misread this as "the dataset is too small" when the actual cause is a rate copied from a full fine-tuning recipe. ## Why roughly ten times The 10x figure is empirical, not derived. It was folklore for years and was pinned down by systematic sweeps published in 2025, which found the optimal LoRA rate sitting about an order of magnitude above the optimal full fine-tuning rate on the same data, with the gap widening to roughly 15x on very short runs where there is less time to accumulate progress. The same work found that under the usual scaling convention — the adapter output is scaled by alpha divided by rank — the optimal rate is close to **rank-independent**, so you do not need to re-tune it when you change adapter size. That property is what makes a single rule of thumb useful at all. ## What the two failure modes look like **Too low (adapter at a full fine-tuning rate):** training loss falls a little and plateaus early; generated output is essentially the base model; held-out loss barely separates from the base model's. The fix is to raise the rate, not to add data or epochs. **Too high (full weights at an adapter rate):** the model degrades in ways that are obvious in samples long before they are obvious in the loss number — stuttering repetition, dropped instruction following, degenerate punctuation. Loss may even look acceptable because the model has found a shortcut for the training distribution while general capability has been damaged. A style-transfer run trained at 2e-4 on full weights can turn a fluent model into a stuttering one in a few hundred steps. **Far too high, either mode:** loss spikes and goes to NaN within the first steps, often masked or delayed by warmup. ## The peak rate is only half the story The number people quote is the *peak* rate. What the optimizer actually experiences is the schedule: a warmup ramp into the peak, then a decay away from it. A rate that diverges under a constant schedule can be perfectly stable with a few percent of warmup and a cosine tail. So a rate is only comparable across runs when the schedule and the effective batch size are held fixed too. ## Searching for it cheaply A full sweep is rarely affordable. The practical method is three short runs on a slice of the data, a factor of three apart, compared at a *fixed step count* rather than a fixed number of epochs, and judged on held-out loss plus a quick look at samples. Choose the largest rate that still trains smoothly; rates near the divergence edge often lead early and lose late. Then scale up to the full run, keeping the schedule shape unchanged.
- Does the ten-times rule change if the run is only a few hundred steps?Yes — short runs favour the upper end, closer to fifteen times the full fine-tuning rate, because there are too few updates to make progress at a conservative rate. The tradeoff is stability: a short, hot run leans harder on warmup and on decaying the rate to near zero by the end, and it is more sensitive to an unlucky early batch. If you cannot afford the instability, add steps rather than heat.
- How would you find a good rate without burning the whole training budget?Run three short probes on a slice of the data at rates a factor of three apart — say 5e-5, 1.5e-4 and 4e-4 for an adapter — and compare held-out loss at a fixed step count with the schedule shape held constant. Pick the largest rate that still trains smoothly, then launch the full job. Judge samples as well as loss: a rate near the divergence edge often looks best early and worst at the end.
- Does the optimal rate need re-tuning when you change adapter size?Under the standard scaling convention, where the adapter's contribution is divided by its rank, the optimal learning rate is roughly rank-independent — that is the whole point of the convention. So a rate found at one adapter size usually transfers. What does move the optimum is a change of base model, a large change in effective batch size, or switching between adapter and full-parameter training.
saying these in an interview costs you the question
- Says the same learning rate works for LoRA and full fine-tuning
- Claims adapters need a lower rate because the base is frozen
- Treats 2e-4 as a safe rate for full-parameter fine-tuning
- Reads a flat adapter loss curve as a data problem
- Assumes the optimal rate scales linearly with adapter size