What does the alpha parameter control in a LoRA adapter, relative to the rank?
answer
- a fixed scalar, never trained
- it multiplies the adapter branch
- the denominator is the rank
- keeps step size stable across a sweep
- overlaps with the learning rate
basics
~20 sAlpha is a scaling constant: a LoRA adapter's output is multiplied by alpha divided by rank before it is added to the frozen layer. It sets how strongly the learned update is applied, and dividing by rank keeps that strength stable as rank changes.
solid answer
~50 sA LoRA layer computes the frozen weight's output plus `(alpha / r) * B(Ax)`. Alpha is a fixed scalar you choose; it never trains. The reason it is divided by the rank is normalization: as you increase r you add more adapter columns, so without the division the adapter's contribution would grow with rank and every rank change would force you to re-tune the learning rate. With the `alpha / r` form, holding alpha fixed while sweeping rank keeps the effective update magnitude roughly stable, which is exactly what you want when rank is the knob under test. Two conventions are common: setting alpha to twice the rank (so the multiplier is a constant 2), or fixing alpha at a constant such as 32 and letting the multiplier fall as rank grows — the latter is what makes the optimal learning rate approximately rank-independent. Practically, alpha and the learning rate are partly redundant knobs; fix one convention and tune the other rather than sweeping both.
go deeper
Know that alpha is a fixed number in the config, not something learned, and that it and the rank appear together as a scaling factor on the adapter's output.
Be able to write the forward pass with the alpha-over-rank multiplier and explain why the rank sits in the denominator: it keeps update magnitude comparable when you change rank.
Show that alpha and learning rate are overlapping knobs, and state a convention you actually use — fixed alpha for rank sweeps, proportional alpha for a settled recipe — with the reason behind it.
Own the experimental hygiene point: a rank sweep with a confounded scaling factor tells you nothing, and adapter artifacts must carry their scaling metadata or downstream merges reproduce a model nobody evaluated.
## What alpha is Alpha is not a learned parameter and not a learning rate. It is a fixed scalar in the adapter's forward pass. A LoRA-adapted layer computes: `h = W x + (alpha / r) * B (A x)` where `W` is the frozen pretrained weight, `A` and `B` are the trained low-rank matrices, and `r` is the rank. The quantity `alpha / r` is usually called the *scaling factor*. It multiplies the adapter branch's contribution before it is summed into the frozen layer's output. Everything alpha does, it does through that one multiplication. ## Why divide by rank at all The division exists to decouple rank from update magnitude. `A` is initialized from a distribution whose per-element variance does not depend on `r`, so as you increase the rank you are summing more terms of similar size into the product `B A`. Left unnormalized, the adapter's contribution — and therefore the effective size of each optimization step at the layer's output — would drift upward as rank grows. Every rank sweep would then be confounded: you would not know whether rank 64 beat rank 8 because it had more capacity or because it was effectively training at a larger step size. With the `1 / r` factor in place, that confound largely disappears. Sweep rank at a fixed alpha and the per-step magnitude at the output stays in the same neighbourhood, so the optimal learning rate is approximately rank-independent and one learning-rate search transfers across the sweep. ## The two conventions you will meet **Alpha proportional to rank** (commonly `alpha = 2r`, e.g. rank 16 with alpha 32). Here the scaling factor is a constant — 2 in this example — regardless of rank. This is a widely used default and reads naturally: "apply the adapter at double strength." Its cost is that the normalization is cancelled out: raise rank and you also raise the effective magnitude, so a rank change may need a learning-rate change. **Alpha fixed** (for example alpha 32 held constant while rank varies from 8 to 256). The scaling factor now genuinely falls as `1 / r`. This is the convention that preserves rank-independence of the learning rate, and it is the one to prefer when rank is the variable you are actually studying. A third variant, rank-stabilized LoRA, divides by the square root of the rank instead of the rank. The motivation is that `1 / r` can over-damp very high ranks, so that increasing rank stops helping; `1 / sqrt(r)` is a gentler normalization that keeps high-rank adapters learning. It is a real, implemented option rather than a curiosity, and worth naming if you are asked how you would run a rank sweep up into the hundreds. ## Alpha and learning rate are partly redundant Because the adapter branch is linear in `B`, multiplying the branch by a constant and multiplying the effective step taken in `B` by a constant are closely related operations. Doubling alpha and doubling the learning rate are not literally identical under an adaptive optimizer, but they push in the same direction, and sweeping both at once produces a two-dimensional search where one dimension would do. The practical discipline is: fix a convention for alpha, then treat the learning rate as the knob you tune. Reporting a run as "rank 16, alpha 32" without the learning rate tells a reader very little. ## The inference-time angle Because the scaling factor is applied at inference too, it acts as a strength dial on a *trained* adapter. Serving the same adapter with a smaller scaling factor blends its behaviour more weakly into the base model's; a larger one amplifies it, usually past the point of degradation. This is occasionally useful for tuning how assertively a style adapter applies without retraining, and it is a genuine reason to keep the trained alpha recorded alongside the adapter weights. It also matters when merging: the merge must add `(alpha / r) * B A` into the base weight, and merging with the wrong scaling silently produces a model that behaves unlike the one you evaluated. ## What alpha does not do Alpha does not add capacity — that is rank's job, and no alpha value lets a rank-4 adapter represent a rank-64 update. It does not regularize in any principled sense, although a very small scaling factor will keep the model close to the base and a very large one will destabilize training. And it is not a per-layer quantity in normal use: one alpha applies across all adapted modules.
- If you keep alpha at 32 and raise the rank from 8 to 64, what happens to the adapter's effective strength?The scaling factor falls from 4 to 0.5 — an eightfold reduction. That is intentional: it offsets the extra adapter columns so the magnitude of the update the layer sees stays roughly comparable, letting the same learning rate carry across the sweep. If you instead set alpha proportional to rank, the factor stays constant and the effective update grows with rank, so a rank change may require re-tuning.
- What is rank-stabilized LoRA changing about this scaling?It divides by the square root of the rank rather than the rank itself. The concern is that the `1 / r` factor damps high-rank adapters so strongly that pushing rank into the hundreds stops delivering gains. The gentler `1 / sqrt(r)` normalization keeps large adapters learning while still removing most of the rank dependence, and it is offered as a switch in mainstream adapter tooling.
- Can you change alpha at inference without retraining?Yes — the scaling factor is applied in the forward pass, so serving with a different value rescales the adapter's influence, acting as a strength dial between base behaviour and full adapter behaviour. It is not free quality: values far from the trained setting move the model away from the configuration you evaluated. Record the trained alpha with the adapter, especially because merging must use it.
saying these in an interview costs you the question
- Calling alpha the adapter's learning rate
- Thinking alpha adds capacity the way rank does
- Believing alpha is trained along with the adapter matrices
- Sweeping alpha and learning rate independently as if orthogonal
- Merging an adapter without applying the scaling factor