Why does scaled dot-product attention divide the QK scores by sqrt(d_k)?
answer
- Scale grows with dimension
- Variance of a sum of d_k terms
- Softmax is exponential, so it saturates
- Saturated softmax has near-zero gradients
- Restore unit variance, not learned
basics
~20 sDot products of high-dimensional vectors grow in magnitude with the dimension, so raw scores get large as d_k grows. Large scores push the softmax toward a near one-hot distribution where gradients almost vanish. Dividing by sqrt(d_k) holds the scores near unit scale.
solid answer
~50 sThe score for a pair of tokens is a dot product of two d_k-dimensional vectors. If the components behave like independent, roughly unit-variance values, that sum of d_k products has variance of about d_k, so its typical magnitude grows like sqrt(d_k). With a head dimension of 128 that means scores swinging by roughly ±11 rather than ±1. Softmax is exponential, so a spread that wide saturates it: one position takes nearly all the mass and the rest get near zero. The gradient of softmax is proportional to p(1-p) terms, which collapse toward zero when p is near 0 or 1 — so the layer stops learning, and early in training it does this before the projections have learned anything useful. Dividing by sqrt(d_k) exactly cancels the dimensional growth and returns the logits to roughly unit variance, keeping the distribution soft and the gradients alive. It is a normalization constant, not a learned parameter.
code
python · 8 linesimport numpy as np
rng = np.random.default_rng(0)
for d in (8, 64, 128, 1024):
q = rng.normal(size=(20000, d))
k = rng.normal(size=(20000, d))
dots = (q * k).sum(-1)
print(d, round(float(dots.std()), 2), round(float((dots / np.sqrt(d)).std()), 2))go deeper
Know the formula includes a division by sqrt(d_k) and that it exists to stop the scores from getting too large before the softmax. Say it is a fixed constant, not something the model learns.
Derive it: a dot product of d_k terms has standard deviation about sqrt(d_k), softmax saturates on wide logit gaps, and saturated softmax has vanishing gradients. Connect the divisor back to the head dimension specifically.
Be able to name the training symptoms of getting this wrong — early loss plateau, near-zero attention entropy, tiny gradients on the query and key projections — and to distinguish the fixed correction from a temperature knob.
Frame it as one instance of a general discipline: keep activation scale invariant as width changes so hyperparameters transfer across model sizes. Be ready to relate it to initialization and normalization choices rather than treating it as an attention-only trick.
## The claim being made The attention formula is softmax(Q K^T / sqrt(d_k)) V. The divisor looks like a detail; it is load-bearing. Remove it and, at the head dimensions real models use, training becomes markedly harder or fails outright. Understanding why is a standard mid-level probe because it requires holding two facts at once: how dot products scale with dimension, and how softmax behaves at large inputs. ## Why raw scores grow with dimension A score is q · k = sum over the d_k components of q_i * k_i. Treat the components as independent with mean 0 and variance 1 — a reasonable model right after initialization, when the projections are random. Each product term then has mean 0 and variance 1, and summing d_k independent terms gives a total with variance d_k. Standard deviation is therefore sqrt(d_k). Put numbers on it. At d_k = 8, scores wander in a band of roughly ±3. At d_k = 64, roughly ±8. At d_k = 128 — the head dimension in many production-scale models — roughly ±11, so the gap between the best-matching and worst-matching key in a row can easily be 20 or more. Nothing about the *meaning* of the match changed; only the arithmetic scale did. ## Why that hurts the softmax Softmax exponentiates. A logit gap of 20 means the top position gets e^20 — about 5 x 10^8 — times the weight of the bottom one. The row is effectively one-hot. Two bad things follow. First, information loss. Attention's usefulness comes from blending several tokens with graded weights. A saturated head reads from exactly one position and ignores every other, which is a strictly weaker function and, early in training, an essentially arbitrary one, because the projections that produced the winning score are still random. Second, and more damaging, vanishing gradients. The Jacobian of softmax has entries of the form p_i(δ_ij − p_j). When one p is near 1 and the rest are near 0, every entry collapses toward zero. Gradient flowing back through the softmax into W_Q and W_K is throttled, so the very projections responsible for the bad scores cannot be corrected. The layer is stuck in a self-reinforcing bad state. This is precisely the regime the scaling factor exists to avoid. ## Why sqrt(d_k) specifically Because the standard deviation grows as sqrt(d_k), dividing by sqrt(d_k) restores unit standard deviation. It is a variance-matching correction, chosen from the same reasoning that motivates initialization schemes: keep the scale of activations stable as width changes so that one hyperparameter set works across model sizes. The constant is fixed, not learned, and costs a single scalar multiply. It is worth being precise about what it does *not* do. It does not stop attention from ever becoming peaked — a well-trained head absolutely can, and often should, put nearly all its mass on one token. The scaling only removes the *dimension-induced* peaking that would happen before any learning occurred. Sharpness after training is a learned property; the model can produce large logits by learning large-norm queries and keys if the task rewards it. ## How this shows up in practice If you ablate the divisor at a realistic head dimension, the symptoms are recognizable: loss plateaus early, attention entropy is near zero from the first steps, and gradient norms on the query and key projections are orders of magnitude smaller than on the value projection. Some implementations fold the constant into the query projection's initialization instead of applying it at runtime — mathematically identical, since scaling Q scales every logit in the row uniformly. A related, separate control is temperature. Dividing logits by a temperature above 1 also flattens the distribution, and some architectures expose a learned per-head scale. The conceptual difference matters in an interview: sqrt(d_k) is a fixed correction for a known statistical artifact of dimension, while temperature is a knob for deliberately choosing how sharp the distribution should be. ## The one-line version to say out loud "Dot products in d_k dimensions have standard deviation about sqrt(d_k). Softmax saturates on large logits and its gradients die there. Dividing by sqrt(d_k) puts the logits back at unit scale so the head starts soft and stays trainable."
- What would you actually observe during training if the divisor were removed at a head dimension of 128?Loss would flatten very early, attention entropy would sit near zero from the first steps, and gradient norms on the query and key projections would be far smaller than on the value projection. The head is stuck: it attends to essentially one arbitrary position, and because the softmax is saturated, almost no gradient flows back to fix the projections that caused it.
- Does the scaling factor prevent a trained head from ever attending sharply to a single token?No. It only removes the peaking caused by dimension alone at initialization. A trained model can learn large-norm queries and keys, producing wide logit gaps and a near one-hot distribution, and for tasks like exact copying that is the desired behaviour. The constant sets the starting scale; training decides the final sharpness.
- How is dividing by sqrt(d_k) different from applying a softmax temperature?Arithmetically they are both a division of the logits, but they answer different questions. The sqrt(d_k) term is a fixed correction for a statistical artifact of vector dimension, chosen so behaviour is stable as head width changes. Temperature is a deliberate control over how concentrated the distribution should be, and is tuned or learned rather than derived.
saying these in an interview costs you the question
- Saying the divisor keeps attention weights summing to one — softmax already does that
- Claiming sqrt(d_k) is a learned parameter tuned during training
- Explaining it as preventing numerical overflow rather than softmax saturation
- Believing it stops trained heads from ever being sharply peaked
- Saying it scales by the sequence length rather than the head dimension