skip to content

Why can a gradient that is nonzero in 32-bit floats round to exactly zero in 16-bit?

level: juniorimportance: must knowfreq 64%

answer

  1. range lives in one field only
  2. five exponent bits is a short ladder
  3. smallest normal near 6e-5
  4. subnormals end near 6e-8
  5. stores as zero, silently, no error

basics

~20 s

A 16-bit float's 5 exponent bits bottom out near 6e-5 for normal values and near 6e-8 once subnormals run out. A gradient smaller than that has no representation, so it stores as exactly zero and that weight stops moving.

solid answer

~50 s

A floating-point number is a sign bit, an exponent field and a mantissa (significand) field, and it is the **exponent width** that fixes the range of magnitudes the format can reach. A 16-bit float spends 5 bits on the exponent, so its smallest positive *normal* value is 2^-14, about 6.1e-5; below that it enters subnormals, which reach down to about 6e-8 while shedding one bit of precision per step. Anything below half of that smallest representable value rounds to exactly zero. A 32-bit float spends 8 bits on the exponent and reaches down to roughly 1.2e-38, so the same gradient is nowhere near its floor. The consequence in a training step is concrete: the gradient for that parameter becomes 0, `w - lr * 0` leaves the weight unchanged, and the parameter is effectively frozen. Nothing raises an error and no NaN appears — underflow is silent.

go deeper

for a junior

Recall that a float is sign, exponent and mantissa, and that the exponent decides how small a number can get. Be able to say that a 16-bit float bottoms out around 6e-5 normal, that below that it degrades into subnormals, and that too-small values become exactly zero with no error.

for a middle

Explain the mechanics: which field sets range versus precision, how subnormals trade a significand bit for reach, why round-to-nearest sends anything below half the smallest representable value to zero, and how a zero gradient turns the update rule into a no-op for that parameter.

for a senior

Show how you would detect it in a real run: gradient-magnitude histograms with mass at exactly zero, parameters bit-identical across checkpoints, a loss curve that looks healthy because the surviving parameters carry it. Be ready to say why nothing in the loop warns you.

for a principal

Own the framing that range and precision are two independent budgets and that a run can fail on either. Be ready to argue where numerical risk should be measured and monitored rather than assumed, and what evidence you would require before signing off a reduced-precision run.

## The layout, and which field does what Every IEEE floating-point number is three fields: a **sign** bit, an **exponent** field, and a **mantissa** (also called the significand or fraction). The value is roughly `sign * 1.mantissa * 2^(exponent - bias)`. Two of those fields do two completely different jobs, and confusing them is the most common mistake on this topic: - The **exponent width decides range** — how large and how small a magnitude the format can reach at all. - The **mantissa width decides relative precision** — how finely spaced the representable values are within a given magnitude. Underflow to zero is purely a *range* failure. It has nothing to do with how many mantissa bits you have. ## The numbers A 16-bit float uses 1 sign bit, 5 exponent bits and 10 stored mantissa bits. Five exponent bits give a very short ladder of magnitudes: - Smallest positive **normal** value: `2^-14`, about `6.1e-5`. - Largest finite value: `65504`. A 32-bit float uses 1 sign bit, 8 exponent bits and 23 stored mantissa bits: - Smallest positive normal value: about `1.2e-38`. - Largest finite value: about `3.4e38`. That is the whole story of underflow. A per-parameter gradient of, say, `1e-8` is an unremarkable number in 32-bit — it sits about thirty orders of magnitude above the floor. In 16-bit it is below the floor. ## Subnormals: a soft landing, not a rescue Between the smallest normal value and zero, IEEE formats do not jump straight to zero. They switch to **subnormal** (denormal) encoding: the implicit leading `1.` of the significand becomes `0.`, the exponent is pinned at its minimum, and the number is represented as `0.mantissa * 2^-14`. This buys extra range downward, but it is paid for one mantissa bit at a time. By the time you reach the smallest subnormal — `2^-24`, about `5.96e-8` — only a single significant bit is left. Below half of that, round-to-nearest has nowhere to go but zero. So the descent looks like this: above `6.1e-5` you have full 11-bit relative precision; between `6.1e-5` and `6e-8` you have progressively fewer significant bits; below about `3e-8` you have zero. ## What it does to a training step The update rule is `w <- w - lr * g`. If `g` has flushed to zero for a parameter, the update term is exactly `0`, the weight is written back bit-for-bit unchanged, and the parameter is frozen for that step. If the gradients for that parameter are chronically in the underflow region, it is frozen for the whole run — it contributes whatever its initialised value contributes and never learns. Two properties make this nasty in practice: 1. **It is silent.** There is no exception, no infinity, no NaN. The IEEE underflow flag exists in hardware, but nothing in a training loop reads it by default. 2. **It is partial.** Only the parameters whose gradients happen to be tiny are affected. The loss keeps going down, driven by the parameters whose gradients survived, so the curve looks healthy while a slice of the model is dead. The usual tell is a histogram of gradient magnitudes with a large spike at exactly zero, or a fraction of parameters whose values are bit-identical between checkpoints. ## Why the direction matters Gradients are the exposed tensor, not activations. Activations in a well-normalised network sit around order 1 — comfortably mid-range for any format. Gradients are routinely several orders of magnitude smaller than the activations that produced them, and a small learning signal makes them smaller still. That asymmetry is why the *backward* values are the ones that hit the floor first while the forward pass looks fine. ## The two-sided view Range has a ceiling as well as a floor. The same 5 exponent bits that stop at `6e-8` below also stop at `65504` above; a value larger than that becomes `inf`, and once an infinity enters an arithmetic chain it propagates (and `inf - inf` or `0 * inf` produces NaN). Underflow and overflow are the same fact about exponent width seen from opposite ends. And note what underflow is *not*. A gradient rounding to zero is a range failure. A perfectly representable small update disappearing when it is added to a much larger weight is a *precision* failure — a different mechanism, with a different fix. Keeping those two apart is the whole reason to know the field layout in the first place.

  • What are subnormal numbers, and why do they only postpone the problem?
    Below the smallest normal value the implicit leading 1 of the significand becomes a 0 and the exponent pins at its minimum, so numbers keep getting smaller by giving up significand bits. In a 16-bit float that buys you from about 6.1e-5 down to about 6e-8, but precision degrades to a single significant bit on the way, and then there is nothing left but zero.
  • Would this show up in the loss curve?
    Usually not. The parameters whose gradients survive keep driving the loss down, so the curve looks normal while a slice of the model is frozen. You find it by looking at the distribution of gradient magnitudes — a spike at exactly zero — or by diffing checkpoints and noticing weights that are bit-identical across many steps.
  • Why are gradients more exposed to underflow than activations?
    Activations in a normalised network sit around order 1, which is mid-range for any format. Gradients are typically several orders of magnitude smaller than the activations that produced them, and a small learning signal shrinks the update further. The floor of the format is therefore reached from the backward side long before the forward side gets anywhere near it.

A ruler with millimetre marks can measure a hair's width badly; a ruler that is only 30 cm long cannot measure a room at all. Exponent bits are the ruler's length, mantissa bits are its markings.

saying these in an interview costs you the question

  • Thinks underflow raises an error or produces a NaN
  • Says a 16-bit float just loses decimal places, not range
  • Assumes any nonzero real number stores as nonzero
  • Confuses the smallest normal value with the smallest representable value
  • Claims 32-bit and 16-bit floats cover the same magnitudes

context