Why can a gradient that is nonzero in 32-bit floats round to exactly zero in 16-bit?
answer
- range lives in one field only
- five exponent bits is a short ladder
- smallest normal near 6e-5
- subnormals end near 6e-8
- stores as zero, silently, no error
basics
~20 sA 16-bit float's 5 exponent bits bottom out near 6e-5 for normal values and near 6e-8 once subnormals run out. A gradient smaller than that has no representation, so it stores as exactly zero and that weight stops moving.
solid answer
~50 sA floating-point number is a sign bit, an exponent field and a mantissa (significand) field, and it is the **exponent width** that fixes the range of magnitudes the format can reach. A 16-bit float spends 5 bits on the exponent, so its smallest positive *normal* value is 2^-14, about 6.1e-5; below that it enters subnormals, which reach down to about 6e-8 while shedding one bit of precision per step. Anything below half of that smallest representable value rounds to exactly zero. A 32-bit float spends 8 bits on the exponent and reaches down to roughly 1.2e-38, so the same gradient is nowhere near its floor. The consequence in a training step is concrete: the gradient for that parameter becomes 0, `w - lr * 0` leaves the weight unchanged, and the parameter is effectively frozen. Nothing raises an error and no NaN appears — underflow is silent.
go deeper
Recall that a float is sign, exponent and mantissa, and that the exponent decides how small a number can get. Be able to say that a 16-bit float bottoms out around 6e-5 normal, that below that it degrades into subnormals, and that too-small values become exactly zero with no error.
Explain the mechanics: which field sets range versus precision, how subnormals trade a significand bit for reach, why round-to-nearest sends anything below half the smallest representable value to zero, and how a zero gradient turns the update rule into a no-op for that parameter.
Show how you would detect it in a real run: gradient-magnitude histograms with mass at exactly zero, parameters bit-identical across checkpoints, a loss curve that looks healthy because the surviving parameters carry it. Be ready to say why nothing in the loop warns you.
Own the framing that range and precision are two independent budgets and that a run can fail on either. Be ready to argue where numerical risk should be measured and monitored rather than assumed, and what evidence you would require before signing off a reduced-precision run.
## The layout, and which field does what Every IEEE floating-point number is three fields: a **sign** bit, an **exponent** field, and a **mantissa** (also called the significand or fraction). The value is roughly `sign * 1.mantissa * 2^(exponent - bias)`. Two of those fields do two completely different jobs, and confusing them is the most common mistake on this topic: - The **exponent width decides range** — how large and how small a magnitude the format can reach at all. - The **mantissa width decides relative precision** — how finely spaced the representable values are within a given magnitude. Underflow to zero is purely a *range* failure. It has nothing to do with how many mantissa bits you have. ## The numbers A 16-bit float uses 1 sign bit, 5 exponent bits and 10 stored mantissa bits. Five exponent bits give a very short ladder of magnitudes: - Smallest positive **normal** value: `2^-14`, about `6.1e-5`. - Largest finite value: `65504`. A 32-bit float uses 1 sign bit, 8 exponent bits and 23 stored mantissa bits: - Smallest positive normal value: about `1.2e-38`. - Largest finite value: about `3.4e38`. That is the whole story of underflow. A per-parameter gradient of, say, `1e-8` is an unremarkable number in 32-bit — it sits about thirty orders of magnitude above the floor. In 16-bit it is below the floor. ## Subnormals: a soft landing, not a rescue Between the smallest normal value and zero, IEEE formats do not jump straight to zero. They switch to **subnormal** (denormal) encoding: the implicit leading `1.` of the significand becomes `0.`, the exponent is pinned at its minimum, and the number is represented as `0.mantissa * 2^-14`. This buys extra range downward, but it is paid for one mantissa bit at a time. By the time you reach the smallest subnormal — `2^-24`, about `5.96e-8` — only a single significant bit is left. Below half of that, round-to-nearest has nowhere to go but zero. So the descent looks like this: above `6.1e-5` you have full 11-bit relative precision; between `6.1e-5` and `6e-8` you have progressively fewer significant bits; below about `3e-8` you have zero. ## What it does to a training step The update rule is `w <- w - lr * g`. If `g` has flushed to zero for a parameter, the update term is exactly `0`, the weight is written back bit-for-bit unchanged, and the parameter is frozen for that step. If the gradients for that parameter are chronically in the underflow region, it is frozen for the whole run — it contributes whatever its initialised value contributes and never learns. Two properties make this nasty in practice: 1. **It is silent.** There is no exception, no infinity, no NaN. The IEEE underflow flag exists in hardware, but nothing in a training loop reads it by default. 2. **It is partial.** Only the parameters whose gradients happen to be tiny are affected. The loss keeps going down, driven by the parameters whose gradients survived, so the curve looks healthy while a slice of the model is dead. The usual tell is a histogram of gradient magnitudes with a large spike at exactly zero, or a fraction of parameters whose values are bit-identical between checkpoints. ## Why the direction matters Gradients are the exposed tensor, not activations. Activations in a well-normalised network sit around order 1 — comfortably mid-range for any format. Gradients are routinely several orders of magnitude smaller than the activations that produced them, and a small learning signal makes them smaller still. That asymmetry is why the *backward* values are the ones that hit the floor first while the forward pass looks fine. ## The two-sided view Range has a ceiling as well as a floor. The same 5 exponent bits that stop at `6e-8` below also stop at `65504` above; a value larger than that becomes `inf`, and once an infinity enters an arithmetic chain it propagates (and `inf - inf` or `0 * inf` produces NaN). Underflow and overflow are the same fact about exponent width seen from opposite ends. And note what underflow is *not*. A gradient rounding to zero is a range failure. A perfectly representable small update disappearing when it is added to a much larger weight is a *precision* failure — a different mechanism, with a different fix. Keeping those two apart is the whole reason to know the field layout in the first place.
- What are subnormal numbers, and why do they only postpone the problem?Below the smallest normal value the implicit leading 1 of the significand becomes a 0 and the exponent pins at its minimum, so numbers keep getting smaller by giving up significand bits. In a 16-bit float that buys you from about 6.1e-5 down to about 6e-8, but precision degrades to a single significant bit on the way, and then there is nothing left but zero.
- Would this show up in the loss curve?Usually not. The parameters whose gradients survive keep driving the loss down, so the curve looks normal while a slice of the model is frozen. You find it by looking at the distribution of gradient magnitudes — a spike at exactly zero — or by diffing checkpoints and noticing weights that are bit-identical across many steps.
- Why are gradients more exposed to underflow than activations?Activations in a normalised network sit around order 1, which is mid-range for any format. Gradients are typically several orders of magnitude smaller than the activations that produced them, and a small learning signal shrinks the update further. The floor of the format is therefore reached from the backward side long before the forward side gets anywhere near it.
A ruler with millimetre marks can measure a hair's width badly; a ruler that is only 30 cm long cannot measure a room at all. Exponent bits are the ruler's length, mantissa bits are its markings.
saying these in an interview costs you the question
- Thinks underflow raises an error or produces a NaN
- Says a 16-bit float just loses decimal places, not range
- Assumes any nonzero real number stores as nonzero
- Confuses the smallest normal value with the smallest representable value
- Claims 32-bit and 16-bit floats cover the same magnitudes