skip to content

In 16-bit floats, why can adding a 1e-7 update to a weight of 1.0 change nothing?

level: middleimportance: should knowfreq 46%

answer

  1. precision is relative to magnitude
  2. spacing near 1.0, not near zero
  3. half an ulp is the threshold
  4. the addend is fine, the sum is not
  5. carry the lost bits, or round randomly

basics

~20 s

A 16-bit float keeps about 11 significand bits, so the next value above 1.0 is roughly 1.001. A 1e-7 update is far below half that gap, so the sum rounds back to 1.0 and the weight never moves.

solid answer

~40 s

This is a **precision** failure, not a range failure: `1e-7` and `1.0` are each perfectly representable, and it is the addition that fails. Mantissa width sets *relative* spacing, so near 1.0 the gap to the next representable value is `2^-10`, about `9.8e-4`, in a 16-bit float against `2^-23`, about `1.2e-7`, in a 32-bit one. Under round-to-nearest, adding anything below half a gap returns the original value unchanged, so `w <- w - lr * g` becomes a silent no-op whenever the update is more than about 5e-4 times smaller than the weight it lands on. The fixes all target the accumulation: keep the parameter copy wider, use **stochastic rounding** so the update survives in expectation, or use **compensated (Kahan) accumulation**, carrying the discarded low bits into the next step.

code

python · 20 lines
python
import math

def half(x):                       # round to 11 significand bits
    m, e = math.frexp(x)
    return math.ldexp(round(m * 2048) / 2048, e)

steps, upd = 20000, 1e-7
w = 1.0
for _ in range(steps):
    w = half(w + upd)              # each sum rounds straight back to 1.0
print("plain      ", w)

w, carry = 1.0, 0.0                # compensated (Kahan) accumulation
for _ in range(steps):
    inc = upd + carry              # carry holds what the last add threw away
    total = half(w + inc)
    carry = inc - (total - w)
    w = total
print("compensated", w)
print("exact      ", 1.0 + steps * upd)

go deeper

for a junior

Recall that floating-point numbers get sparser as they get larger, and that adding a very small number to a much larger one can leave it unchanged. Know that this is about precision, not about the small number being unrepresentable.

for a middle

Explain the mechanics: exponents are aligned, the sum is rounded to nearest, and an addend below half an ulp of the larger operand disappears. Be able to state the ulp at 1.0 for a 16-bit and a 32-bit float and work the 1.0 plus 1e-7 example in both.

for a senior

Demonstrate that you would suspect this when a run flatlines after a learning-rate decay, and that you would check whether gradients are nonzero while weights are bit-identical. Be ready to compare wider accumulation, stochastic rounding and compensated accumulation on cost and on what each actually guarantees.

for a principal

Own the tradeoff between throughput and accumulation precision as a policy rather than a per-run tweak: where in a training system exactness must be preserved, what it costs in memory, and whether unbiased noise from stochastic rounding is an acceptable substitute for exactness in your domain.

## Relative, not absolute Floating-point precision is *relative* to magnitude. The mantissa (significand) field stores a fixed number of significant bits; the exponent slides those bits up and down the number line. So the spacing between neighbouring representable values — one **ulp**, unit in the last place — scales with the value itself. For a 16-bit float with 10 stored mantissa bits (11 significant bits including the implicit leading 1): - near 1.0, one ulp is `2^-10` ≈ `9.8e-4` - near 1000, one ulp is about `1` - near 0.01, one ulp is about `1e-5` For a 32-bit float with 23 stored mantissa bits, near 1.0 one ulp is `2^-23` ≈ `1.2e-7`. The quantity often quoted as machine epsilon in the half-ulp (unit roundoff) convention is `2^-11` ≈ `4.9e-4` for 16-bit and `2^-24` ≈ `6e-8` for 32-bit — that is the worst-case relative error of a single rounded operation. ## The addition rule that bites To add two floats, the hardware aligns exponents, adds, then rounds the result to the nearest representable value. If the smaller addend is below half an ulp of the larger, the exact sum lies closer to the larger operand than to any other representable value, so the rounded result *is* the larger operand. The addend vanishes completely. Work the example. In 16-bit, `1.0 + 1e-7`: half an ulp at 1.0 is about `4.9e-4`, and `1e-7` is four thousand times smaller, so the result is `1.0` exactly. In 32-bit the same sum survives, but barely: half an ulp at 1.0 is about `6e-8`, `1e-7` is larger than that, so the sum rounds up to the next representable value, `1 + 1.19e-7`. One format loses the update entirely; the other is within a factor of two of losing it. ## Why this is worse than it sounds Three properties make stagnation a real failure mode rather than a curiosity. **It is magnitude-dependent, so it is per-parameter.** The same absolute update that moves a weight of 0.01 easily is swallowed by a weight of 1000, because the ulp there is a hundred thousand times larger. In one tensor, small weights keep learning while large ones freeze. **It gets worse exactly when you want fine adjustment.** Learning-rate decay shrinks `lr * g` deliberately in the late phase of training. The moment the update drops below half an ulp of the weight, decay stops meaning "smaller steps" and starts meaning "no steps". Lowering the learning rate to fix a wobbling late-stage run can therefore silently stop the run instead. **Repetition does not save you.** Intuition says a thousand tiny nudges must eventually add up. They do not: each add is rounded independently, and each one returns the original value. The weight sits at 1.0 forever no matter how many steps you take. This is what separates the failure from ordinary rounding noise — the error is not random and averaging, it is a systematic, always-in-the-same-direction loss. ## The three fixes **Wider accumulation.** Do the arithmetic that has to be exact — the parameter update and the running optimizer statistics — in a format with enough mantissa bits that the update exceeds half an ulp, and keep the narrow copy only for the throughput-critical work. This costs memory and is the reason accumulation precision is discussed separately from storage precision. **Stochastic rounding.** Instead of always rounding to the nearest representable value, round up with probability equal to how far the exact result sits between the two neighbours. An update that is one five-thousandth of an ulp then moves the weight up one ulp roughly one step in five thousand, and the expected value of the accumulated weight is correct. It converts a systematic bias into unbiased noise — usually a good trade, because optimisation tolerates noise and does not tolerate a frozen parameter. **Compensated (Kahan) accumulation.** Keep a small extra variable holding the low-order bits the last addition threw away, and add it back into the next update. The discarded remainders accumulate in the compensation term until together they exceed half an ulp, at which point the weight finally steps. It costs one extra value per parameter and some arithmetic, and it recovers most of the lost signal deterministically. ## Keep the two failures apart A gradient that is too small to represent at all and becomes zero is an exponent-range failure. A perfectly representable update that disappears into a larger weight is a mantissa-precision failure. They look identical from outside — a weight that does not move — and they have different fixes. The diagnostic question that separates them is simply: is the *gradient* zero, or is the gradient fine and the *sum* unchanged?

  • How does stochastic rounding recover the update?
    It rounds up with probability equal to the exact result's fractional position between the two neighbouring representable values. An update one five-thousandth of an ulp large moves the weight one ulp about one step in five thousand, so the accumulated weight is correct in expectation. Systematic loss becomes unbiased noise, which optimisation tolerates far better.
  • Would lowering the learning rate help here?
    No — it is the fastest way to make it worse. The failure triggers when the update falls below half an ulp of the weight, so shrinking the update pushes more parameters below that threshold. A late-stage run that stops improving after a decay step is a classic presentation of this, and the fix is accumulation precision, not the schedule.
  • Why does the same absolute update survive on some weights and not others?
    Because spacing scales with magnitude. In a 16-bit float the gap between representable values near 0.01 is about 1e-5, while near 1000 it is about 1. A fixed update of 1e-4 moves the small weight and is invisible to the large one, so within a single tensor part of the parameters keep learning and part freeze.

A kitchen scale that reads to the nearest gram will show the same number whether or not you add a grain of salt, no matter how many grains you add one at a time.

saying these in an interview costs you the question

  • Says the update is too small to be represented
  • Suggests lowering the learning rate as the fix
  • Assumes many repeated tiny adds eventually accumulate
  • Confuses this with a gradient underflowing to zero
  • Treats floating-point precision as absolute rather than relative

context