skip to content

Why did bf16 displace fp16 as the default 16-bit format for LLMs?

level: middleimportance: must knowfreq 62%

answer

  1. Same width, different split
  2. Range versus precision, one axis
  3. Eight exponent bits, like fp32
  4. Loss scaling disappears with the wider exponent
  5. 65504 is where fp16 stops

basics

~20 s

Both formats use 16 bits, but bf16 spends more of them on the exponent and fewer on the mantissa. That gives bf16 the same dynamic range as fp32, so large activations and tiny gradients neither overflow nor vanish — at the cost of coarser precision.

solid answer

~50 s

fp16 splits its 16 bits as 1 sign, 5 exponent, 10 mantissa; bf16 splits them 1, 8, 7. Identical size, different axis. The 8-bit exponent makes bf16's range essentially fp32's, so converting an fp32 tensor to bf16 is a truncation of precision only — nothing overflows to infinity or underflows to zero. fp16 tops out around 65504 and loses small magnitudes below roughly 6e-5, which is exactly where large activations and small gradients live, so fp16 training needs loss scaling and careful range management. bf16 removes that whole class of failure, and modern accelerators give it the same throughput. The price is about three fewer mantissa bits — fewer significant digits per value — which matters much less in practice because deep nets tolerate noise far better than they tolerate infinities. That is why open-weight checkpoints ship in bf16 and quantization pipelines treat bf16 as the reference.

code

python · 5 lines
python
import torch

x = torch.tensor([1e5, 1e-8], dtype=torch.float32)
print(x.to(torch.float16))   # overflows to inf and underflows to 0
print(x.to(torch.bfloat16))  # both magnitudes survive, with fewer digits

go deeper

for a junior

Be able to say that bf16 and fp16 are both 16 bits, that bf16 trades precision for a much wider range, and that bf16 is what modern LLM checkpoints ship in.

for a middle

Give the actual bit splits — 1/5/10 versus 1/8/7 — and explain concretely what fp16's 65504 ceiling and ~6e-5 floor do to activations and gradients, and why that made loss scaling necessary.

for a senior

Argue the tradeoff from failure modes: precision loss is bounded noise, range loss is infinity and NaN. Know when casting between the two is unsafe and how you would detect saturation on real traffic.

for a principal

Own the format policy across training, serving and quantization: what the reference precision is, what every lower width is measured against, and which hardware targets in your fleet constrain the choice.

## Two ways to spend sixteen bits A floating-point number is sign x mantissa x 2^exponent. The bit budget splits between exponent bits, which set how far the format reaches (dynamic range), and mantissa bits, which set how finely it resolves values within that reach (precision). The three formats in play: - **fp32**: 1 sign, 8 exponent, 23 mantissa — huge range, high precision, 4 bytes. - **fp16 (IEEE half)**: 1 sign, 5 exponent, 10 mantissa — 2 bytes. - **bf16 (brain float)**: 1 sign, 8 exponent, 7 mantissa — 2 bytes. bf16 is deliberately fp32 with the bottom 16 mantissa bits chopped off. That design choice is the whole story. ## What the exponent split actually costs fp16's 5-bit exponent gives a maximum finite magnitude of 65504 and a smallest normal magnitude around 6.1e-5 (subnormals reach a little lower, around 6e-8, with degraded precision). Those bounds are not comfortably outside the range neural networks produce. A single large activation in a badly conditioned layer can exceed 65504 and become infinity; from there NaN propagates through the whole forward pass. In backprop the opposite failure dominates: gradients frequently sit below 1e-5, and in fp16 they round to zero, so the update silently disappears. The historical workaround was **loss scaling**: multiply the loss by a large constant before backprop so gradients land in fp16's representable window, then divide the gradients back before the optimizer step, with dynamic adjustment when infinities appear. It works, but it is an extra moving part that can and does misfire. bf16's 8-bit exponent matches fp32's, giving a range up to roughly 3.4e38 and down into the e-38 region. Nothing a transformer produces overflows it. Loss scaling becomes unnecessary. The conversion fp32 -> bf16 -> fp32 is lossy only in low-order mantissa bits, and the top bits of the bit pattern are literally identical to fp32's, which also makes conversion trivially cheap in hardware. ## What you give up bf16 has 7 explicit mantissa bits (8 with the implicit leading one), giving roughly 2-3 significant decimal digits, versus fp16's ~3-4. Consecutive representable bf16 values near 1.0 are about 0.0078 apart; fp16's are about 0.00098. So bf16 is noticeably grainier value-by-value. The reason this rarely hurts is that error in a deep network behaves differently along the two axes. Precision loss shows up as small, roughly unbiased noise that accumulation over many terms partly averages out, and that training itself is robust to. Range loss shows up as infinity or zero — catastrophic, non-recoverable, and it destroys the tensor rather than perturbing it. Given a fixed 16-bit budget, buying range with mantissa bits is the better trade for this workload. Accumulation inside matmul units happens in fp32 regardless, which further limits the damage. ## Why this matters to quantization work specifically Three practical consequences: 1. **bf16 is the reference point.** Open-weight checkpoints ship in bf16, so "full precision" in an LLM serving conversation almost always means bf16, not fp32, and the 2-bytes-per-parameter baseline follows from that. 2. **Quality comparisons are measured against bf16.** When a 4-bit or 8-bit format is described as near-lossless, the baseline is the bf16 checkpoint. 3. **The range-versus-precision axis repeats all the way down.** The same tension explains the two fp8 variants — one leaning toward precision, one toward range — and it explains why 4-bit formats attach a per-block scale: with only two exponent bits per element, range has to be restored externally. ## Hardware and the current picture Every datacenter accelerator generation in current use supports bf16 matmul at fp16-equivalent throughput, so there is no speed reason to prefer fp16. fp16 survives mainly in older or consumer-oriented stacks, in some inference-only paths where the range risk is lower than in training, and in legacy checkpoints. Some CPU and embedded targets still lack bf16 and force fp16 or fp32. One nuance worth stating in an interview: for pure inference, fp16's narrower range is less dangerous than in training because there is no backward pass and no tiny gradients. But a model *trained* in bf16 can contain weights or produce activations that overflow fp16, so casting a bf16 checkpoint to fp16 is not automatically safe. If you must, check for saturation rather than assuming.

  • If bf16 is strictly better for range, why does fp16 still exist at all?
    Because precision is not worthless. fp16's three extra mantissa bits give finer resolution, which helps where values are known to sit in a narrow, well-conditioned band — some inference paths, image and signal workloads, and hardware that supports fp16 but not bf16, including older consumer GPUs and many embedded targets. fp16 is also the IEEE standard half, so it has broader non-ML tooling support.
  • Is it safe to cast a bf16 checkpoint down to fp16 for serving?
    Not automatically. A model trained in bf16 may hold weights, and will certainly produce intermediate activations, whose magnitudes exceed fp16's 65504 ceiling; those become infinities and then NaNs. Inference is more forgiving than training because there is no tiny-gradient underflow, but you have to check for saturation on real inputs rather than assume. If the hardware supports bf16, keep bf16.
  • Does the same range-versus-precision tension appear at 8 bits?
    Yes, and it is resolved by shipping two formats rather than one. One 8-bit float variant spends more bits on the mantissa for precision, the other more on the exponent for range, and deployments pick per tensor according to whether that tensor's distribution is wide-tailed or tightly clustered. The axis you trade along never changes; only the total bit budget does.

fp16 and bf16 are two rulers of the same length: one has finer tick marks but is too short to measure the room, the other reaches wall to wall with coarser ticks. For a room, reach beats tick spacing.

saying these in an interview costs you the question

  • Saying bf16 is more accurate than fp16 rather than wider-ranged
  • Claiming bf16 uses more than 16 bits
  • Thinking loss scaling is needed because fp16 rounds badly, not because it underflows
  • Assuming any bf16 checkpoint casts safely to fp16
  • Treating fp32 as the normal baseline for modern LLM checkpoints

context