skip to content

In 4-bit block-scaled formats like MXFP4, what does the shared per-block scale buy?

level: middleimportance: should knowfreq 42%

answer

  1. Sixteen codes cannot span a tensor
  2. One scale per small group
  3. Outliers confined to their block
  4. Count the scale in the bit budget
  5. Thirty-two elements, eight-bit scale, 4.25

basics

~20 s

Four bits can only encode about sixteen distinct values, which is far too coarse for a whole tensor. A scale shared by a small block of weights re-centres those sixteen levels on that block's own magnitude range, so a large weight elsewhere in the tensor cannot flatten this block to zero.

solid answer

~50 s

An FP4 element in MXFP4 is E2M1 — one sign, two exponent, one mantissa bit — which represents only magnitudes 0, 0.5, 1, 1.5, 2, 3, 4 and 6. On its own that grid is useless for real weights spanning many orders of magnitude. MXFP4 therefore groups 32 consecutive elements into a block with one shared 8-bit power-of-two scale (E8M0); the stored value is element x scale. Because each block gets its own scale, a heavy-tailed tensor cannot destroy the resolution of well-behaved blocks: outliers are confined to the block they live in. The cost is that the scale is real storage, so the effective width is 4 + 8/32 = 4.25 bits per weight, not 4. Smaller blocks localise outliers better and quantize more faithfully but cost more overhead per weight — a 16-element block with an 8-bit scale is 4.5 bits. Block size is precisely that dial.

code

python · 6 lines
python
def effective_bits(element_bits, scale_bits, block_size):
    return element_bits + scale_bits / block_size

print(effective_bits(4, 8, 32))   # MXFP4: 4.25 bits per weight
print(effective_bits(4, 8, 16))   # NVFP4 block term: 4.5 bits per weight
print(70e9 * effective_bits(4, 8, 32) / 8 / 1e9, "GB for 70B")

go deeper

for a junior

Know that a 4-bit weight is stored as a small code plus a scale shared by a group of neighbours, and that the scale is what lets sixteen codes represent real magnitudes.

for a middle

Compute effective bits per weight from element width, scale width and block size, and explain why block scaling limits the damage a single outlier weight can do.

for a senior

Argue the block-size tradeoff concretely — adaptation versus overhead — and connect it to whether the hardware applies scales natively in the matmul or forces an unpack step.

for a principal

Own format selection across a fleet: which widths your accelerators execute natively, what the real bytes-per-weight is after scales, and how that changes the capacity plan versus 8-bit.

## Why 4 bits alone cannot work A 4-bit element has 16 codes. In the E2M1 floating-point layout used by microscaling formats — 1 sign bit, 2 exponent bits, 1 mantissa bit — those codes cover magnitudes 0, 0.5, 1, 1.5, 2, 3, 4 and 6, plus their negatives. A transformer's weights inside a single tensor span several orders of magnitude, with a long tail of rare large values. Mapping that entire distribution onto eight positive magnitudes with one global scale gives a catastrophic outcome: the scale must be large enough to represent the largest weight, which pushes almost every ordinary weight into the bottom one or two codes, or to zero. ## The block-scaling idea Instead of one scale per tensor, cut the tensor into small contiguous blocks and give each block its own scale. Each stored element is a 4-bit code; the reconstructed weight is `code_value x block_scale`. Three things follow: 1. **Local range adaptation.** A block whose weights all sit near 0.01 gets a small scale and uses its full 16-code grid across that narrow band. A block containing a big outlier gets a large scale and quantizes coarsely — but only that block suffers. 2. **Outlier containment.** Weight outliers are the dominant source of low-bit quantization error. Blocking bounds their blast radius to 32 (or 16) weights rather than a whole matrix. 3. **Cheap dequantization.** Reconstruction is one multiply per element with a value shared across the block, which vectorizes well and, on recent hardware, is handled inside the tensor cores rather than as a separate pass. ## The arithmetic of effective bit width The scale is stored, so it counts. Effective bits per weight = element_bits + scale_bits / block_size. - **MXFP4** (the OCP microscaling format): 32 elements, E2M1 elements, one E8M0 scale. E8M0 is 8 bits of exponent only — no sign, no mantissa — so the scale is a pure power of two, making dequantization a fast exponent add rather than a multiply. Effective width: 4 + 8/32 = **4.25 bits**. - **NVFP4**: 16-element blocks with an E4M3 (8-bit float) per-block scale, plus a second-level FP32 scale for the whole tensor. Effective width: 4 + 8/16 = **4.5 bits**, ignoring the negligible per-tensor term. So a 70B model at MXFP4 is about 70e9 x 4.25 / 8 = 37 GB rather than the clean 35 GB — the kind of 5% that decides whether a deployment fits. ## The block-size dial Smaller blocks mean better local adaptation and lower quantization error, at higher overhead per weight and slightly more bookkeeping. Larger blocks are cheaper but let a single outlier degrade more neighbours. That is the entire tradeoff, and it is why the two mainstream 4-bit microscaling formats picked different points: 32 with a coarse power-of-two scale, or 16 with a finer float scale. A subtlety worth naming: a power-of-two scale can only shift the grid by whole binades, so the largest weight in a block is often not exactly representable and the block wastes part of its range. A float scale can land the grid precisely on the block's maximum, recovering some of that loss — which is a large part of why the finer-scale variant reports better fidelity at the same nominal width. ## Why this matters beyond the memory number The reason block scaling became a shipping format rather than a research trick is hardware. Recent accelerator generations execute 4-bit block-scaled matmuls natively, applying the per-block scale inside the math pipeline. That turns 4-bit from a memory trick that you pay for in dequantization overhead into a genuine compute format. Before native support, the weights were unpacked to 16-bit before every matmul, so you saved memory and bandwidth but not arithmetic. ## What to say in an interview Lead with the reason the scale exists — 16 codes cannot span a tensor — not with the format acronyms. Then give the effective-bit arithmetic, because it shows you know the scale is not free. Then name the block-size tradeoff. That sequence answers the question at the level asked, and it is robust to the format names changing next year.

  • MXFP4 uses a power-of-two scale while NVFP4 uses an FP8 scale plus a second per-tensor scale. Why two levels?
    A power-of-two scale can only align the grid to whole binades, so a block often wastes part of its range. An FP8 per-block scale lands the grid precisely on the block's maximum, cutting error at the same nominal width — but an 8-bit float has its own limited range, so a per-tensor FP32 scale is applied first to bring every block scale into that window. The second level exists to make the first level representable.
  • Would shrinking the block to 4 elements keep improving quality?
    Quality improves, but the economics stop working. At 4 elements an 8-bit scale costs 2 extra bits per weight, so the effective width reaches 6 bits — at which point you would generally be better off using a genuinely wider element format. There is also a floor: below some block size you are mostly storing scales, and the metadata itself dominates bandwidth.
  • Does block scaling apply only to weights?
    No. The same construction is applied to activations and, increasingly, to the KV cache, since those are also heavy-tailed and also dominated by memory traffic. Activations are harder because their distributions shift with the input rather than being fixed at build time, so the scales must be computed on the fly rather than baked into the artifact.

It is like measuring a hilly landscape with a short ruler: you cannot span it in one go, so you re-zero the ruler on each small patch and record the offset alongside each measurement.

saying these in an interview costs you the question

  • Quoting MXFP4 as exactly 4 bits per weight
  • Saying the scale is metadata that costs nothing
  • Believing one scale per tensor would work at 4 bits
  • Thinking smaller blocks are strictly better with no cost
  • Confusing the shared block scale with a per-element exponent

context