At 8 bits, what differs between int8 and fp8 for representing model tensors?
answer
- Uniform grid versus non-uniform grid
- Absolute error versus relative error
- Outliers stretch a uniform scale
- Zero-point belongs to the integer side
- Weights tolerate integers, activations prefer floats
basics
~20 sint8 lays 256 values on an evenly spaced grid set by a scale, and optionally a zero-point offset. fp8 spends its 8 bits on an exponent and mantissa, so its values crowd near zero and spread out at large magnitudes — a much better match for heavy-tailed tensors.
solid answer
~50 sBoth cost one byte, but they distribute their 256 codes differently. int8 is uniform: reconstructed value is `(code - zero_point) * scale`, so every step is the same size and the whole range is set by the largest magnitude present. That makes int8 acutely outlier-sensitive — one extreme activation stretches the scale and coarsens every ordinary value in the group. fp8 is non-uniform: E4M3 (4 exponent, 3 mantissa bits, max magnitude 448) and E5M2 (5 exponent, 2 mantissa, wider range but coarser) resolve small values finely and large values coarsely, which mirrors how weights and activations are actually distributed. In practice int8 remains strong for weights, which are well-behaved and can use symmetric per-channel scales; fp8 is preferred for activations and increasingly for the KV cache, where outliers are the norm. Hardware matters too: recent accelerator generations execute fp8 matmuls natively, so the format is a compute win rather than only a storage one.
code
python · 7 linesimport numpy as np
x = np.array([0.01, 0.02, 0.05, 12.0]) # three ordinary values, one outlier
scale = np.abs(x).max() / 127 # symmetric int8 over the whole group
q = np.round(x / scale).astype(np.int8)
print(q, np.round(q * scale, 4)) # small values collapse onto few codesgo deeper
Know that both are one byte per value, that int8 spaces its values evenly while fp8 packs them closer near zero, and that a stored scale converts codes back to real numbers.
Explain the dequantization rule with scale and zero-point, and why one large outlier in a group ruins a uniform grid but hurts a float grid far less.
Choose per tensor with reasons: integer for weights with per-channel symmetric scales, float for activations and cache, and check what the target hardware executes natively before committing.
Own the format policy across a heterogeneous fleet — datacenter accelerators with native float paths versus edge targets with only integer units — and account for the tooling and revalidation cost of maintaining two quantized artifacts.
## Same byte, different grid An 8-bit code has 256 possibilities either way. The question is where you place them on the number line. **Integer quantization** places them uniformly. The dequantization rule is `x ~= (q - z) * s`, where `q` is the stored integer, `s` is a scale and `z` is a zero-point. In the *symmetric* variant `z` is fixed at zero and the grid is centred on the origin — the usual choice for weights, which are roughly zero-centred. In the *asymmetric* variant `z` shifts the grid so it can hug a one-sided distribution such as post-ReLU activations, at the cost of storing one extra value per group and adding a cross-term to the matmul. **Float quantization** places them non-uniformly. Splitting 8 bits into exponent and mantissa yields values whose spacing doubles at every binade: dense near zero, sparse far out. The two standard variants are **E4M3** (1 sign, 4 exponent, 3 mantissa; max magnitude 448) and **E5M2** (1 sign, 5 exponent, 2 mantissa; wider range, less precision). E4M3 is the usual choice for forward-pass tensors; E5M2's extra range is aimed at values with wider spread. ## Why the shape of the grid decides the winner Transformer tensors are heavy-tailed. Most values cluster tightly around zero; a small number of channels carry magnitudes one or two orders larger, and those outliers matter to the output. With a uniform grid the scale is set by the maximum magnitude in the group. One outlier at 100x the typical value means every ordinary value is quantized on a grid 100x too coarse, and the bulk of the distribution collapses onto a handful of codes. The mitigations are all about restricting the group: per-channel instead of per-tensor scales, or small groups. That works well for weights, whose distributions are fixed and inspectable at build time, which is why weight-only int8 is a durable, high-quality option. With a float grid the outlier still forces a wide range, but the codes remaining near zero are still finely spaced, because precision is relative rather than absolute. That property is exactly what activations need — they are generated at runtime from unpredictable inputs, so their outliers cannot be tamed offline the way weights' can. ## Relative versus absolute error The cleanest framing: integer formats give roughly *constant absolute* error (half a step, everywhere); float formats give roughly *constant relative* error (a fixed fraction of the value's own magnitude). If what you care about is small values being represented proportionally well — and in a network where activations get multiplied and summed, you usually do — relative error is the better guarantee. If your values genuinely occupy a known narrow band, the integer grid uses its codes more efficiently and wins. ## Where each is used in practice - **Weights**: int8 with symmetric per-channel scales is well established and quality is very close to the 16-bit baseline. fp8 weights are also common on hardware with native support. - **Activations**: fp8 is generally preferred. Making int8 activations work requires extra machinery to move the outlier burden off activations before quantizing. - **KV cache**: 8-bit float storage is routine, and lower widths are appearing, because the cache is bandwidth-bound and its distributions are activation-like. - **Below 8 bits**: the distinction narrows. A 4-bit element has so few codes that a shared block scale is mandatory either way, and the mainstream 4-bit formats are float-shaped (E2M1) with block scales. ## Hardware is part of the answer A format only pays fully when the matrix units execute it. Recent datacenter accelerator generations added native fp8 tensor cores, and the newest add native 4-bit block-scaled paths. int8 matmul has had native support for far longer and remains ubiquitous, including on edge and CPU targets where fp8 support is absent. So the honest answer to "int8 or fp8?" starts with what your hardware runs natively; a format the hardware must emulate saves memory but converts back to 16-bit before every matmul, forfeiting the compute win. ## Answering well Do not present this as one format being better. Present it as uniform-versus-non-uniform spacing, map that onto weight-versus-activation distributions, then add the hardware constraint. That answer stays correct regardless of which vendor ships what next.
- What exactly does the zero-point add in asymmetric integer quantization, and what does it cost?It is an integer offset that shifts the representable grid so it can cover a one-sided range — useful for tensors that are never negative, since a symmetric grid would waste half its codes. The cost is an extra stored value per group and an additional cross-term when the quantized matmul is expanded, which some kernels handle inefficiently. Weights are usually symmetric precisely to avoid that term.
- When is int8 clearly the right choice over fp8?When the hardware lacks native fp8 — many edge accelerators, CPUs and older GPUs — because an emulated format saves memory but not compute. Also when the tensor is weights rather than activations: weight distributions are fixed, inspectable offline and tameable with per-channel scales, so a uniform grid loses very little. Integer paths also have far more mature tooling on non-datacenter targets.
- Does the integer-versus-float distinction still matter at 4 bits?Much less. With only sixteen codes, neither grid spans a real tensor, so a shared per-block scale becomes mandatory either way and the scale does most of the range adaptation. The mainstream 4-bit formats use a float element layout with block scales, so the practical question shifts from integer-versus-float to block size and scale precision.
An integer grid is a ruler with evenly spaced millimetre marks; a float grid is a slide rule, packing marks near the small end. Measuring dust specks and boulders with one instrument favours the slide rule.
saying these in an interview costs you the question
- Saying int8 and fp8 differ only in how the bits are named
- Claiming a uniform grid handles outliers as well as a float grid
- Forgetting that the scale (and zero-point) is stored alongside the codes
- Ignoring whether the hardware executes the format natively
- Assuming fp8 is strictly better and int8 is obsolete