skip to content

Full-int8 TFLite conversion tanked your model's accuracy. How do you diagnose it?

level: seniorimportance: should knowfreq 50%

answer

  1. bisect: evaluate the float tflite first
  2. preprocessing parity before blaming int8
  3. calibration coverage decides every scale
  4. find the layer, do not guess
  5. selective float, 16-bit activations, or QAT

basics

~20 s

Bisect first: evaluate the float32 .tflite to separate conversion bugs from quantization loss. Then check that calibration data used production preprocessing, widen it if it was thin, and use per-layer error metrics to find the sensitive layer before choosing a remedy.

solid answer

~50 s

Work in order. **Isolate**: convert the same model without quantization and evaluate the float `.tflite`. If it already differs from the TensorFlow model, the problem is conversion or preprocessing, not quantization. **Check the boundary**: if you set an int8 input type, verify the app applies the recorded scale and zero-point — a preprocessing mismatch looks exactly like quantization damage. **Check calibration**: was the representative dataset real production-preprocessed data, a few hundred samples wide, covering the classes and conditions you serve? Noise or too-narrow data produces scales that clip real activations. **Localize**: `tf.lite.experimental.QuantizationDebugger` with `QuantizationDebugOptions` runs float and quantized graphs side by side and reports per-layer error, which usually points at one layer with a huge dynamic range. **Remedy**: fix the calibration set, denylist the worst ops back to float, move to 16-bit activations with `tf.lite.OpsSet.EXPERIMENTAL_TFLITE_BUILTINS_ACTIVATIONS_INT16_WEIGHTS_INT8`, or retrain with quantization-aware training.

code

python · 9 lines
python
import tensorflow as tf

debugger = tf.lite.experimental.QuantizationDebugger(
    converter=converter,
    debug_dataset=representative_dataset,
)
debugger.run()
with open("layer_stats.csv", "w") as f:
    debugger.layer_statistics_dump(f)

go deeper

for a junior

Know that converting to 8-bit integers can cost accuracy and that the sample data you supply during conversion is what sets the ranges. Being able to say "check the calibration data first" is enough here.

for a middle

Explain the bisect step — evaluate the float .tflite before blaming quantization — and describe what makes a calibration set good: real inputs, production preprocessing, a few hundred samples with real coverage.

for a senior

Show a measured workflow rather than a list of tricks: isolate, verify preprocessing parity, get per-layer error numbers, then pick the remedy that matches the cause and re-validate on the actual device.

for a principal

Own the accuracy gate: what regression threshold blocks a release, whether the fleet's hardware permits a mixed-precision graph at all, and when the organization should pay for a quantization-aware training cycle instead of tuning conversion forever.

## Step 1: bisect before you theorize "Quantization broke my model" is an assumption. Convert the same SavedModel with no `optimizations` at all and evaluate that float32 `.tflite` on the same held-out set you used in TensorFlow. Three outcomes: - Float `.tflite` matches TensorFlow, int8 does not → genuine quantization loss. Continue below. - Float `.tflite` already differs → the problem is conversion, an op-fusion difference, or the evaluation harness. Quantization is innocent and you are looking in the wrong place. - Both differ from your training numbers → suspect the evaluation pipeline itself. This costs ten minutes and saves days. ## Step 2: rule out the boundary and preprocessing The most common "quantization" bug is not quantization. If you set `inference_input_type` to `tf.int8`, the model no longer quantizes the input for you and the app must apply the tensor's recorded scale and zero-point. Feeding raw values into that input produces confidently wrong outputs with no error. Equally common: the calibration generator normalized differently from the serving path. Calibrate on `[0, 255]` and serve `[-1, 1]` and every recorded activation range is wrong. Compare the two code paths line by line — same resize, same channel order, same normalization constants. ## Step 3: interrogate the calibration set Calibration decides every activation scale in the model, and it is fed by you. - **Size.** Under a hundred samples is usually too few to see the tails; the standard guidance is a few hundred. More helps only up to a point. - **Coverage.** The samples must span what production looks like: all classes, the lighting or noise conditions, the sequence lengths. If your calibration images are all bright studio shots and production is dim phone photos, the recorded ranges do not describe the tensors you will actually see. - **Realness.** Random noise is the worst case — implausibly wide ranges stretch the scale so that real values occupy a handful of levels. Re-running conversion with a better calibration set is the cheapest remedy and often the whole fix. ## Step 4: localize the damage Whole-model accuracy tells you there is a problem, not where. `tf.lite.experimental.QuantizationDebugger`, configured with `tf.lite.experimental.QuantizationDebugOptions`, runs the float and quantized versions over your data and reports per-layer statistics — how large each tensor's range is, and how far the quantized tensor diverges from the float one. Almost always a small number of layers dominate. What you typically find: a layer whose activations have a very wide range, so 256 levels spread thin. Concatenations of tensors with mismatched ranges are a classic, as are layers immediately after an unbounded activation, and depthwise convolutions where per-channel weight magnitudes differ by orders of magnitude. Note that convolution weights are quantized per output channel by default, which already mitigates the weight side; the activation side is per-tensor, which is where the pain usually is. ## Step 5: choose the remedy that matches the cause **Better calibration data** if step 3 found a gap. Free, and fixes a surprising share of cases. **Selective quantization.** Leave the worst offenders in float by denylisting those ops or nodes in the debug options, then re-convert. You keep most of the size and speed win and recover most of the accuracy. The catch: a mixed graph may not be acceptable to an integer-only accelerator, so check the target before choosing this. **16-bit activations.** Adding `tf.lite.OpsSet.EXPERIMENTAL_TFLITE_BUILTINS_ACTIVATIONS_INT16_WEIGHTS_INT8` to `target_spec.supported_ops` keeps int8 weights (so the size win survives) while giving activations 16 bits of headroom. Accuracy usually recovers close to float. It is slower than pure int8 and has narrower kernel and hardware coverage, so verify your target supports it. **Quantization-aware training.** The heavyweight option: fine-tune the model with simulated quantization in the loop so the weights adapt to the rounding. It costs a training run but is the reliable answer when post-training quantization simply cannot hit your accuracy bar. **Model surgery.** Sometimes the honest fix is architectural — replacing an unbounded activation with a bounded one, or splitting a concatenation whose branches have incompatible ranges. ## Step 6: verify on device, not just on the desktop Finally, measure the shipped configuration end to end: the real app preprocessing, the real runtime, the real hardware. Desktop evaluation of the `.tflite` file validates the model; only the on-device run validates the deployment. Keep a small fixed set of inputs whose expected outputs you check on every build, so a preprocessing regression is caught before release rather than by users.

  • Which layer patterns tend to be the worst int8 offenders?
    Anything with a wide or unbounded activation range: layers after unbounded activations, concatenations whose branches have very different scales, and depthwise convolutions with wildly varying per-channel magnitudes. Convolution weights are already quantized per output channel, which helps the weight side; the per-tensor activation scales are usually where the precision is lost.
  • What does moving to 16-bit activations with 8-bit weights buy, and what does it cost?
    Activations get 16 bits of range, which typically recovers most of the lost accuracy, while weights stay int8 so you keep most of the file-size win. The cost is speed and portability: 16x8 kernels are slower than pure int8 and the op and hardware coverage is narrower, so an integer accelerator may not support the resulting model.
  • When do you stop tuning post-training quantization and move to quantization-aware training?
    When calibration is already good, the sensitive layers are identified, and selective float or 16-bit activations either miss the accuracy bar or are unacceptable to the target hardware. At that point the weights themselves need to adapt to rounding, which only a training run with simulated quantization gives you. Budget a fine-tuning cycle and a re-validation.

saying these in an interview costs you the question

  • Blames int8 without evaluating the float tflite first
  • Never compares calibration and serving preprocessing
  • Adds more calibration samples without checking coverage
  • Guesses at layers instead of measuring per-layer error
  • Jumps straight to QAT before fixing calibration

context