skip to content

Which tf.lite post-training quantization modes require a representative_dataset, and why?

level: middleimportance: must knowfreq 72%

answer

  1. weights only versus weights plus activations
  2. constants can be read, activations cannot
  3. Optimize.DEFAULT alone is dynamic range
  4. float16 is a cast, not a calibration
  5. a few hundred real preprocessed samples

basics

~20 s

Only full-integer quantization needs converter.representative_dataset. Dynamic-range and float16 touch weights alone, which the converter can read from the file. Integer activations need real sample inputs so the converter can observe each tensor's value range and pick a scale and zero-point.

solid answer

~50 s

There are three post-training modes on the converter. Setting `converter.optimizations = [tf.lite.Optimize.DEFAULT]` alone gives **dynamic-range** quantization: weights are stored int8, activations stay float and are quantized on the fly at run time. Adding `converter.target_spec.supported_types = [tf.float16]` gives **float16** weights — a straight cast, roughly half the size. Assigning `converter.representative_dataset` to a generator gives **full-integer** quantization: the converter runs those samples through the graph, records the min/max of every activation tensor, and bakes a fixed scale and zero-point into each one. Weights it can inspect statically; activations only exist when data flows, so calibration data is the only way to learn their ranges. A few hundred samples drawn from real training or validation data, with the same preprocessing as production, is the usual recipe. Adding `tf.lite.OpsSet.TFLITE_BUILTINS_INT8` to `target_spec.supported_ops` makes conversion fail loudly instead of silently leaving an unsupported op in float.

code

python · 12 lines
python
import tensorflow as tf

converter = tf.lite.TFLiteConverter.from_saved_model("saved_model/")
converter.optimizations = [tf.lite.Optimize.DEFAULT]

def representative_dataset():
    for sample in calibration_images[:200]:      # real, preprocessed inputs
        yield [sample[tf.newaxis, ...]]

converter.representative_dataset = representative_dataset
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
open("model_int8.tflite", "wb").write(converter.convert())

go deeper

for a junior

Know that quantization is opt-in via converter.optimizations and that the fully integer mode is the one that asks you for sample data. Being able to name the three modes is enough at this level.

for a middle

Explain why weights need no data but activations do, what the calibration pass records, and what the generator must yield. Interviewers expect the scale-and-zero-point framing here.

for a senior

Demonstrate judgment about the calibration set itself — size, coverage, and preprocessing parity with production — and about forcing the integer-only op set so unsupported ops fail at build time instead of on device.

for a principal

Own the mode choice as a hardware and process decision: which targets in the fleet demand a fully integer graph, who curates and versions the calibration set, and how accuracy is gated at each conversion step.

## The three post-training modes, in terms of the API Everything is driven off `converter.optimizations = [tf.lite.Optimize.DEFAULT]`. What you add alongside it decides the mode. **Dynamic-range** — `optimizations` alone. Weights are converted to int8 and stored that way, which cuts file size about 4x versus float32. Activations remain float32 in the graph; at inference the kernel quantizes them on the fly, does an integer matmul or convolution, and dequantizes the result. No calibration data is needed because weights are constants sitting in the file — the converter reads them and computes their ranges directly. **Float16** — `optimizations` plus `target_spec.supported_types = [tf.float16]`. Weights are cast to 16-bit float, roughly halving the file. This is a pure numeric cast with no range estimation involved, so again no calibration data. Accuracy impact is usually negligible, and float16 kernels are attractive when the target has good half-precision hardware. **Full-integer** — `optimizations` plus `representative_dataset`. Now both weights and activations become int8, and the graph can run end to end in integer arithmetic. This is the only mode that needs data. ## Why activations are the ones that need data Quantizing a tensor to int8 means choosing an affine mapping between real values and the 256 available integer levels: `real ≈ scale * (q - zero_point)`. To choose `scale` you need to know the range of values the tensor takes. For a weight tensor that is trivial: the values are baked into the model, so the converter reads the array and takes its min and max. Convolution weights are typically quantized per output channel (per-axis), which handles channels with wildly different magnitudes. An activation tensor has no values until something is fed through the network. The output range of layer 12 depends entirely on what the inputs look like. The only way to learn it is to run inference on representative samples and watch. That is exactly what the calibration pass does: it executes the float graph over your generator, records min/max statistics per tensor, and freezes them into the model as static quantization parameters. At inference, no range estimation happens at all — the scales are constants, which is precisely what makes integer-only accelerators able to run the graph. ## What the generator must yield ``` def representative_dataset(): for sample in calibration_samples: # ~100-500 samples yield [sample] # a list, one entry per model input ``` The generator yields a **list** of inputs per call, one element per model input, each shaped like a single batch. Two rules matter more than the mechanics: 1. **Use real data.** Random noise produces ranges that no real input ever produces. Typically that means implausibly wide ranges, which stretch the scale and throw away precision where the real values actually live — accuracy collapses and nothing errors. 2. **Use production preprocessing.** If your serving path normalizes to [-1, 1] but calibration feeds raw 0-255 pixels, every recorded range is wrong. This is the single most common cause of "quantization destroyed my model." Labels are irrelevant — calibration is unsupervised, it only watches tensors. A few hundred samples is the standard guidance; the returns flatten quickly, but too few samples miss the tails. ## Forcing the integer-only op set By default, if the converter meets an op with no int8 kernel it may leave that op in float, producing a mixed graph that still runs on CPU but cannot be handed wholesale to an integer accelerator. Adding ``` converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8] ``` turns that silent fallback into a conversion error naming the offending op. On any project targeting an integer-only NPU or DSP, set it — you want the failure at build time, not a mysterious performance cliff on device. ## Choosing between the modes The deciding factor is usually the target hardware, not the accuracy budget. Integer-only accelerators and microcontrollers require full-int8; there is no other option. Float16 suits targets with fast half-precision float units and is the least disruptive to accuracy. Dynamic-range is the cheap default: it gets the size win with no data pipeline and no calibration risk, but the per-inference quantize/dequantize work means it will not match full-int8 latency, and it does not unlock integer-only hardware. A sensible order of work: convert float32 and verify parity, then try dynamic-range, then full-int8 with a proper calibration set, measuring accuracy on a held-out set at each step so you know which step cost you what.

  • How many samples should the representative dataset contain, and where should they come from?
    Roughly 100 to 500 samples drawn from the training or validation set, passed through the exact preprocessing used in production. Coverage matters more than count: the samples should span the classes, lighting, lengths or whatever axes your real inputs vary along, because the recorded min/max is what every future inference is clipped to. Labels are not used.
  • What does adding tf.lite.OpsSet.TFLITE_BUILTINS_INT8 to target_spec.supported_ops change?
    It restricts conversion to the integer-only built-in kernels, so any op lacking an int8 implementation raises a conversion error instead of quietly staying float. You want that on targets whose accelerator only accepts fully integer graphs — a silent float island there means the graph is split or rejected on device, which shows up as a latency cliff rather than an error.
  • Why can dynamic-range quantization still be slower than full-int8 on the same device?
    Dynamic-range quantizes activations at every inference: each kernel measures the incoming float tensor, converts it, computes in integer, then converts back. That per-call overhead is pure work full-int8 has already done at conversion time. It also leaves float tensors in the graph, so an integer-only accelerator cannot take the model and it stays on the CPU.

Weights are like the printed prices on a menu — you can read the range straight off the page. Activations are like the day's takings: you have to actually run the restaurant for a while before you know what a typical total looks like.

saying these in an interview costs you the question

  • Thinks Optimize.DEFAULT alone produces full int8
  • Feeds random noise as the representative dataset
  • Says float16 conversion needs calibration data
  • Believes the calibration samples must be labelled
  • Calibrates with preprocessing different from production
  • Uses a handful of samples and expects stable ranges

context