skip to content

In TFLite conversion, what changes when you set converter.inference_input_type to tf.int8?

level: middleimportance: should knowfreq 44%

answer

  1. default keeps float edges on an int8 model
  2. quantize and dequantize ops at the boundary
  3. scale and zero-point move into your code
  4. integer-only accelerators reject float edges
  5. wrong scaling fails silently, not loudly

basics

~20 s

It removes the float boundary. By default even a fully quantized model exposes float32 input and output with quantize/dequantize ops at the edges; setting inference_input_type (and inference_output_type) to tf.int8 makes the model take and return raw integers, so the caller must apply the recorded scale and zero-point itself.

solid answer

~50 s

When you convert with a `representative_dataset`, the interior of the graph becomes int8 but the converter still wraps it: it inserts a quantize op on the input and a dequantize op on the output, so the exposed signature is float32 and callers can keep feeding normalized floats. Setting `converter.inference_input_type = tf.int8` and `converter.inference_output_type = tf.int8` (or `tf.uint8`) strips those wrapper ops. Now the model's input tensor is int8, and it is your code's job to convert: `q = round(real / scale) + zero_point`, using the scale and zero-point the converter recorded on that tensor. You do this because many integer-only accelerators and microcontroller runtimes refuse a graph with float edges, and because it saves a conversion pass per inference on large inputs. The trap is silent: feed raw values into an int8 input without applying scale and zero-point and nothing errors, the predictions are just wrong.

code

python · 9 lines
python
import tensorflow as tf

converter = tf.lite.TFLiteConverter.from_saved_model("saved_model/")
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_dataset
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.int8    # strips the input quantize op
converter.inference_output_type = tf.int8   # strips the output dequantize op
open("int8_io.tflite", "wb").write(converter.convert())

go deeper

for a junior

Know that a quantized model can still accept plain float input by default, and that changing the input type means your own code has to convert values to integers first.

for a middle

Explain the quantize/dequantize ops the converter inserts at the boundary, the affine mapping with scale and zero-point, and why the two IO-type fields only apply to a fully integer conversion.

for a senior

Show that you know this failure is silent: no error, just wrong predictions. Describe the parity check you would run between the desktop float model and the on-device integer model before shipping.

for a principal

Frame it as an interface contract between the model artifact and the application: who owns the quantize helper, how scale and zero-point travel with a re-calibrated model, and why an int8 boundary is a hardware requirement rather than a micro-optimization.

## The default: an integer model with float doors A full-integer conversion needs `converter.optimizations = [tf.lite.Optimize.DEFAULT]` and a `representative_dataset`. After that, every weight and activation inside the graph is int8. But the converter is pragmatic about the boundary: unless told otherwise it leaves `inference_input_type` and `inference_output_type` at float32 and inserts a quantize op right after the input and a dequantize op right before the output. That means the model you get back still looks float from the outside. Existing calling code that feeds normalized float arrays keeps working unchanged, and the internal speed and size benefits are already there. For most CPU deployments this default is the right one and you should not touch it. ## What setting the type to tf.int8 does ``` converter.inference_input_type = tf.int8 converter.inference_output_type = tf.int8 ``` Those two wrapper ops disappear. The graph now begins and ends in integer space. Three things follow. First, **the model's input tensor dtype changes**, so a buffer of float32 values is no longer an acceptable input — the runtime expects int8 bytes. Second, **quantization moves into your application code**. Each quantized tensor carries an affine mapping recorded at conversion time: a `scale` (a float) and a `zero_point` (an integer). To feed a real value you compute `q = round(real / scale) + zero_point` and clamp to the int8 range. To interpret an output you invert it: `real = scale * (q - zero_point)`. Those parameters are stored in the model file per tensor, and you read them from the model rather than hard-coding them, because they change every time you re-calibrate. Third, **the model becomes acceptable to integer-only hardware**. Many NPUs, DSPs and microcontroller runtimes cannot execute a float op at all. A float input op is enough to make them reject the model or split the graph, so integer edges are a hard requirement rather than an optimization there. `tf.uint8` is the other legal choice and exists mostly for older accelerator stacks and for camera pipelines that already hand you unsigned bytes; the zero-point differs accordingly (commonly around 128 for uint8 against 0-ish for int8). ## Why this is a silent-failure question The classic bug: a vision model calibrated on images normalized to [-1, 1], converted with int8 input, and then fed raw camera bytes 0..255 reinterpreted as int8. Every value above 127 wraps negative, none of them mean what the model expects, and the runtime happily returns confident nonsense. There is no dtype error, because the bytes are the right size, and no shape error, because the shape is right. The only symptom is bad accuracy on device while the desktop evaluation looked fine. The defence is to keep one quantize/dequantize helper in the app, derive it from the model's recorded parameters rather than constants in a header, and validate on-device outputs against the desktop float model for a fixed set of inputs before shipping. ## When to set it and when not to Set it when the target is integer-only hardware, when you are on a microcontroller build where the float kernels are not even compiled in, or when the input is large enough that a per-inference float-to-int conversion of the whole tensor is measurable (large images, long audio buffers). Leave the default when you deploy to a general CPU and want the model to be a drop-in replacement for the float one. The float boundary costs one pass over the input and one over the output; on a small model that is noise, and the reduced chance of a preprocessing mismatch is worth more. A useful middle position for teams: ship the float-boundary model first, confirm parity end to end, then switch the boundary to int8 as a separate, independently validated change. That way, if the numbers move, you know which change moved them. ## Interaction with the rest of the conversion These fields only make sense on a fully quantized conversion. Dynamic-range quantization leaves activations float by design, so there is no integer boundary to expose; float16 conversion likewise keeps a float interface. Attempting to force an int8 boundary without a representative dataset does not give you a quantized model — the calibration step is what produced the scales in the first place, and without it there is nothing to record on the input tensor.

  • Where do the scale and zero-point for the input tensor come from?
    From the calibration pass. Running the representative dataset records each tensor's observed min/max, and the converter turns that into an affine mapping stored on the tensor in the model file. Your app should read those values from the loaded model rather than hard-coding them, because every re-calibration or retrain can change them.
  • When would you choose tf.uint8 over tf.int8 for the boundary?
    Mostly for compatibility. Some older accelerator stacks and legacy pipelines were built around unsigned 8-bit tensors, and image sources often hand you unsigned bytes already. The arithmetic is the same affine mapping with a different zero-point convention, so the choice is dictated by what the surrounding hardware and driver expect, not by accuracy.
  • Does setting inference_input_type to tf.int8 make a dynamic-range model integer-only?
    No. Dynamic-range quantization deliberately keeps activations in float and quantizes them per inference, so there is no calibrated integer boundary to expose. These fields are meaningful only on a full-integer conversion produced with a representative dataset, which is what creates the scales in the first place.

saying these in an interview costs you the question

  • Thinks a fully quantized model always has int8 inputs
  • Feeds raw pixel bytes into an int8 input unscaled
  • Hard-codes scale and zero-point in application code
  • Believes the setting itself makes the model quantized
  • Expects a dtype error when the scaling is wrong

context