skip to content

TFLite

TensorFlow Lite converts a trained model into a compact flatbuffer, applies quantization, and runs it through hardware delegates on mobile and embedded targets. Expect questions on post-training versus quantization-aware training and on what a delegate actually accelerates.

on this pageshow

questions

12

How do you convert a TensorFlow SavedModel into a .tflite file with tf.lite.TFLiteConverter?

level: juniorimportance: must knowfreq 62%

answer

  1. build a converter, then call convert
  2. from_saved_model is the preferred entry
  3. output is a buffer, not a written file
  4. float32 unless you opt in
  5. inference-only artifact, variables folded

basics

~10 s

Build a converter with tf.lite.TFLiteConverter.from_saved_model(path), then call convert(). It returns the model as a bytes FlatBuffer that you write to a .tflite file yourself. Nothing is quantized unless you set converter.optimizations.

solid answer

~40 s

The converter has three entry points: `from_saved_model(saved_model_dir)`, `from_keras_model(model)` and `from_concrete_functions(...)`. `from_saved_model` is the recommended one, because a SavedModel already carries frozen variables and concrete signatures, so the converter has everything it needs without re-tracing your Python. `convert()` returns a `bytes` object holding the FlatBuffer — it does not write anything to disk, so you open a file in `"wb"` mode and write it. The result is an inference-only artifact: variables are folded into constants, training ops are dropped, and some op patterns are fused. Conversion is float32 by default; quantization only happens if you set `converter.optimizations = [tf.lite.Optimize.DEFAULT]` and the related target-spec fields. Note the toolchain split in TensorFlow 2.21: the converter under `tf.lite` is current, while the old `tf.lite.Interpreter` runtime is deprecated in favour of LiteRT.

code

python · 7 lines
python
import tensorflow as tf

converter = tf.lite.TFLiteConverter.from_saved_model("saved_model/")
tflite_model = converter.convert()  # bytes, not a file

with open("model.tflite", "wb") as f:
    f.write(tflite_model)

go deeper

for a junior

Be able to write the four lines from memory: construct with from_saved_model, call convert(), open a file in binary write mode, write the bytes. Say plainly that the output is float32 unless you ask for quantization.

for a middle

Explain why from_saved_model is the safer entry point, what the converter folds and fuses, and that the .tflite artifact is inference-only. Know that the converter is still tf.lite while the runtime moved to LiteRT.

for a senior

Show that you treat conversion as a build step: the SavedModel is the source of truth, the .tflite is regenerated in CI, and float conversion is validated against the TensorFlow evaluation before any quantization is layered on.

for a principal

Own the artifact story across the fleet: how models are versioned, which shapes and signatures are frozen at export, how conversion failures surface before release, and how you avoid the trap of a hand-converted .tflite nobody can reproduce.

## What conversion actually is You never train in TFLite. You train with TensorFlow or Keras, then *convert* the trained graph into a `.tflite` file, which is a FlatBuffer: a flat, memory-mappable binary containing the operator list, the tensor shapes and dtypes, and the weight buffers. FlatBuffers can be read without parsing or allocating, which is why the format was chosen for phones and microcontrollers — the runtime can mmap the file and start reading tensors immediately. ## The three entry points ``` converter = tf.lite.TFLiteConverter.from_saved_model("saved_model/") converter = tf.lite.TFLiteConverter.from_keras_model(model) converter = tf.lite.TFLiteConverter.from_concrete_functions([fn], model) ``` `from_saved_model` is preferred. A SavedModel on disk already has its variables frozen into a checkpoint and its entry points recorded as signatures, so the converter reads a stable artifact rather than re-tracing live Python objects. `from_keras_model` is convenient in a notebook but depends on the in-memory model tracing cleanly. With Keras 3 (the default Keras in TensorFlow 2.x today), the robust path is to call `model.export("saved_model/")` to write a SavedModel and then use `from_saved_model`. `from_concrete_functions` is for when you want to convert one specific traced `tf.function` signature, typically to pin input shapes. ## convert() returns bytes, not a file This trips people up in interviews and in code review: ``` tflite_model = converter.convert() # bytes with open("model.tflite", "wb") as f: f.write(tflite_model) ``` There is no `output_file` argument. `convert()` runs the MLIR-based conversion pipeline in-process and hands you the buffer. If conversion fails, it raises — most often with a message naming an operator the built-in kernel set does not implement. ## What the converter changes about your graph Conversion is not a format rename. The pipeline folds variables into constants, removes anything that only existed for training (optimizer slots, gradient ops, dropout in inference mode), constant-folds subgraphs, and fuses recognizable patterns — a convolution followed by batch normalization and a ReLU typically collapses into one fused kernel with a built-in activation. That fusion is a large part of why a `.tflite` model is faster than the same graph in TensorFlow, and it is also why the converted op list will not match your layer list one-for-one when you inspect it. Because the output is inference-only, you cannot resume training from a `.tflite` file. It is a build artifact: keep the SavedModel or the checkpoint as the source of truth and treat the `.tflite` as something you can always rebuild. ## Defaults By default the output keeps float32 weights and float32 activations, so the file is roughly the same size as the float model. Size and latency work is opt-in through `converter.optimizations`, `converter.target_spec.supported_types` and `converter.representative_dataset`. Shipping a float32 model is a perfectly valid first step: convert, confirm the numbers match your TensorFlow evaluation, and only then start quantizing, so that if accuracy moves you know which step moved it. ## Shapes and signatures A converted model records concrete input shapes. If your SavedModel signature has a dynamic batch dimension, the `.tflite` model carries a placeholder batch size (usually 1) which the runtime can resize before allocating. Fully dynamic non-batch dimensions are a common source of conversion errors, because many built-in kernels need static shapes; pinning shapes at export time with an `input_signature` on the `tf.function` avoids a class of late failures. ## The current toolchain, and what moved TensorFlow Lite was rebranded LiteRT, and the two halves moved at different times. Conversion still lives in `tf.lite`: `tf.lite.TFLiteConverter` is not deprecated and `.tflite` is still the file format. The *runtime* moved — TensorFlow's own source warns that `tf.lite.Interpreter` is deprecated in favour of the interpreter shipped in the `ai_edge_litert` package, and the old standalone `tflite-runtime` wheel has stopped shipping. So a current answer is: convert with `tf.lite.TFLiteConverter`, run with LiteRT. Candidates who say "tf.lite is dead, everything moved" are half right and get corrected. ## Metadata A `.tflite` file can carry an optional metadata blob describing input normalization, label files and model provenance, which mobile codegen tools read to produce typed wrappers. It is packed into the FlatBuffer after conversion, not produced by `convert()` itself, and it is entirely optional — a model without metadata runs fine.

  • Why is from_saved_model preferred over from_keras_model?
    A SavedModel is already frozen on disk with explicit signatures, so the converter reads a stable artifact instead of re-tracing a live Python object. That removes a whole class of tracing failures and makes conversion reproducible in CI. With Keras 3 the usual recipe is `model.export(dir)` to write a SavedModel first, then `from_saved_model(dir)`.
  • Is tf.lite deprecated now that LiteRT exists?
    Only half of it. `tf.lite.TFLiteConverter` is current and still the supported way to produce a `.tflite` FlatBuffer. The runtime is what moved: `tf.lite.Interpreter` is deprecated in favour of the interpreter in the `ai_edge_litert` package, and the separate `tflite-runtime` wheel is no longer maintained. Convert with `tf.lite`, run with LiteRT.
  • What happens to training-only parts of the graph during conversion?
    They are dropped. Variables are folded into constant buffers, optimizer state and gradient ops disappear, and layers with training/inference branches convert in their inference form. Recognizable patterns such as conv plus batch-norm plus activation get fused into a single kernel, so the converted operator list will not line up one-for-one with your Keras layers.

saying these in an interview costs you the question

  • Thinks convert() writes the .tflite file itself
  • Assumes conversion quantizes the model by default
  • Expects to continue training from a .tflite file
  • Calls the .tflite output a protobuf SavedModel
  • Assumes every TensorFlow op converts unchanged

context

open as a page

In LiteRT, what calls does a minimal Interpreter inference pass need?

level: juniorimportance: must knowfreq 72%

basics

~10 s

Construct the Interpreter with the .tflite file, call allocate_tensors(), read the input's index from get_input_details(), set_tensor(index, data), invoke(), then get_tensor() on the output index. Without allocate_tensors() there are no buffers to write into.

open as a page

Which tf.lite post-training quantization modes require a representative_dataset, and why?

level: middleimportance: must knowfreq 72%

basics

~20 s

Only full-integer quantization needs converter.representative_dataset. Dynamic-range and float16 touch weights alone, which the converter can read from the file. Integer activations need real sample inputs so the converter can observe each tensor's value range and pick a scale and zero-point.

open as a page

In LiteRT, what happens to ops a GPU or NNAPI delegate cannot run?

level: middleimportance: must knowfreq 66%

basics

~20 s

The graph is partitioned. The delegate claims the largest contiguous runs of ops it supports; everything else stays on CPU kernels. Nothing errors — but each partition boundary copies tensors between CPU and accelerator memory, so a fragmented graph can run slower than pure CPU.

open as a page

In TFLite conversion, what changes when you set converter.inference_input_type to tf.int8?

level: middleimportance: should knowfreq 44%

basics

~20 s

It removes the float boundary. By default even a fully quantized model exposes float32 input and output with quantize/dequantize ops at the edges; setting inference_input_type (and inference_output_type) to tf.int8 makes the model take and return raw integers, so the caller must apply the recorded scale and zero-point itself.

open as a page

How does converting a QAT Keras model to TFLite differ from post-training int8?

level: middleimportance: should knowfreq 40%

basics

~10 s

A quantization-aware-trained model already carries fake-quant nodes holding learned ranges, so conversion needs converter.optimizations but no representative_dataset. The converter reads those recorded ranges and turns the simulated quantization into real int8 kernels.

open as a page

In LiteRT, what does the Interpreter's num_threads actually control?

level: middleimportance: should knowfreq 46%

basics

~20 s

It sets how many threads the CPU kernels use inside a single invoke() — intra-op parallelism, not concurrent requests. Float CPU inference goes through the XNNPACK delegate by default. One Interpreter still runs one inference at a time, so concurrency means one interpreter per thread.

open as a page

Why is tf.lite.Interpreter deprecated while tf.lite.TFLiteConverter is not?

level: middleimportance: should knowfreq 38%

basics

~20 s

Only the runtime moved. Google split TFLite: on-device execution now ships as LiteRT in the ai_edge_litert package, and TensorFlow's own warning points tf.lite.Interpreter users there. Conversion stayed in tf.lite, and the .tflite flatbuffer format is unchanged.

open as a page

Full-int8 TFLite conversion tanked your model's accuracy. How do you diagnose it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Bisect first: evaluate the float32 .tflite to separate conversion bugs from quantization loss. Then check that calibration data used production preprocessing, widen it if it was thin, and use per-layer error metrics to find the sensitive layer before choosing a remedy.

open as a page

What does adding tf.lite.OpsSet.SELECT_TF_OPS to a TFLite conversion cost you?

level: seniorimportance: should knowfreq 46%

basics

~20 s

It unblocks conversion by allowing selected TensorFlow kernels into the model, at the price of a much larger app binary, CPU-only execution for those ops, a fragmented graph that accelerators cannot take whole, and ops that generally will not quantize.

open as a page

How do you prove a LiteRT GPU delegate actually sped up your model on device?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Run the benchmark_model binary on the real device: a baseline pass, then --use_gpu=true, with warmup runs so delegate init and kernel compilation are excluded. Read the delegated-node and partition counts, use --enable_op_profiling for the per-op breakdown, and check outputs still match CPU.

open as a page

How do you choose LiteRT delegates for a diverse Android device fleet?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat delegates as a per-device decision, not a build-time constant. Ship a CPU path that always works, enable GPU or NNAPI where measurement on that device class justifies it, verify accuracy as well as latency, and keep a remote switch to disable a delegate when a driver misbehaves.

open as a page