skip to content

Interpreter and Delegates

On device you allocate tensors, fill the input buffers, and invoke the interpreter — a deliberately small API, now shipped as LiteRT rather than tf.lite. Speed then comes from delegates that offload supported subgraphs to GPU or NPU, with silent CPU fallback for anything they can't take.

on this pageshow

questions

6

In LiteRT, what calls does a minimal Interpreter inference pass need?

level: juniorimportance: must knowfreq 72%

answer

  1. Six calls, and one of them is setup
  2. Nothing works before the buffers exist
  3. Indices come from a dict, not counting
  4. Reshape invalidates the memory plan
  5. Names instead of indices for multi-input models

basics

~10 s

Construct the Interpreter with the .tflite file, call allocate_tensors(), read the input's index from get_input_details(), set_tensor(index, data), invoke(), then get_tensor() on the output index. Without allocate_tensors() there are no buffers to write into.

solid answer

~40 s

The LiteRT runtime API is deliberately tiny. You build an `Interpreter` over the `.tflite` flatbuffer, then call `allocate_tensors()` — that walks the graph, works out every intermediate tensor's size, and allocates the arena the kernels will write into. Then `get_input_details()` and `get_output_details()` each return a list of dicts describing the model's inputs and outputs: `index`, `shape`, `dtype`, `quantization`. You copy your data in with `set_tensor(input_details[0]["index"], array)` — note that the index is the *tensor* index from the details dict, not the position of the input — run `invoke()`, and read results with `get_tensor(output_details[0]["index"])`. The array you pass must already match the declared shape and dtype; the interpreter does no casting or resizing for you. If you change the input shape with `resize_tensor_input()`, you must call `allocate_tensors()` again before the next `invoke()`.

code

python · 13 lines
python
import numpy as np
from ai_edge_litert.interpreter import Interpreter

interpreter = Interpreter(model_path="model.tflite")
interpreter.allocate_tensors()

inp = interpreter.get_input_details()[0]
out = interpreter.get_output_details()[0]

x = np.zeros(inp["shape"], dtype=inp["dtype"])
interpreter.set_tensor(inp["index"], x)
interpreter.invoke()
print(interpreter.get_tensor(out["index"]))

go deeper

for a junior

Memorize the order: construct, allocate_tensors, read the details, set_tensor, invoke, get_tensor. Say plainly that the index comes from get_input_details() and that the input array must already match the declared shape and dtype.

for a middle

Explain what allocate_tensors() actually does — resolve shapes and allocate the shared tensor arena — and why a resize_tensor_input() call forces you to run it again. Mention the signature runner as the named-input alternative.

for a senior

Show you have hit the real failures: hard-coded indices breaking after a reconversion, int8 models whose quantization scale you must apply yourself, and copy overhead in a per-frame loop. Know that one interpreter is single-inference-at-a-time.

for a principal

Frame the interpreter as a thin, allocation-explicit contract and decide where it belongs in an app: interpreter lifetime and pooling, whether inference sits behind a queue or a per-thread instance, and how model reconversion is versioned so index or signature drift cannot ship silently.

## Why the API is this small LiteRT (the runtime formerly exposed as `tf.lite.Interpreter`) is designed to run inside a phone app or on a microcontroller-class device. There is no session, no graph builder, no autograd. The model arrives as a `.tflite` flatbuffer that already fixes the operator list and, usually, the tensor shapes. The runtime's whole job is: lay out memory, copy input bytes in, execute the kernel list, hand output bytes back. ## The call sequence 1. **Construct.** `Interpreter(model_path="model.tflite")` from `ai_edge_litert.interpreter`. There is also a `model_content=` form for bytes you already hold in memory. Construction parses the flatbuffer and builds the kernel list but allocates almost nothing. 2. **`allocate_tensors()`.** This is the step everyone forgets first time. It resolves shapes through the graph and allocates the tensor arena — one block of memory that the runtime reuses across intermediate tensors whose lifetimes do not overlap. Until it runs, the input tensors have no backing buffer, so `set_tensor()` has nowhere to write. 3. **`get_input_details()` / `get_output_details()`.** Each returns a list of dicts, one per model input/output. The fields you actually use are `index` (the tensor's index inside the graph), `shape` (a NumPy array of the static shape), `shape_signature` (with `-1` for dynamic dimensions), `dtype`, and `quantization` / `quantization_parameters` (the scale and zero-point for an int8 model). 4. **`set_tensor(index, value)`.** Copies your NumPy array into the runtime's buffer. The `index` argument is the tensor index taken from the details dict — passing `0` because "it is the first input" happens to work sometimes and silently writes into the wrong tensor other times. 5. **`invoke()`.** Runs the kernels synchronously on the calling thread. 6. **`get_tensor(index)`.** Returns a copy of the output buffer as a NumPy array. ## The details dict is the contract Because conversion can reorder or renumber tensors, hard-coding indices is the classic source of a model that "works until we reconvert it". Always read `index` from the details. The same dict tells you the dtype: a full-int8 model expects `uint8`/`int8` input, and passing float32 will either raise or be reinterpreted as garbage. For a quantized model, the `quantization` pair `(scale, zero_point)` is how you map real values to the integer domain: `q = round(real / scale) + zero_point`. ## Shapes: static by default A converted model normally has a fixed batch dimension — often 1. To run a different shape you call `resize_tensor_input(index, [8, 224, 224, 3])`, and then **`allocate_tensors()` again**, because the arena sizes computed earlier are now wrong. Skipping the re-allocation is the second classic bug: you either get an error about the tensor not being allocated, or you read a stale buffer. Whether resizing works at all depends on the model: if the graph was exported with a dynamic dimension (`shape_signature` shows `-1`), the shapes propagate; if a downstream op baked a constant shape at conversion time, the resize is rejected. ## Named inputs: SignatureRunner Models converted from a SavedModel carry signatures. `get_signature_runner("serving_default")` returns a callable that takes and returns **named** tensors, so you write `runner(images=x)` rather than juggling indices. It is the more robust surface for anything with more than one input, and it is the one to reach for when a model exposes several signatures (for example a stateful model with separate `init` and `step` entry points). ## Copy semantics `set_tensor()` copies; `get_tensor()` copies. If you are running in a tight loop and profiling shows the copies matter, `tensor(index)` returns a callable giving a NumPy view onto the runtime buffer, which you can fill in place — but that view is invalidated by any `allocate_tensors()` or resize, so treat it as an optimization you reach for last. ## Threading note One `Interpreter` instance is not safe to `invoke()` from multiple threads at once. `num_threads` parallelizes *inside* one inference, not across concurrent requests; concurrency means one interpreter per thread.

  • When would you use get_signature_runner() instead of set_tensor() and indices?
    When the model came from a SavedModel and carries signatures. The signature runner takes and returns named tensors, so `runner(images=x)` replaces index bookkeeping, and it is the only clean way to address a model that exposes several entry points. It also survives reconversion, which can renumber tensor indices and silently break code that hard-codes them.
  • What does the quantization field in get_input_details() tell you?
    For an integer model it gives the scale and zero-point that map real values to the integer domain: `q = round(real / scale) + zero_point`. You need it to prepare input and to interpret output, unless the converter added quantize/dequantize ops at the boundary so the model takes and returns float directly.
  • Is get_tensor() returning a view or a copy of the runtime buffer?
    A copy. `tensor(index)` instead returns a callable producing a NumPy view onto the arena, which avoids the copy but is invalidated by any `allocate_tensors()` or resize, and can be overwritten by the next `invoke()`. Use the view only when profiling shows the copies actually cost you.

saying these in an interview costs you the question

  • Calling invoke() without allocate_tensors() first
  • Passing 0 as the tensor index because it is the first input
  • Assuming the interpreter casts or reshapes your input for you
  • Resizing an input and not re-running allocate_tensors()
  • Calling invoke() on one interpreter from several threads

context

open as a page

In LiteRT, what happens to ops a GPU or NNAPI delegate cannot run?

level: middleimportance: must knowfreq 66%

basics

~20 s

The graph is partitioned. The delegate claims the largest contiguous runs of ops it supports; everything else stays on CPU kernels. Nothing errors — but each partition boundary copies tensors between CPU and accelerator memory, so a fragmented graph can run slower than pure CPU.

open as a page

In LiteRT, what does the Interpreter's num_threads actually control?

level: middleimportance: should knowfreq 46%

basics

~20 s

It sets how many threads the CPU kernels use inside a single invoke() — intra-op parallelism, not concurrent requests. Float CPU inference goes through the XNNPACK delegate by default. One Interpreter still runs one inference at a time, so concurrency means one interpreter per thread.

open as a page

Why is tf.lite.Interpreter deprecated while tf.lite.TFLiteConverter is not?

level: middleimportance: should knowfreq 38%

basics

~20 s

Only the runtime moved. Google split TFLite: on-device execution now ships as LiteRT in the ai_edge_litert package, and TensorFlow's own warning points tf.lite.Interpreter users there. Conversion stayed in tf.lite, and the .tflite flatbuffer format is unchanged.

open as a page

How do you prove a LiteRT GPU delegate actually sped up your model on device?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Run the benchmark_model binary on the real device: a baseline pass, then --use_gpu=true, with warmup runs so delegate init and kernel compilation are excluded. Read the delegated-node and partition counts, use --enable_op_profiling for the per-op breakdown, and check outputs still match CPU.

open as a page

How do you choose LiteRT delegates for a diverse Android device fleet?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat delegates as a per-device decision, not a build-time constant. Ship a CPU path that always works, enable GPU or NNAPI where measurement on that device class justifies it, verify accuracy as well as latency, and keep a remote switch to disable a delegate when a driver misbehaves.

open as a page