skip to content

In LiteRT, what calls does a minimal Interpreter inference pass need?

level: juniorimportance: must knowfreq 72%

answer

  1. Six calls, and one of them is setup
  2. Nothing works before the buffers exist
  3. Indices come from a dict, not counting
  4. Reshape invalidates the memory plan
  5. Names instead of indices for multi-input models

basics

~10 s

Construct the Interpreter with the .tflite file, call allocate_tensors(), read the input's index from get_input_details(), set_tensor(index, data), invoke(), then get_tensor() on the output index. Without allocate_tensors() there are no buffers to write into.

solid answer

~40 s

The LiteRT runtime API is deliberately tiny. You build an `Interpreter` over the `.tflite` flatbuffer, then call `allocate_tensors()` — that walks the graph, works out every intermediate tensor's size, and allocates the arena the kernels will write into. Then `get_input_details()` and `get_output_details()` each return a list of dicts describing the model's inputs and outputs: `index`, `shape`, `dtype`, `quantization`. You copy your data in with `set_tensor(input_details[0]["index"], array)` — note that the index is the *tensor* index from the details dict, not the position of the input — run `invoke()`, and read results with `get_tensor(output_details[0]["index"])`. The array you pass must already match the declared shape and dtype; the interpreter does no casting or resizing for you. If you change the input shape with `resize_tensor_input()`, you must call `allocate_tensors()` again before the next `invoke()`.

code

python · 13 lines
python
import numpy as np
from ai_edge_litert.interpreter import Interpreter

interpreter = Interpreter(model_path="model.tflite")
interpreter.allocate_tensors()

inp = interpreter.get_input_details()[0]
out = interpreter.get_output_details()[0]

x = np.zeros(inp["shape"], dtype=inp["dtype"])
interpreter.set_tensor(inp["index"], x)
interpreter.invoke()
print(interpreter.get_tensor(out["index"]))

go deeper

for a junior

Memorize the order: construct, allocate_tensors, read the details, set_tensor, invoke, get_tensor. Say plainly that the index comes from get_input_details() and that the input array must already match the declared shape and dtype.

for a middle

Explain what allocate_tensors() actually does — resolve shapes and allocate the shared tensor arena — and why a resize_tensor_input() call forces you to run it again. Mention the signature runner as the named-input alternative.

for a senior

Show you have hit the real failures: hard-coded indices breaking after a reconversion, int8 models whose quantization scale you must apply yourself, and copy overhead in a per-frame loop. Know that one interpreter is single-inference-at-a-time.

for a principal

Frame the interpreter as a thin, allocation-explicit contract and decide where it belongs in an app: interpreter lifetime and pooling, whether inference sits behind a queue or a per-thread instance, and how model reconversion is versioned so index or signature drift cannot ship silently.

## Why the API is this small LiteRT (the runtime formerly exposed as `tf.lite.Interpreter`) is designed to run inside a phone app or on a microcontroller-class device. There is no session, no graph builder, no autograd. The model arrives as a `.tflite` flatbuffer that already fixes the operator list and, usually, the tensor shapes. The runtime's whole job is: lay out memory, copy input bytes in, execute the kernel list, hand output bytes back. ## The call sequence 1. **Construct.** `Interpreter(model_path="model.tflite")` from `ai_edge_litert.interpreter`. There is also a `model_content=` form for bytes you already hold in memory. Construction parses the flatbuffer and builds the kernel list but allocates almost nothing. 2. **`allocate_tensors()`.** This is the step everyone forgets first time. It resolves shapes through the graph and allocates the tensor arena — one block of memory that the runtime reuses across intermediate tensors whose lifetimes do not overlap. Until it runs, the input tensors have no backing buffer, so `set_tensor()` has nowhere to write. 3. **`get_input_details()` / `get_output_details()`.** Each returns a list of dicts, one per model input/output. The fields you actually use are `index` (the tensor's index inside the graph), `shape` (a NumPy array of the static shape), `shape_signature` (with `-1` for dynamic dimensions), `dtype`, and `quantization` / `quantization_parameters` (the scale and zero-point for an int8 model). 4. **`set_tensor(index, value)`.** Copies your NumPy array into the runtime's buffer. The `index` argument is the tensor index taken from the details dict — passing `0` because "it is the first input" happens to work sometimes and silently writes into the wrong tensor other times. 5. **`invoke()`.** Runs the kernels synchronously on the calling thread. 6. **`get_tensor(index)`.** Returns a copy of the output buffer as a NumPy array. ## The details dict is the contract Because conversion can reorder or renumber tensors, hard-coding indices is the classic source of a model that "works until we reconvert it". Always read `index` from the details. The same dict tells you the dtype: a full-int8 model expects `uint8`/`int8` input, and passing float32 will either raise or be reinterpreted as garbage. For a quantized model, the `quantization` pair `(scale, zero_point)` is how you map real values to the integer domain: `q = round(real / scale) + zero_point`. ## Shapes: static by default A converted model normally has a fixed batch dimension — often 1. To run a different shape you call `resize_tensor_input(index, [8, 224, 224, 3])`, and then **`allocate_tensors()` again**, because the arena sizes computed earlier are now wrong. Skipping the re-allocation is the second classic bug: you either get an error about the tensor not being allocated, or you read a stale buffer. Whether resizing works at all depends on the model: if the graph was exported with a dynamic dimension (`shape_signature` shows `-1`), the shapes propagate; if a downstream op baked a constant shape at conversion time, the resize is rejected. ## Named inputs: SignatureRunner Models converted from a SavedModel carry signatures. `get_signature_runner("serving_default")` returns a callable that takes and returns **named** tensors, so you write `runner(images=x)` rather than juggling indices. It is the more robust surface for anything with more than one input, and it is the one to reach for when a model exposes several signatures (for example a stateful model with separate `init` and `step` entry points). ## Copy semantics `set_tensor()` copies; `get_tensor()` copies. If you are running in a tight loop and profiling shows the copies matter, `tensor(index)` returns a callable giving a NumPy view onto the runtime buffer, which you can fill in place — but that view is invalidated by any `allocate_tensors()` or resize, so treat it as an optimization you reach for last. ## Threading note One `Interpreter` instance is not safe to `invoke()` from multiple threads at once. `num_threads` parallelizes *inside* one inference, not across concurrent requests; concurrency means one interpreter per thread.

  • When would you use get_signature_runner() instead of set_tensor() and indices?
    When the model came from a SavedModel and carries signatures. The signature runner takes and returns named tensors, so `runner(images=x)` replaces index bookkeeping, and it is the only clean way to address a model that exposes several entry points. It also survives reconversion, which can renumber tensor indices and silently break code that hard-codes them.
  • What does the quantization field in get_input_details() tell you?
    For an integer model it gives the scale and zero-point that map real values to the integer domain: `q = round(real / scale) + zero_point`. You need it to prepare input and to interpret output, unless the converter added quantize/dequantize ops at the boundary so the model takes and returns float directly.
  • Is get_tensor() returning a view or a copy of the runtime buffer?
    A copy. `tensor(index)` instead returns a callable producing a NumPy view onto the arena, which avoids the copy but is invalidated by any `allocate_tensors()` or resize, and can be overwritten by the next `invoke()`. Use the view only when profiling shows the copies actually cost you.

saying these in an interview costs you the question

  • Calling invoke() without allocate_tensors() first
  • Passing 0 as the tensor index because it is the first input
  • Assuming the interpreter casts or reshapes your input for you
  • Resizing an input and not re-running allocate_tensors()
  • Calling invoke() on one interpreter from several threads

context