skip to content

In LiteRT, what does the Interpreter's num_threads actually control?

level: middleimportance: should knowfreq 46%

answer

  1. Inside one inference, not across many
  2. One arena means one inference at a time
  3. The default float CPU path is already delegated
  4. Big and little cores are not equal
  5. Core count is a bad default

basics

~20 s

It sets how many threads the CPU kernels use inside a single invoke() — intra-op parallelism, not concurrent requests. Float CPU inference goes through the XNNPACK delegate by default. One Interpreter still runs one inference at a time, so concurrency means one interpreter per thread.

solid answer

~50 s

`num_threads` is passed at interpreter construction and controls parallelism **within one inference**: the CPU kernels split work such as a convolution across that many threads. It does not let you serve several requests at once — a single `Interpreter` executes one `invoke()` at a time and is not safe to call concurrently, so real concurrency means an interpreter per thread or a small pool. The CPU path itself is not the naive reference implementation: XNNPACK is applied by default for float models and is where most of the threaded speedup comes from. Scaling is sublinear and often non-monotonic on phones: big.LITTLE cores differ wildly, threads compete with your UI and camera pipeline, and sustained load triggers thermal throttling, so 8 threads frequently loses to 4. Measure on the target device rather than defaulting to core count, and remember that thread count interacts with delegates — once a GPU or NNAPI delegate has claimed most of the graph, `num_threads` only affects the leftover CPU partitions.

code

python · 7 lines
python
from ai_edge_litert.interpreter import Interpreter

# num_threads is a construction-time setting: it parallelises the CPU
# kernels inside a single invoke(), not across concurrent invocations.
interpreter = Interpreter(model_path="model.tflite", num_threads=4)
interpreter.allocate_tensors()
interpreter.invoke()

go deeper

for a junior

Know that num_threads is set when you construct the Interpreter and speeds up a single inference by splitting CPU work, and that it does not make the interpreter safe to use from several threads at once.

for a middle

Draw the line between intra-op parallelism and inter-request concurrency, and say that XNNPACK is the default float CPU backend where that parallelism actually lives. Explain why scaling is sublinear on heterogeneous mobile cores.

for a senior

Demonstrate measured tuning: a thread sweep on the real device class, sustained runs that expose throttling, and awareness that inference threads compete with the camera and UI pipeline. Note that once a hardware delegate takes the graph, thread count stops mattering.

for a principal

Own the resource contract: how many interpreters exist per process, whether inference is pooled or serialised, the memory each arena costs, and the CPU budget inference is allowed so it never degrades the interactive path on the weakest supported device.

## Two different kinds of parallelism The confusion this question exists to catch is between: - **Intra-op parallelism** — one inference, split across threads. A large convolution's output tiles are computed by several threads at once. This is what `num_threads` controls. - **Inter-request concurrency** — several inferences running at the same time. LiteRT gives you nothing here. One `Interpreter` holds one tensor arena and one execution state; `invoke()` on it is not thread-safe. Two threads invoking the same interpreter is a data race on the intermediate buffers, and the symptom is corrupted outputs rather than a clean crash. If you need concurrency, you create N interpreters — each with its own arena and therefore its own memory cost — and hand them out from a pool. On a memory-constrained device that memory cost is often the reason you serialise inference behind a single worker thread instead. ## What the CPU path actually is A modern LiteRT build does not run the textbook reference kernels for float models. **XNNPACK is applied by default** as a CPU delegate: it supplies highly optimised, SIMD-vectorised, multi-threaded implementations of the common float operators, and it is the component that turns extra threads into actual speedup. This has two consequences worth stating in an interview: 1. "CPU inference" already means "delegated inference" on the default path — XNNPACK is a delegate like any other, subject to the same partial-delegation rules, and ops it does not cover fall back to the built-in kernels. 2. Benchmarks that compare "GPU delegate versus CPU" are really comparing GPU versus XNNPACK, which is a much stronger baseline than people expect. A GPU delegate that beats reference kernels may lose to XNNPACK with four threads, particularly on small models where transfer overhead dominates. ## Why more threads is not better - **Heterogeneous cores.** Phone SoCs mix a few large cores with several small ones. Scheduling a synchronised parallel kernel across both means every step waits for the slowest core; adding little cores can make an inference *slower*. - **Contention.** Your app's UI thread, camera pipeline, and audio path need CPU too. An inference that saturates every core produces dropped frames, which the user perceives as a much worse regression than 15 ms of latency. - **Thermal throttling.** A benchmark's first few runs are fast and the hundredth is not. Sustained-load numbers are the ones that predict field behaviour. - **Diminishing returns.** Parallel efficiency falls as the per-thread tile shrinks; small models with small tensors often see nothing past two threads. The practical procedure is to sweep the thread count on the actual device class and pick the knee of the curve, not the maximum — and to run long enough to be past thermal transients. ## Interaction with hardware delegates Once you attach a GPU or NNAPI delegate, the thread setting applies only to the partitions still executing on CPU. If delegation captured most of the graph, changing `num_threads` will barely move the number, and a candidate who keeps tuning it is measuring the wrong thing. Conversely, if you see thread count still mattering a lot with a delegate enabled, that is evidence the delegate captured less of the graph than you assumed. ## Measuring it The `benchmark_model` tool takes `--num_threads` alongside `--use_xnnpack`, `--use_gpu` and friends, plus `--warmup_runs` and `--num_runs`, so a thread sweep on device is a few shell lines. Because the tool reports mean and percentile latency over many runs, it also exposes the throttling behaviour that a single timed inference in your app will hide. ## Summary answer shape Say: intra-op, not inter-request; XNNPACK is the default float CPU backend and where the parallelism lives; one interpreter is one inference at a time, so concurrency is a pool; and the right thread count is measured on the target device, usually below core count.

  • Can two threads call invoke() on the same Interpreter?
    No. One interpreter owns one tensor arena and one execution state, so concurrent invocations race on the intermediate buffers and produce corrupted results rather than a clean error. For concurrency you create one interpreter per worker thread — paying the arena memory for each — or serialise inference behind a single worker, which on a memory-constrained device is often the better trade.
  • Why is a GPU delegate sometimes slower than four CPU threads?
    Because the CPU baseline is XNNPACK, not reference kernels — vectorised, multi-threaded implementations of the common float operators. For small models the GPU's host/device transfers, synchronisation, and first-use kernel compilation can exceed the compute it saves. The GPU wins on large, compute-dense graphs it can take in one partition, not on every model.
  • How would you choose a thread count for a phone app?
    Sweep it with benchmark_model on the actual device class, over enough runs to get past thermal transients, and pick the knee rather than the maximum. Then sanity-check the app end to end: a setting that shaves 10 ms of latency but competes with the camera or UI thread and drops frames is a regression, however good the microbenchmark looks.

saying these in an interview costs you the question

  • Thinking num_threads lets one interpreter serve concurrent requests
  • Setting threads to the device core count by default
  • Assuming CPU inference means unoptimised reference kernels
  • Calling invoke() on one interpreter from two threads
  • Benchmarking threads without warm-up or sustained load

context