skip to content

In LiteRT, what happens to ops a GPU or NNAPI delegate cannot run?

level: middleimportance: must knowfreq 66%

answer

  1. It is a negotiation, not a switch
  2. Nothing throws; the number just moves
  3. Count the seams, not the nodes
  4. Each boundary is a memory crossing
  5. One odd op splits the whole graph

basics

~20 s

The graph is partitioned. The delegate claims the largest contiguous runs of ops it supports; everything else stays on CPU kernels. Nothing errors — but each partition boundary copies tensors between CPU and accelerator memory, so a fragmented graph can run slower than pure CPU.

solid answer

~50 s

Delegation is **partial**, not all-or-nothing. When you attach a GPU, NNAPI, Core ML or XNNPACK delegate, the runtime asks it which nodes it can execute, then partitions the graph into delegated subgraphs and CPU subgraphs. Each delegated subgraph is replaced by a single custom node; the remaining ops run on the built-in CPU kernels. That is why a model with one unsupported op in the middle does not fail — it silently splits into three pieces. The cost is at the seams: crossing between CPU and accelerator memory means copies, and on GPU also synchronisation that stalls the pipeline. Ten small partitions can easily be slower than never enabling the delegate at all. Ops get rejected for unsupported operator types, unsupported dtypes or quantization schemes, dynamic shapes, or tensor ranks the backend does not handle. The runtime logs how many nodes were delegated and into how many partitions, and `benchmark_model` exposes both, plus a `--max_delegated_partitions` cap.

go deeper

for a junior

Know that a delegate takes only the ops it supports and the rest quietly runs on CPU, so turning a delegate on is not a guarantee that the model runs on the GPU or NPU.

for a middle

Explain graph partitioning: supported nodes are grouped into subgraphs replaced by one delegate node each, and every boundary costs a memory transfer plus synchronisation. Name concrete rejection reasons — unsupported op, dtype, dynamic shape.

for a senior

Show how you diagnose it: read the delegated-node and partition counts, use per-op profiling to find which ops stayed on CPU, and treat a single fragmenting layer as a model-architecture fix rather than a runtime flag. Verify numerics, not just latency.

for a principal

Own the policy: which delegates are enabled for which device classes, how a per-device fallback and kill switch work, and how much model-architecture freedom you trade for delegate coverage across the fleet you actually ship to.

## Delegation is a negotiation, not a switch A LiteRT delegate is a plugin that says "give me these nodes and I will execute them on this hardware". When an interpreter is created with a delegate, the runtime runs a partitioning pass: it asks the delegate which nodes it supports, groups the supported nodes into maximal connected subgraphs, and replaces each subgraph with a single delegate node in the execution plan. Everything the delegate refused stays on the built-in CPU kernels. The key property — and the reason this question is asked constantly — is that failure is **silent and graceful**. You do not get an exception for an unsupported operator; you get a slower model. Enabling a delegate never breaks correctness in an obvious way, so the only signal is the number, and candidates who have only read about delegates assume `use_gpu=true` means "runs on GPU". ## Why the seams cost so much Every boundary between a delegated subgraph and a CPU subgraph is a data transfer: - **GPU delegate.** The tensor must be moved between host memory and GPU-visible memory (or, on mobile, between different representations of shared memory), and the CPU has to wait for the GPU to finish before it can run the next CPU node. That synchronisation destroys the pipelining that made the GPU worth using. - **NNAPI.** Crossing into the vendor driver has its own marshalling and, on many devices, a per-invocation setup cost. One partition wrapping 90% of the compute is a win. Six partitions around a handful of tiny ops is usually a loss, because you pay a transfer for each and the accelerator never gets a large enough chunk of work to amortise it. The practical rule: **count partitions, not delegated ops.** ## Why an op gets rejected - **Operator not implemented** by that backend — the long tail of TFLite operators is only partly covered by any given delegate, and custom ops or `SELECT_TF_OPS` fallback ops are never delegated. - **Dtype or quantization scheme.** A GPU backend that runs float16/float32 may refuse per-channel int8 tensors, or accept them only by inserting dequantize steps. - **Dynamic shapes.** Backends that compile a fixed kernel plan need static shapes; anything with an unknown dimension gets left on CPU. - **Rank or size limits.** Some backends cap tensor rank or dimension sizes. - **Attribute combinations** — a supported op with an unusual stride, dilation, or padding mode can still be rejected. The consequence for model design is that a single exotic layer sitting in the middle of an otherwise standard convolutional stack is disproportionately expensive: it does not just run on CPU, it *splits the graph in two*. Moving that op to the very start or end of the network, or replacing it with a supported equivalent, often buys more than any tuning flag. ## Delegate creation can fail outright Separately from per-op rejection, constructing the delegate can fail — no GPU driver, an NNAPI implementation the device does not really support, a delegate library missing from the build. Well-written code treats that as expected and falls back to CPU rather than crashing the app, because your fleet will contain devices where it happens. ## How to see what actually happened The runtime logs the outcome of partitioning: how many nodes were delegated, how many remained, and how many partitions resulted. The `benchmark_model` tool surfaces the same information alongside latency, and `--enable_op_profiling=true` gives a per-operator breakdown so you can see exactly which ops stayed behind. `--max_delegated_partitions` lets you cap the partition count, which is a blunt but useful experiment: if forcing fewer partitions makes the model faster, fragmentation was your problem. ## What good engineering looks like here Always benchmark delegate-on against delegate-off on real target hardware, not on an emulator. Treat a delegate as a *hypothesis* about a specific device class rather than a global optimisation. And check outputs as well as latency: a GPU backend computing in float16 will produce slightly different numbers than the CPU float32 path, which is usually harmless for a classifier's argmax and occasionally not harmless at all for regression or detection thresholds.

  • Why can enabling the GPU delegate make a model slower than CPU-only?
    Two reasons. Fragmentation: if the delegate takes several small subgraphs, every boundary costs a host/device transfer and a synchronisation, and the accelerator never gets enough contiguous work to amortise them. And fixed costs: delegate initialisation and kernel compilation happen on first use, so a model invoked rarely, or benchmarked without warm-up runs, pays setup on every measurement.
  • What kinds of ops most commonly get refused by a delegate?
    Custom ops and any operator pulled in via the SELECT_TF_OPS fallback are never delegated. Beyond those: operators the backend simply has not implemented, tensors in a dtype or quantization scheme it does not handle, dynamic shapes when the backend needs a static plan, and unusual attribute combinations such as odd dilation or padding on an otherwise supported convolution.
  • What should your code do when the delegate cannot be created at all on a device?
    Fall back to the CPU path and keep serving. Missing drivers, unsupported NNAPI implementations, or a delegate library absent from the build are normal across a real fleet, so delegate construction must be treated as fallible, logged with the device model, and never allowed to crash inference. A remote kill switch for the delegate on specific SoCs is the mature version of this.
  • Besides latency, what should you verify after enabling a delegate?
    Numerics. Accelerator backends often compute in float16 or use different kernel implementations, so outputs will not be bit-identical to the CPU float32 path. Compare against a CPU reference on a real evaluation set and check the metric you actually ship — top-1, detection mAP, or a threshold-crossing rate — rather than eyeballing one tensor.

saying these in an interview costs you the question

  • Assuming a delegate runs the entire model on the accelerator
  • Thinking an unsupported op raises an error
  • Judging delegation by delegated op count, not partition count
  • Benchmarking delegates on an emulator instead of the device
  • Ignoring that GPU float16 changes the outputs

context