skip to content

How do you choose LiteRT delegates for a diverse Android device fleet?

level: principalimportance: should knowfreq 30%

answer

  1. A floor that always works, then optional wins
  2. The matrix is unbounded; decide locally
  3. Ship the switch, not just the setting
  4. Faster inference can mean a worse app
  5. The most durable lever is upstream

basics

~20 s

Treat delegates as a per-device decision, not a build-time constant. Ship a CPU path that always works, enable GPU or NNAPI where measurement on that device class justifies it, verify accuracy as well as latency, and keep a remote switch to disable a delegate when a driver misbehaves.

solid answer

~60 s

There is no delegate that wins everywhere, so the architecture has to admit that. Start from a **CPU baseline that always works** — XNNPACK with a tuned thread count — because delegate creation genuinely fails on real devices with missing drivers or broken vendor implementations, and inference must degrade rather than crash. Layer optional acceleration on top, decided per device class from benchmarks you actually ran on that silicon: latency, cold-start cost, partition count, and accuracy delta against the CPU path. Encode that as a remotely updatable allow/deny list keyed on SoC or device model, so a bad driver on one chipset can be switched off within hours instead of a release cycle. Some teams add a first-launch acceptance probe — run the delegate once with a timeout and a numeric check, cache the verdict — which handles devices your lab never saw. Weigh the whole cost too: NNAPI's behaviour varies enormously by vendor and Google has been steering newer Android away from it, GPU delegation competes with your rendering, and each extra delegate adds binary size.

go deeper

for a junior

Know that delegate support differs by device, that a GPU or NNAPI delegate may not be available at all, and that your code must fall back to CPU rather than fail when one cannot be created.

for a middle

Explain that the choice must be measured per device class, that delegate creation is fallible, and that GPU float16 numerics mean accuracy needs verifying alongside latency.

for a senior

Design the runtime decision: a CPU baseline, a per-device allow-list, an acceptance probe for untested hardware, telemetry keyed by SoC and OS, and benchmarks covering cold start and sustained warm latency rather than a single cold measurement.

for a principal

Own the whole trade: binary size versus coverage, GPU contention with rendering, battery and thermals, NNAPI's deprecation trajectory, the operator constraints you impose on modelling to keep graphs delegable, and the remote-config lever that lets you retract a delegate without a release.

## The premise: heterogeneity is the problem, not the model On a server you choose hardware. On a phone fleet the hardware chooses you: hundreds of SoCs, a dozen GPU driver lineages, vendor NNAPI implementations of wildly varying quality, and devices from budget tiers where the "accelerator" is slower than four CPU cores. A delegate decision that is a build-time constant is therefore wrong for a large fraction of your users by construction. The senior-to-principal shift is from "which delegate is fastest?" to "what system decides per device, and how do I change my mind after shipping?" ## Layer one: a baseline that cannot fail XNNPACK with a measured thread count is the floor. It is the default float CPU path, it is present everywhere, and it is a strong competitor — vectorised and multi-threaded. Every acceleration path must be written as an *optimisation over* this baseline, with delegate creation treated as a fallible operation: missing GPU driver, unsupported NNAPI, delegate library absent from a stripped build. Catch the failure, log it with device identifiers, run on CPU. An app that crashes because a delegate would not initialise is a defect on somebody else's silicon. ## Layer two: evidence per device class What you need per tier is a small table: CPU baseline latency, delegate latency, cold-start cost, delegated-node and partition counts, and the accuracy delta. Four facts drive the decision: - **Steady-state latency** — does the delegate win at all against XNNPACK? - **Cold start** — delegate init and kernel compilation are paid on first use. A per-frame camera model amortises them instantly; a model invoked once per session may never recover them. - **Fragmentation** — many partitions mean transfer overhead dominates, and the fix is the model, not the flag. - **Numerics** — GPU backends commonly compute in float16, so validate the shipped metric, not a tensor diff. ## Layer three: a decision you can change remotely The deployment mechanism matters as much as the measurement. A per-SoC allow/deny list delivered by remote config lets you enable a delegate on the chipsets you tested and disable it on one that turns out to produce garbage or hang, without waiting for an app-store release. This is not hypothetical: vendor driver regressions ship in OS updates, and your model is the thing that surfaces them. For the long tail you never tested, a **first-launch acceptance probe** is the pragmatic extension: on first run, execute the model once through the delegate with a fixed input, under a timeout, compare against a cached expected output within tolerance, and persist the verdict. Devices that pass get acceleration; the rest quietly stay on CPU. It costs one inference at install time and converts an unbounded compatibility matrix into a local decision. ## Layer four: costs that are not latency - **Binary size.** Each bundled delegate adds to the APK. On a fleet where install conversion matters, that trade is real, and Play delivery mechanisms that fetch components conditionally are worth considering. - **Contention.** The GPU is also drawing your UI. A model that saturates it may deliver lower inference latency and a visibly worse app. Measure frame rate alongside inference time. - **Power and heat.** Sustained accelerator use drains battery and throttles the device, so a benchmark's steady state must be a *warm* steady state. - **NNAPI's trajectory.** NNAPI was always the most variable option — the same call routes to completely different vendor implementations — and Google has been deprecating it on newer Android in favour of vendor-supplied and Play-delivered acceleration paths. Treat it as legacy support for older devices rather than the strategic direction, and say so if asked. ## Layer five: keep the model delegate-friendly The most durable lever is upstream. A model built from widely supported ops with static shapes delegates cleanly on nearly every backend; one exotic layer in the middle bisects the graph on all of them at once. Constraining the architecture to a supported operator set — and testing partitioning as part of model CI, since converter settings change partitioning between versions — buys more fleet-wide performance than any per-device tuning. ## The telemetry loop Finally, ship instrumentation: inference latency percentiles, delegate-used, delegate-init-failed, and cold-start, all keyed by device model and OS version. That is what tells you a chipset regressed after an OS update, and it turns the allow-list from a static artefact into a maintained one.

  • How would you handle devices your lab never tested?
    A first-launch acceptance probe: run one inference through the delegate under a timeout with a fixed input, compare against a cached expected output within tolerance, and persist the verdict for that install. Pass means acceleration, fail or timeout means the CPU path. It costs one inference and converts an unbounded device matrix into a local, self-correcting decision.
  • Why not just enable NNAPI everywhere on Android and let it pick the best accelerator?
    Because the same NNAPI call routes to entirely different vendor implementations, whose quality, operator coverage and per-invocation overhead vary enormously — on many devices it loses to XNNPACK. Google has also been steering newer Android away from NNAPI toward vendor and Play-delivered acceleration. Treat it as legacy support for older devices, justified by measurement, not as a universal default.
  • What telemetry would you ship to keep the delegate policy honest after launch?
    Inference latency percentiles, which delegate was used, delegate-initialisation failures, cold-start duration, and any acceptance-probe verdicts — all keyed by device model, SoC and OS version. That is what surfaces a chipset regressing after an OS update and turns your allow-list into a maintained artefact rather than a snapshot of what your lab measured once.
  • How does model architecture affect this decision?
    More than per-device tuning does. A graph built from widely supported operators with static shapes delegates cleanly almost everywhere, while one unusual layer in the middle bisects the graph on every backend simultaneously. Constraining the architecture to a supported operator set, and asserting partition counts in model CI, buys fleet-wide performance that no runtime flag can recover.

saying these in an interview costs you the question

  • Picking one delegate at build time for the whole fleet
  • Letting a delegate creation failure crash inference
  • Assuming NNAPI always beats optimised CPU kernels
  • Judging a delegate on latency without checking accuracy
  • No way to disable a delegate without an app release

context