skip to content

How do you prove a LiteRT GPU delegate actually sped up your model on device?

level: seniorimportance: should knowfreq 50%

answer

  1. Two runs, one variable changed
  2. Exclude the first inference, then price it separately
  3. Latency alone is half the verdict
  4. The per-op table names the culprit
  5. Warm device, real phone, real percentiles

basics

~20 s

Run the benchmark_model binary on the real device: a baseline pass, then --use_gpu=true, with warmup runs so delegate init and kernel compilation are excluded. Read the delegated-node and partition counts, use --enable_op_profiling for the per-op breakdown, and check outputs still match CPU.

solid answer

~50 s

Measure it on the target hardware with `benchmark_model`, not in an emulator and not with one timed call in your app. Push the binary and the `.tflite` alongside it, run a CPU baseline with `--num_threads=N`, then the same model with `--use_gpu=true` (or `--use_nnapi=true`), always with `--warmup_runs` set high enough that first-invocation delegate initialisation and kernel compilation are excluded from the steady-state number — and separately record that cold-start cost, because a model invoked once per session pays it every time. Then read the partitioning output: how many nodes the delegate took and how many partitions resulted. `--enable_op_profiling=true` gives a per-operator breakdown that shows exactly which ops stayed on CPU, which is what turns "the GPU didn't help" into an actionable model change. Finally compare outputs against the CPU path on real inputs: GPU backends often compute in float16, so latency parity is only half the verdict.

code

bash · 14 lines
bash
adb push benchmark_model /data/local/tmp/
adb push model.tflite /data/local/tmp/
adb shell chmod +x /data/local/tmp/benchmark_model

# CPU baseline (this is XNNPACK, not reference kernels)
adb shell /data/local/tmp/benchmark_model \
  --graph=/data/local/tmp/model.tflite \
  --num_threads=4 --warmup_runs=20 --num_runs=200

# GPU delegate, with the per-operator breakdown
adb shell /data/local/tmp/benchmark_model \
  --graph=/data/local/tmp/model.tflite \
  --use_gpu=true --warmup_runs=20 --num_runs=200 \
  --enable_op_profiling=true

go deeper

for a junior

Know that benchmark_model is the tool for this and that you compare a plain run against a --use_gpu run on a real device rather than timing one call inside the app.

for a middle

Describe the controlled comparison: same model and binary, one variable changed, warmup runs to exclude kernel compilation, and enough runs for a stable distribution. Mention that the output reports how many nodes were delegated.

for a senior

Show the diagnostic loop end to end — partition count and per-op profiling pointing at the op that fragments the graph, cold-start cost measured separately, sustained runs exposing throttling, and an accuracy comparison against the CPU path before anything ships.

for a principal

Turn it into a gate: a per-device-tier benchmark report with latency, cold start, partitioning and accuracy delta, stored as a regression baseline, feeding a documented policy on which delegates are enabled where and what evidence is required to change it.

## Why the app-side stopwatch is not enough Timing `invoke()` in your own app conflates several things: delegate creation, kernel compilation on first use, the steady-state inference, thermal state, and whatever else the device was doing. `benchmark_model` exists to separate them. It is a standalone binary you push to the device, so it measures the runtime rather than your app, and it reports distribution statistics over many runs instead of one sample. ## The measurement protocol **1. Baseline first.** Run the model with no hardware delegate, at the thread count you would actually ship. Remember this baseline is XNNPACK, the optimised float CPU path — a much stronger competitor than people assume, and the number a GPU has to beat. **2. Delegate run.** Same binary, same model, `--use_gpu=true` (or `--use_nnapi=true`, or `--use_coreml` on iOS). Everything else held constant. **3. Warm up.** `--warmup_runs` must be large enough to cover delegate initialisation and shader/kernel compilation, which happen at first use and can be tens or hundreds of milliseconds. Without warm-up you are benchmarking compilation. But *record* that cost too: it is the number that decides whether delegation is worth it for a model that runs once per session versus one that runs on every camera frame. **4. Enough runs, long enough.** `--num_runs` should be high enough to expose thermal throttling. Steady-state sustained latency on a warm device is the number that predicts field behaviour; a cold-device best case is marketing. **5. Real device, real device class.** Emulators do not model the GPU driver, NNAPI vendor implementation, or thermal behaviour of the phone you ship to. Benchmark across the device tiers in your fleet, because delegate performance varies more between SoCs than between models. ## Reading the output, not just the latency The run prints how the graph was partitioned: how many nodes the delegate accepted and how many separate partitions resulted. That is the diagnostic. High latency with one partition means the accelerator is genuinely slower for this workload. High latency with many partitions means fragmentation — every seam costs a memory transfer and a synchronisation — and the fix is a model change, not a flag. `--enable_op_profiling=true` adds a per-operator table: time per node and which ones executed where. This is how you find the single offending op that split your graph. Once identified, the options are to replace it with a supported equivalent, move it to the model boundary so it does not bisect the graph, or handle it outside the model in pre/post-processing. `--max_delegated_partitions` is a useful experiment: capping partitions and seeing latency improve confirms fragmentation as the cause. ## Correctness is part of the verdict A GPU backend commonly computes in float16 and may use different kernel implementations, so outputs will not be bit-identical to the CPU float32 path. For a classifier the argmax usually survives; for detection thresholds, regression outputs, or anything feeding a downstream numerical stage, it may not. The right check is a real evaluation set run through both paths, comparing the metric you actually ship — top-1, mAP, threshold-crossing rate — not a single tensor diff. Ship the delegate only when both the latency win and the metric parity hold. ## What the full report looks like A senior answer ends with the artefact: per device tier, CPU baseline latency, delegate latency, cold-start cost, delegated-node and partition counts, and the accuracy delta. That table is what lets a team decide to enable the delegate for some SoCs and not others, and it is also the regression baseline for the next model version — because the same model recompiled with different converter settings can partition completely differently. ## Common mistakes Benchmarking on an emulator; forgetting warm-up and concluding the GPU is terrible; comparing against reference CPU kernels instead of the real XNNPACK default; reporting a single mean with no percentiles or thermal soak; and declaring victory on latency without ever checking that the model still predicts the same things.

  • Why do warmup runs matter so much when benchmarking a GPU delegate?
    Delegate initialisation and kernel or shader compilation happen on first use and can cost tens to hundreds of milliseconds. Without warm-up you fold that into the average and conclude the GPU is slow. But the cost is real for infrequent inference, so measure it deliberately — a per-frame model amortises it instantly, a once-per-session model may never recover it.
  • The delegate run shows high latency and one partition. What does that tell you?
    That fragmentation is not the problem — the accelerator took the graph whole and is simply slower for this workload, typically a small or memory-bound model where transfer and synchronisation exceed the compute saved, against an already-vectorised multi-threaded CPU baseline. The response is to stop tuning delegate flags and either accept CPU or change the model's size and shape.
  • What accuracy check would you run before shipping a delegated model?
    Run a real evaluation set through both the CPU and delegated paths and compare the shipped metric — top-1, mAP, or threshold-crossing rate — not a raw tensor diff. GPU backends often compute in float16, so small numeric drift is expected; what matters is whether decisions change. Enable the delegate only when latency wins and the metric holds within an agreed tolerance.
  • Why is an emulator a poor place to evaluate delegates?
    Because the things that decide the result are absent: the vendor GPU driver, the device's NNAPI implementation, the actual memory architecture, and thermal behaviour under sustained load. Emulated results routinely invert on real silicon, and delegate performance varies more between SoCs than between model architectures, so per-device-tier benchmarking is the only trustworthy evidence.

saying these in an interview costs you the question

  • Timing a single invoke() inside the app and calling it a benchmark
  • Skipping warmup and blaming the GPU for compilation cost
  • Comparing the delegate against unoptimised reference CPU kernels
  • Benchmarking delegates on an emulator
  • Shipping on a latency win without checking accuracy

context