How do you measure a model's on-device latency so another team can trust the number?
answer
- A number needs its conditions
- Discard the first iterations
- Keep running until it heats up
- Report a distribution, not an average
- One request is not items per second
basics
~20 sState the conditions, then measure under load. Fix device, input shape, precision and batch size; discard warm-up iterations; run long enough to expose thermal throttling; and report a median and a 95th percentile rather than a mean or a best run.
solid answer
~50 sA latency number is only meaningful with its conditions attached: which device and power state, what input shape, what numeric precision, what batch size, and whether pre- and post-processing are inside the measurement. Then measure properly. Discard the first iterations — lazy initialization, first-touch allocation and clock ramp-up make them unrepresentative. Run hundreds of iterations and, on a phone, keep running for minutes, because sustained load triggers thermal throttling and steady-state numbers can be far worse than the first ten seconds. Report the **median and the 95th percentile**; a mean hides a long tail and a minimum reports a condition no user experiences. Finally, keep single-request latency and throughput as separate deliverables. At batch one the ranking of two models can be the opposite of their ranking at batch 64, so a throughput win is not a latency win.
go deeper
Be ready to say that one timed call is not a measurement: you need repeated runs, a discarded warm-up, and the conditions written down alongside the number.
Explain why the first iterations are unrepresentative, why a median and a 95th percentile beat a mean, and why batch size changes what the number even means.
Show device experience: sustained-load runs that expose thermal throttling, end-to-end scope including pre- and post-processing, and a clear separation between single-request latency and throughput.
Own what counts as evidence. Define the reporting standard a model must meet before it can be accepted against a latency budget, so teams cannot ship on a peak number taken from a cold, idle device.
## A latency number is a claim about conditions "The model runs in 12 ms" is not a fact until you say under what. A trustworthy report names, at minimum: - **Device and state** — which hardware, which compute unit, plugged in or on battery, what else was running, whether performance governors were pinned. - **Input shape and precision** — resolution or sequence length, and what numeric format weights and activations run in. - **Batch size** — and therefore whether the number is a per-request latency or a per-item throughput figure. - **Scope** — model compute only, or the end-to-end path including decoding the input, resizing, normalization and post-processing. Products pay for the end-to-end path; benchmarks often quote only the middle. - **Cold or warm** — first request after load, or steady state. Omitting these is how two teams end up with numbers that differ by 3x and both are right. ## Warm-up The first iterations are systematically unrepresentative: one-time initialization, memory allocated on first touch, caches cold, and processors sitting at low clocks until demand ramps them up. Standard practice is to run a set of warm-up iterations and discard them, then time the rest. If you also care about the **cold-start** case — an app that loads a model and immediately runs one inference — measure that separately and report it as its own number. It is a real user experience, but it is a different number, and averaging it into steady-state timings corrupts both. ## Sustained load and thermal throttling On phones and small edge devices, a benchmark that runs for ten seconds measures the device at its best and tells you nothing about the product. Under continued load the device heats, clocks are reduced to stay inside a thermal envelope, and latency degrades — often materially, after a minute or two of sustained work. If your feature runs continuously (a camera pipeline, a live filter), the steady-state throttled number *is* the number, and the peak figure is marketing. Report both, and say how long the device ran before the steady-state figure was taken. Ambient temperature and case material matter enough that the same phone gives different answers on different days. ## Percentiles, not a mean Measure many iterations and report a distribution. The **median** describes the typical request. The **95th percentile** describes what a meaningful fraction of users experience and is what latency budgets are usually written against. A **mean** is dragged around by a few slow samples and describes nobody; a **minimum** or "best of N" describes a condition no user will see. Reporting the number of iterations and the spread lets a reader judge whether your measurement is stable at all. ## Latency and throughput are different deliverables At batch size one, per-layer fixed costs dominate and hardware sits partly idle; at batch 64, the same hardware runs closer to its efficient regime and per-item cost falls. Crucially, the **ranking can flip**: model A can win at batch one and lose at batch 64, because the properties that make a model good at small batch — few operations, shallow graphs — are not the ones that make it good when there is abundant parallel work. So a single number cannot serve both an interactive request path and a bulk offline job. Decide which one the product needs, measure that regime, and label it. If someone quotes items-per-second at large batch to answer a question about how long one user waits, that is an answer to a different question. ## A workable protocol Pin the device state and record it. Load the model, run warm-up iterations, discard them. Run several hundred timed iterations at the target batch size and shape. Continue under load for a few minutes and record the steady-state distribution separately from the early one. Report median and p95 for each phase, plus the conditions above. Re-measure whenever the device, the runtime, the precision or the input shape changes — none of those numbers transfer. ## What a strong answer sounds like Lead with "a latency number needs its conditions", then give the mechanics — warm-up, many iterations, sustained load for throttling, median and p95 — and close by separating single-request latency from throughput. Candidates who have actually shipped to devices reach for throttling and percentiles unprompted; candidates who have not tend to describe timing a single call.
- Why report the 95th percentile rather than the mean?Latency distributions are right-skewed, so a mean is pulled by a handful of slow samples and describes no actual request. The median describes the typical request and the 95th percentile describes the bad-but-not-rare one, which is where user-visible budgets and timeouts are set. Reporting both bounds the experience; reporting a mean alone hides exactly the tail that causes complaints.
- When is a cold-start measurement the number that matters?Whenever the product's first inference happens right after the model loads and the user is waiting for it — an app opening a scanner, a feature invoked rarely. Then initialization, allocation and low starting clocks are part of the user's wait, and steady-state timings understate it badly. Measure and report it as its own figure rather than folding it into the warm distribution.
- Two models are compared and their ranking flips between batch one and batch 64. What do you report?Whichever regime the product runs in, labelled as such. An interactive single-request path is a batch-one question and a bulk offline job is a throughput question, and they select different models. Reporting one number for both invites someone to choose a model on a benchmark that does not describe their workload. If both paths exist, report both regimes explicitly.
Quoting a model's fastest single run is like quoting a car's top speed on a closed track: true, unreachable in traffic, and useless for planning a commute.
saying these in an interview costs you the question
- Times a single inference call and calls it done
- Reports a mean or the fastest observed run
- Never mentions warm-up or discarded iterations
- Benchmarks for ten seconds and ignores thermal throttling
- Quotes throughput to answer a single-request latency question
- Omits device, batch size, precision and input shape