On a phone latency and app-size budget, do you quantize or ship a smaller model?
answer
- the two levers compose, not compete
- set the stopping rule before optimising
- one reduces bytes, the other reduces operations
- footprint includes runtime and kernels
- the cheapest lever that clears the budget wins
basics
~20 sUsually both, in order: fix the budgets and accuracy floor first, take int8 quantization because it is the cheapest lever, and change architecture when quantization alone misses the target. They compose — a smaller model quantizes too.
solid answer
~50 sThey are not alternatives. Int8 quantization cuts weight bytes roughly fourfold and unlocks integer kernels, and it is cheap: a post-training pass with no retraining and no architecture risk. Shrinking the architecture cuts FLOPs and activation memory, which quantization does not, but it costs a training cycle and usually more accuracy per unit of saving. So the sane order is budgets first (p95 latency on the worst device you support, app-size delta, accuracy floor), then the cheapest lever that meets them: dynamic PTQ, static PTQ, then a smaller backbone, then QAT if the accuracy floor is still missed. Two caveats shape the decision: app size includes the runtime and kernel libraries, not just the `.pte`, and quantization only pays off where the backend actually executes int8 — otherwise you have bought rounding error and no speed.
go deeper
Know that int8 quantization is the first thing to try for size and speed, and that it does not reduce how many operations the model performs.
Explain the distinct effects: quantization reduces bytes and enables integer kernels, while a smaller architecture reduces FLOPs and activation memory — and the two stack.
Show that you profile before choosing, measure the installed-app delta rather than the model file, and validate on the weakest supported device rather than a flagship.
Own the sequencing and the cost of ownership: agreed budgets as a stopping rule, cheapest lever first, an explicit decision about per-tier artifacts, and a CI gate that keeps the result from regressing.
## Reframe the question "Quantize or shrink?" is a false choice — a smaller architecture can also be quantized, and most shipped mobile models are both. The real question is **which lever to pull first, given a budget and a team**, and that is a sequencing and cost-of-ownership decision rather than a technical one. ## Fix the budget before touching the model Three numbers should exist before any optimisation work starts: 1. **Latency**, expressed as a percentile on the *worst device you support*, not the median device and not the engineer's phone. A p95 on a four-year-old mid-range handset is a different target from a p50 on a flagship. 2. **Size**, as the delta to the installed app, including the runtime and kernel libraries — not just the model file. Store listings and cellular-download thresholds make this a product constraint, not an engineering preference. 3. **Accuracy floor**, expressed on a metric the product actually cares about, with a stated tolerance. Without these, optimisation has no stopping rule and teams grind on the model long after they stopped mattering. ## What each lever actually buys **Int8 quantization** shrinks weights about 4x versus float32 and lets the CPU do integer arithmetic with better memory behaviour. Post-training, it is one pipeline pass: no retraining, no architecture change, reversible if it disappoints. What it does *not* do is reduce the number of operations — a model with too many FLOPs is still too slow, just with cheaper ones — and it does little for peak activation memory in a static graph. **A smaller architecture** — fewer channels, lower input resolution, a mobile-oriented backbone, or distillation into a compact student — reduces FLOPs *and* activation memory. That is the only lever that helps a model which is compute-bound rather than memory-bound. The costs are real: a training cycle, a fresh hyperparameter search, and a genuine accuracy drop that no amount of numerics work recovers. **QAT** is not a size lever at all; it is a way to buy back accuracy you lost to quantization, at the price of a permanent training-pipeline commitment. A useful heuristic: if profiling shows the model is dominated by large weight matrices and memory traffic, quantization is your lever. If it is dominated by convolution FLOPs at high resolution, architecture (or input resolution) is your lever. Profile before choosing. ## Binary size is not just the model A principal-level answer includes the parts juniors forget. The shipped footprint is the model artifact plus the runtime plus every operator kernel linked in. Selective build matters here: the on-device runtime can link only the kernels a given model actually uses rather than a full kernel library, which is often a larger saving than a few megabytes of weights. Multiple delegates, multiple architectures, and float fallback paths all add up. Measure the installed-app delta on a real build, not the file size of the exported artifact. ## Device fleet reality One answer rarely fits the fleet. Low-end devices, which are where latency budgets fail, benefit most from integer arithmetic. High-end devices may have accelerators that prefer float16 or have their own supported numeric formats, so an int8 model tuned for CPU is not automatically the fastest artifact everywhere. That opens a strategy question: ship one artifact and accept it is suboptimal at both ends, or ship per-tier artifacts and pay for the extra build, test and support matrix. The single-artifact answer is usually right until the fleet data says otherwise, because every extra variant multiplies validation cost. ## Cost of ownership is the deciding factor The technical comparison usually ends in a tie; the operational one does not. A PTQ pipeline is a script anyone can rerun after a model update. A QAT pipeline requires training data access, labels, compute and a revalidation pass on every retrain — and if the person who built it leaves, it silently stops being rerun. A new architecture means the research team owns a second model line forever. So the judgment is: buy the cheapest lever that clears the budget, and stop. Reach for the expensive ones only with an explicit accounting of who maintains them and how often. ## Whatever you choose, gate it Optimisation without a gate regresses. The mechanism is a device-tier benchmark and an accuracy check that run in CI on the shipped artifact, with thresholds derived from the three budget numbers. That gate is what makes the decision durable — it turns "we quantized once and it looked fine" into a property the team keeps.
- How do you decide whether a model is memory-bound or compute-bound on a phone?Profile on the target device rather than reasoning from parameter counts. Compare per-operator time against theoretical FLOPs and weight bytes moved: models dominated by large linear layers with little reuse are memory-bound and gain most from int8, while high-resolution convolution stacks are compute-bound and need fewer operations, not cheaper ones.
- Why can app-size savings from quantization disappoint relative to the model shrink?Because the shipped footprint is model plus runtime plus linked operator kernels, and quantization only touches the first. A 4x smaller weight blob against a fixed runtime and kernel library is a much smaller relative win. Selective linking of only the kernels a model uses often saves more than the weights did.
- When would you ship different artifacts for different device tiers?Only when fleet telemetry shows a single artifact fails the budget at one end and the affected population is large enough to justify the cost. Each variant multiplies build, benchmark, accuracy-validation and support work, so the bar is a measured user-impact number, not a suspicion that a flagship could go faster.
- What stops an optimised on-device model from silently regressing later?A gate on the shipped artifact: device-tier latency benchmarks and an accuracy check in CI, with thresholds derived from the agreed budgets. Without it, the next model retrain or dependency bump quietly reintroduces float fallbacks or a bigger backbone and nobody notices until users do.
saying these in an interview costs you the question
- Optimises with no latency, size or accuracy target agreed
- Treats quantization and a smaller model as mutually exclusive
- Counts only the model file as the app-size cost
- Benchmarks solely on a flagship developer phone
- Adopts QAT without accounting for who reruns it