When is ten-crop plus flip test-time augmentation worth its ten-fold inference cost?
answer
- five crops, each also mirrored
- an ensemble over views, not over models
- variance reduction on the prediction
- ten passes per request, not ten models
- select the threshold under the same procedure
basics
~20 sOnly when the accuracy gain is worth ten forward passes per input. Averaging predictions over ten label-preserving views cuts prediction variance for a fixed model, which suits offline or batch scoring far better than a latency-bound online request path.
solid answer
~50 sTen-crop means four corner crops plus a centre crop, each also horizontally flipped, giving ten views of one input; all ten go through the model and the predicted probabilities are averaged. It is an ensemble over views rather than over models - no retraining, no extra weights - and it typically buys a small but real accuracy gain and steadier behaviour on borderline inputs. The bill is ten times the compute, and on a serial request path ten times the latency, which is usually where it dies. So it earns its place in offline batch scoring and in evaluation, and rarely in a tight online budget, where the same compute often buys more as a bigger model or higher input resolution. Two discipline points: the views must be label-preserving and match invariances the model trained with, and any threshold you select must be selected with the averaging turned on.
go deeper
Know that augmentation normally runs at training time only, and that test-time augmentation is the named exception where several views of one input are averaged into a single prediction.
Be able to state what ten-crop is, that the ten probability vectors are averaged rather than voted, and that the weights are untouched - it is an ensemble over views, not over models.
Argue the cost: ten passes per request, tail latency, and the fact that selection and thresholds must be done under the same inference procedure that ships. Know the cheaper variants.
Frame it as a compute-allocation decision - would the same ten-fold budget buy more as a bigger model, higher resolution, or a real ensemble - and decide where in the product a small accuracy gain justifies the serving bill.
## What test-time augmentation is Everywhere else on this topic, augmentation is a training-only device and evaluation runs on a fixed deterministic pipeline. Test-time augmentation is the deliberate exception. At inference you build several views of the same input, run each through the unchanged model, and combine the outputs into one prediction. Ten-crop is the classic recipe for image classification: take four corner crops and one centre crop of the resized image, then add the horizontal mirror of each. Five crops times two orientations is ten views. Each view is passed through the network and the ten output probability vectors are averaged into the final prediction. Nothing about the weights changes. TTA is a decision-rule change, not a training change. ## Why it helps A single centre crop is one sample of a somewhat arbitrary choice - which part of the frame the model sees and in which orientation. The prediction inherits variance from that choice. Averaging over a set of views is variance reduction on the prediction, the same statistical idea that makes an ensemble of models work, applied to an ensemble of inputs instead. It is cheapest exactly where the model is uncertain: an object that falls near a crop boundary is captured properly by at least some of the views. The gain is usually small. It is also usually real, and it costs no retraining, which is why it appears in evaluation protocols so often. ## Aggregation choices Average the predicted probabilities. Two alternatives are worse or different: - **Hard voting** over the ten argmaxes throws away the margin. Nine views at 0.51 and one at 0.99 for the other class is information you should keep, and votes can tie. - **Averaging logits** before the softmax is a defensible but distinct operation - roughly a geometric rather than arithmetic mean of probabilities. It weights confident views more heavily and shifts calibration. Pick one deliberately and keep it consistent between evaluation and serving. Whatever you choose, the aggregation is now part of the model's contract. Calibration measured on single-crop outputs does not describe the averaged output. ## The cost, stated honestly Ten views is ten forward passes. If they run serially, request latency multiplies by roughly ten and your tail latency is where the pain lands. If you batch the ten views into one forward pass you convert most of the latency hit into a throughput hit, but the compute bill is unchanged: your serving fleet does ten times the work per request, so cost per prediction is ten times higher. That framing is the interview answer. The question is not *does TTA help* - it usually helps a little - but *is this the best use of a ten-fold compute budget?* Often it is not. The same compute spent on a larger model, a higher input resolution, or a genuine ensemble of two or three differently-seeded models typically buys more accuracy than ten crops of one model. ## When it is the right call - **Offline or batch scoring**, where throughput matters and per-item latency does not. - **Evaluation and reporting**, provided you say clearly that the number is a TTA number, since it is not comparable to single-pass results. - **High-stakes, low-volume inference**, where a small accuracy gain is worth real money and traffic is thin. - **Cascades**: run a single pass, and fall back to TTA only for inputs whose confidence lands in an ambiguous band. Most traffic pays one pass; the hard tail pays ten. And when it is not: any tight online latency budget, any high-volume path where compute cost dominates, and any case where a cheaper single-pass change - better training, a longer schedule, a bigger backbone - is still on the table. ## Two discipline points people get wrong **The views must be label-preserving, and should match the trained invariances.** Corner crops assume the object survives cropping; if your objects routinely span the whole frame, corner crops feed the model partial evidence and TTA can hurt. Applying a flip at test time when the model was trained without flips - or, worse, in a domain where mirroring changes the answer - degrades the average. The training policy and the test-time view set should be consistent. **Select with TTA on if you serve with TTA on.** Checkpoint selection, decision thresholds and calibration all depend on the exact inference procedure. Tuning a threshold on single-crop validation outputs and then serving the ten-view average silently shifts the operating point. Whichever procedure ships must be the procedure the validation numbers came from. ## Cheaper relatives If the ten-fold bill is refused but you still want some of the effect: use two views - centre crop and its mirror - for most of the gain at a fifth of the cost; batch views together; or distil the TTA-averaged outputs into a single-pass student so the averaging cost is paid once at training time rather than on every request.
- Why average the predicted probabilities across the ten views rather than take a majority vote?Averaging keeps the margin: nine views at 0.51 and one at 0.99 for the other class carry information a vote discards, and votes can tie. Averaging logits instead is a legitimate but different rule - closer to a geometric mean - that weights confident views more and shifts calibration. Pick one and use it in both evaluation and serving.
- How would you keep most of the benefit without paying ten forward passes per request?Cut the view set to the centre crop plus its mirror, which captures much of the gain at a fifth of the cost. Batch the views into one pass to convert latency into throughput. Or run a cascade: single pass for everything, ten views only for inputs in an ambiguous confidence band. Distilling the averaged outputs into a single-pass model moves the cost to training entirely.
- Can test-time augmentation ever make accuracy worse?Yes. If the views are not label-preserving for the domain, or the model never saw those transforms in training, the extra predictions are drawn from inputs the model handles badly and they drag the average down. Corner crops also hurt when objects routinely fill the frame, since each crop then shows partial evidence.
saying these in an interview costs you the question
- Reports the accuracy gain without measuring latency or cost
- Uses test-time transforms the model was never trained on
- Takes a hard vote over crops and discards confidence
- Tunes the decision threshold single-pass then serves averaged
- Believes test-time augmentation retrains or changes the weights