Five checkpoints from one run: ensemble their predictions or collapse them into one averaged weight vector?
answer
- training cost is identical either way
- the decision lives at inference
- diversity is what the ensemble sells
- one artifact versus five
- compare against seed-to-seed spread
basics
~20 sPrediction ensembling usually scores highest but multiplies inference cost and memory by five. Collapsing the same checkpoints into one averaged weight vector keeps single-model serving cost and captures part of the gain. Decide from the serving budget, then measure both.
solid answer
~50 sBoth options are free at training time - the snapshots come from one budget - so the decision is entirely about inference. Prediction averaging keeps five parameter sets and runs five forward passes per request, buying the larger accuracy gain because the snapshots make partly independent errors; it costs five times the memory and, unless you can parallelize, roughly five times the latency. Weight averaging collapses them into one vector at exactly the original serving cost, but it captures only part of the ensemble's gain and requires that the snapshots be linearly connected plus a pass to re-estimate normalization statistics. My default is to measure both against the single best checkpoint, and against seed-to-seed variance, before believing either. If the ensemble's extra gain is real and the latency budget cannot absorb it, distilling the ensemble into one model is the third option worth costing.
go deeper
Know the two ways to use several checkpoints - average their predictions, or average their weights into one model - and that only the second keeps inference cost unchanged.
Be able to state the mechanism behind each: output averaging cancels uncorrelated errors and needs diversity, weight averaging cancels optimization noise and needs the checkpoints to sit in one region.
Show the measurement discipline - three candidates scored through the serving path, a seed-variance noise floor, a snapshot-disagreement diagnostic, and the statistics pass after collapsing.
Own the decision as a serving-budget and operational-surface call rather than an accuracy contest, name distillation as the fallback when the accurate option does not fit, and state the default plus the evidence that would overturn it.
## The setup A snapshot ensemble spends one training budget and saves the model at several points along the way - for instance five checkpoints harvested at the ends of successive cycles of a rate that is repeatedly raised again. The appeal is that the checkpoints are free: no additional training runs, no additional data. What is not free is what you do with them at serving time, and that is the decision. ## Option A - average the predictions Keep all five parameter sets. At inference, run all five and average their output distributions. - **Accuracy:** normally the best of the options. It works because the snapshots make partly different errors; averaging cancels the uncorrelated part. The gain scales with how much they disagree. - **Cost:** five parameter sets resident in memory, five forward passes per request. On a throughput-bound service you can sometimes batch across the five and hide part of the cost; on a latency-bound service you cannot hide it unless you have spare parallel capacity. - **Operational weight:** five artifacts to version, load, warm and roll back together. Any one of them going missing changes the model silently. ## Option B - collapse into one weight vector Average the five parameter vectors into one and serve that. - **Accuracy:** typically better than the single best checkpoint, typically worse than the prediction ensemble. It cancels the optimization noise but cannot exploit functional diversity, because there is only one function left. - **Cost:** identical to the original single model - one artifact, one forward pass, no memory increase. That is the entire argument for it. - **Preconditions:** the checkpoints must lie in one low-loss region connected by straight lines, which holds for checkpoints from one run after the trajectory has settled but not in general; and any normalization running statistics must be re-estimated with one forward-only pass over training data afterwards, or the collapsed model can score near chance. ## How to decide 1. **Start from the serving constraint, not the leaderboard.** If the request budget has no room for four extra forward passes and you cannot run them in parallel, option A is not on the table and the discussion is over. State this before anyone runs an experiment. 2. **Measure three numbers, not two.** The single best checkpoint, the collapsed average, and the prediction ensemble - on the same held-out data, through the same serving path. 3. **Measure the noise floor too.** Retrain with two or three seeds and look at the spread of the single-model score. Any gain smaller than that spread is not a result. This is the step most often skipped, and it is why many reported ensembling gains evaporate. 4. **Check diversity before you pay for it.** If the snapshots agree on almost every input, the ensemble cannot help much. The disagreement rate between snapshot pairs is a one-line diagnostic and predicts the ensemble gain better than intuition does. Snapshots taken after the rate has been raised again diverge far more than consecutive epochs do. 5. **Price the whole cost, not just compute.** Five artifacts mean more failure modes, more warm-up, a bigger image, slower deploys and a rollback story with five moving parts. A one-artifact model that is slightly worse is often the better engineering decision. ## The third option If the ensemble is genuinely better and the budget genuinely cannot absorb it, train a single student model to match the ensemble's output distribution. It costs an extra training run but yields one artifact at one forward pass, often recovering a good share of the ensemble's advantage. Whether that extra run is worth it depends on how long the model will be in production - a model serving for a year amortizes a distillation run trivially; a weekly-retrained model may not. ## What to say in an interview The answer an interviewer wants is not 'ensemble, it is more accurate'. It is: the training cost is identical so the decision is a serving-cost decision; here is what each option costs in latency, memory and operational surface; here is the measurement plan including the seed-variance baseline; and here is the fallback if the accurate option does not fit. Then state a default - collapse to one vector unless the measured ensemble gain clearly exceeds the noise floor and the budget has room - and say what evidence would change it.
- What determines whether prediction-averaging the snapshots is worth its extra cost?How much the snapshots disagree. Averaging outputs cancels the uncorrelated part of their errors, so if they agree on nearly every input there is almost nothing to cancel and you pay five forward passes for noise. Measure the pairwise disagreement rate on held-out data before committing. Snapshots harvested after the learning rate is raised again are far more diverse than checkpoints from consecutive epochs.
- You measure a 0.3-point accuracy gain from the ensemble. What do you check before shipping it?Whether 0.3 points is larger than the spread you get from retraining the single model with different seeds. If two seeds of the baseline differ by half a point, the ensemble gain is inside the noise and you have measured nothing. Also confirm the gain survives on a second held-out split and through the real serving path, not just in the evaluation notebook.
- The collapsed average is far worse than any individual snapshot. What are the two likely causes?First, stale normalization running statistics - the averaged weights never ran a forward pass, so the stored estimates do not match; recompute them with one forward-only pass over training data. Second, the snapshots were not linearly connected, for example harvested too early or across a restart that moved the run elsewhere; interpolate between two of them and look for a loss barrier.
- When would you spend an extra training run distilling the ensemble instead of shipping either option?When the ensemble's gain is real and needed, the latency budget cannot absorb five passes, and the model will be in production long enough to amortize the run. Distillation trains one student to match the ensemble's output distribution, giving one artifact at one forward pass. For a model retrained weekly the extra run recurs and the arithmetic is much less attractive.
saying these in an interview costs you the question
- Assumes the collapsed average always matches the ensemble
- Forgets that five checkpoints mean five forward passes
- Reports a gain without a seed-variance baseline
- Ignores the memory and rollback cost of five artifacts
- Skips re-estimating statistics after collapsing
- Proposes averaging weights from checkpoints of different runs