With a fixed training budget for a bake-off of eight variants, how do you split runs between variants and seeds?
answer
- two stages with two different jobs
- screen broadly, confirm narrowly
- a one-seed table is not a ranking
- the winner was picked partly for luck
- enter the baseline twice as a control
basics
~20 sScreen broadly, confirm narrowly. Give every variant one short run to eliminate clear losers, then promote two or three survivors to a paired multi-seed confirmation. Never ship on the screening ranking - a one-seed leaderboard reorders when rerun.
solid answer
~50 sTreat it as two stages with different jobs. Screening runs cheaply and broadly - one seed, a shortened schedule - and its only job is to eliminate variants that are clearly worse than the noise floor; it must never be reported as a ranking, because the top row of a one-seed table was chosen partly for a lucky seed and will reorder on rerun. Confirmation takes the two or three survivors and spends the rest of the budget on paired multi-seed runs against the baseline, with a pre-declared seed list. Buy resolution the cheap ways first - larger evaluation sets, fixed stopping and checkpoint rules, shared data order across arms - because seeds buy it quadratically badly. My favourite control is a placebo row: enter the baseline twice under different seeds. If the placebo beats itself by the margin someone is claiming, the bake-off cannot resolve that margin, and everyone can see it.
go deeper
Understand that a table of one run per variant does not reliably say which variant is best, and that promising candidates need rerunning before anyone acts on the ordering.
Be able to describe the two-stage shape - cheap short screening runs to eliminate, then multi-seed paired confirmation on the few survivors - and why the second stage is where the decision happens.
Show the operating detail: conservative elimination cuts, cheap variance reductions before extra compute, fixed stopping and checkpoint rules, and a placebo entry that calibrates what the harness can resolve.
Own the allocation and the norms - how the budget splits, what effect size the product actually needs, what the team is allowed to publish from a screening run, and whether a volatile winner should be shipped at all.
## The failure this question is really about Eight variants, one run each, sorted by score, ship the winner. It looks efficient and it is close to worthless. The row that comes first was selected on a quantity that mixes real skill with seed luck, and selection on a noisy quantity favours whatever was luckiest. Two consequences follow: the winner's screening number is an *overestimate* of what it will deliver on a fresh seed, and the ordering of the table is unstable - rerun every row under a second seed and the winner frequently changes. Anyone who has run the same bake-off twice has seen the table shuffle. ## Stage the budget **Stage 1 - screening.** Every variant gets one seed and a shortened schedule. The job is *elimination, not ranking*: discard variants that lose by clearly more than the noise floor, and label everything else as surviving rather than ordered. Short runs are acceptable here precisely because you are looking for large effects; be aware that a shortened schedule can be unfair to variants that need longer to pay off, which is a reason to keep the screening cut conservative rather than aggressive. **Stage 2 - confirmation.** Two or three survivors go head-to-head with the baseline on a shared, pre-declared seed list at full schedule, analysed as per-seed differences. The rest of the budget lives here. If the budget cannot fund a real confirmation, the correct call is to screen fewer variants, not to skip the confirmation. Successive halving generalizes this idea: many candidates at short budgets, progressively fewer at longer ones. The important part is not the exact schedule but the discipline that *the decision is made in the final stage only*. ## Why the split leans toward screening breadth Seeds are an expensive way to buy resolution. The standard error of a mean falls with the square root of the run count, so the runs required scale with the square of the noise-to-effect ratio - halving the effect you can resolve quadruples the compute. Spreading seeds thinly across all eight variants therefore buys almost nothing: you get eight badly-resolved estimates instead of one well-resolved comparison. Concentrating seeds on the few candidates that could plausibly win is strictly better use of the same compute. Before spending compute on seeds at all, spend the cheap variance reductions: - A larger or less noisy evaluation set, so evaluation noise stops competing with training noise. - Identical data order across arms within a seed, so part of the nuisance variation cancels. - A fixed stopping rule and a fixed checkpoint-selection rule; picking the best of many validation evaluations inflates every arm, and inflates the noisier arm most. - Averaging the reported metric over the last few evaluations instead of taking the peak. ## The placebo row Enter the baseline into the bake-off twice, under two different seeds, unlabelled. It costs one run and it calibrates the whole table: the gap between the two placebo entries is a direct, visible measurement of what the harness cannot distinguish. If they differ by 0.4 points, then a 0.3-point claim about any other row is not a claim the bake-off can adjudicate, and no one has to win that argument in the abstract. This single control changes the conversation from statistical argument to shared observation, which is why it belongs in a lead's toolkit. ## The norms to set, not just the arithmetic - **Pre-declare the seed list** for confirmation runs, so nobody drops the seed where their variant lost. - **Label screening tables provisional** in the tool that renders them, not in a footnote people skip. - **Require an effect that matters, not merely one that is detectable.** With enough compute anything measurable is detectable; the question the business is asking is whether the gain justifies the complexity, latency and maintenance the variant adds. - **Publish the spread alongside every number.** A culture that reports single runs will keep producing improvements that vanish on retraining, and the credibility cost lands on the team. - **Count the variance as a property of the candidate.** Production ships one seed; a volatile winner is a worse product than a stable near-winner, and that tradeoff is a leadership call rather than a statistical one. ## What a strong answer sounds like Name the two stages and what each is for; say explicitly that the screening table is not a ranking and explain why selection on noise inflates the winner; give the quadratic argument for concentrating seeds rather than spreading them; list the cheap variance reductions you would take before buying more runs; and finish with the process norms - pre-declared seeds, published spread, a placebo control, and an effect-size bar set by the product rather than by the statistics.
- Why is the top row of a one-seed leaderboard an optimistic estimate of its own skill?Because it was selected on a score that mixes skill with seed luck, and selecting the maximum of noisy quantities preferentially picks the luckiest draw. On a fresh seed the same variant regresses toward its true mean, so its screening number does not reproduce - and the ordering below it often changes too.
- How would you defend spending the budget unevenly to a team that wants every variant treated fairly?Fairness is in the elimination rule, not in equal compute. Every variant gets the same screening chance and the same conservative cut; the extra seeds go to whichever candidates survive, because a well-resolved comparison of the plausible winners is worth more than eight unresolved ones. Equal compute spent everywhere buys nobody a decision.
- What would make you recommend the stable second-place variant over the volatile winner?That production ships one seed. If the winner's spread is wide, the deployed draw may land below the runner-up, and the number promised to stakeholders came from a different draw. When the mean gap is inside that spread, the lower-variance candidate is the better product and the easier system to operate and retrain.
saying these in an interview costs you the question
- Ships the winner of a single-seed leaderboard
- Spreads seeds evenly across all variants and resolves nothing
- Confuses screening with a decision-ready ranking
- Buys compute before taking the free variance reductions
- Requires only statistical significance, ignoring whether the gain matters