Five seeds of a ranking model span 0.6 AUC points - how do you judge a claimed +0.3 gain?
answer
- compare distributions, not two numbers
- the effect is smaller than the noise
- same seed list for both arms
- paired differences cancel shared luck
- seeds needed scale with the squared ratio
basics
~20 sA single run cannot resolve an effect half the size of the seed spread. Rerun both the baseline and the candidate over the same list of seeds, look at the per-seed differences, and report that distribution rather than the best run.
solid answer
~50 sThe claim is unmeasurable as stated: one run of each arm cannot separate a +0.3 effect from a spread of 0.6 points across seeds. Rerun both arms on a *shared, pre-declared* seed list and analyse the per-seed differences, not the difference of the best runs - pairing cancels whatever the seed controls in common and usually shrinks the noise you have to beat. A range of 0.6 over five runs implies a per-arm standard deviation around 0.25 points, so resolving 0.3 needs roughly five to eight seeds per arm and resolving 0.2 needs around a dozen; the count scales with the square of the ratio, so halving the effect you want to detect quadruples the runs. Also freeze everything else - data snapshot, evaluation set, stopping rule, checkpoint-selection rule - because picking the best checkpoint per run adds its own optimistic bias. And treat the spread itself as a finding: you ship one seed.
go deeper
Recall that a model's score depends on its random seed, so a single run of two models cannot settle which is better when the seed-to-seed differences are as big as the gap.
Explain the mechanics of a paired comparison: one shared seed list, per-seed differences, and the standard error shrinking as one over the square root of the number of seeds.
Show you would design and defend the protocol - fixed data snapshot, fixed stopping and checkpoint rules, pre-declared seeds, a stated seed count derived from the spread, and a report built on the difference distribution.
Own the decision framing: what effect size is worth shipping, how much compute the organization will spend to resolve it, and whether a volatile candidate should be stabilized before it is even considered for adoption.
## Why one run against one run says nothing Every trained model is one draw from a distribution induced by initialization, data order and per-step stochastic choices. When five seeds of the baseline span 0.6 AUC points, that distribution is wide relative to the effect being claimed. A candidate that scores +0.3 above one baseline run is entirely consistent with the candidate being identical to the baseline and simply having drawn a good seed - and equally consistent with it being worse and having drawn a very good seed. The measurement has less resolution than the quantity being measured. ## Turn a range into a usable spread A range is not a standard deviation. For a bell-shaped spread, the expected range of five draws is roughly 2.3 standard deviations, so a 0.6-point range over five seeds implies a per-arm standard deviation of about 0.25 points. Work in standard deviations from there, because that is what the arithmetic below needs. With more seeds, just compute the standard deviation directly rather than inferring it from the extremes - the range is a noisy statistic that grows as you add runs. ## Design the comparison as paired Run both arms over the *same* pre-declared seed list, and compute the difference per seed rather than comparing summaries. Two reasons: - **Variance cancellation.** If seed 7 draws an unusually favourable data order, both arms benefit, and the difference is unaffected. The standard deviation of the paired differences is often meaningfully smaller than the per-arm standard deviation, which buys resolution for free. The cancellation is only partial - two architectures consume randomness differently, so the same seed does not mean the same nuisance draws - but partial is still worth having. - **Honesty.** A pre-declared list makes it impossible to quietly drop the seed where the candidate lost. ## How many seeds The standard error of the mean paired difference is `s_d / sqrt(n)`, where `s_d` is the standard deviation of the per-seed differences. Treating "about two standard errors" as the resolution threshold, detecting an effect `d` needs roughly n ~ (2 * s_d / d)^2 With `s_d` around 0.3 points, detecting 0.3 needs about five to eight runs per arm, and detecting 0.2 needs around a dozen. The quadratic is the part to internalize: **halving the effect you want to resolve quadruples the runs**. This is why a team that argues about 0.1-point differences on single runs is not doing measurement at all - the budget for that claim is an order of magnitude beyond what they spent. ## Freeze everything that is not the treatment Seed variance is the noise you can see. Silent confounders are worse: - **Data snapshot and version** - identical for both arms. - **Evaluation set and metric computation** - identical, and large enough that evaluation noise is small relative to training noise. - **Stopping rule** - fixed in advance, not "stopped when it looked good". - **Checkpoint selection** - the rule that picks the reported checkpoint must be identical, and it is itself a source of optimism: choosing the best of fifty validation evaluations per run inflates every arm, and inflates the noisier arm more. - **Hyperparameters** - if the candidate got a tuning pass and the baseline did not, you are measuring tuning, not architecture. ## What to report Report the paired differences: their mean, an interval, and how many seeds favoured the candidate. Three of five seeds favouring the candidate by less than the spread is not a result. Reporting the maximum over seeds is the failure mode to avoid - the maximum is an estimate of the best seed's luck, not of the model's skill, and it grows as you add seeds even when nothing improves. ## The variance is itself a result Production ships one model from one seed. A candidate whose seed-to-seed spread is 0.6 points is a fragility finding regardless of its mean: whichever seed you ship is a draw from that spread, and the number you promised stakeholders came from a different draw. Sometimes the right recommendation is not "adopt" or "reject" but "reduce the variance first" - a longer warmup, a lower final learning rate, more data per step, or averaging over the last checkpoints - because a stable mediocre model can be worth more than a volatile slightly-better one. ## The answer an interviewer wants Say plainly that the comparison as presented cannot support the claim; describe the paired multi-seed rerun with a fixed protocol; give the rough seed count and note that it scales quadratically; and mention that the spread itself is information about the candidate, not just an obstacle to measuring it.
- How many seeds would a 0.2-point claim need here?With a per-seed difference standard deviation near 0.3 points, resolving 0.2 needs roughly a dozen runs per arm, using n about (2 * s_d / d)^2 as the rule of thumb. The scaling is quadratic, so a 0.1-point claim needs about four times that again. If the budget will not stretch that far, the honest move is to stop claiming effects at that resolution.
- You can only afford three extra runs in total. What do you do?Spend them on the single comparison you would actually act on, both arms on shared seeds, and reduce noise elsewhere instead of buying it back with compute: a larger evaluation set, a fixed stopping rule, averaging the metric over the last few checkpoints rather than picking the best one. Then report the result as provisional and say so explicitly.
- Is a large seed spread ever the finding rather than the obstacle?Yes. Production ships one seed, so a 0.6-point spread means the deployed model is a lottery draw around the mean you reported. That can justify recommending stabilization - longer warmup, lower final learning rate, more data per step - over adopting a candidate whose mean is slightly better but whose variance is untouched.
- Why is reporting the best seed's score particularly misleading?The maximum over seeds estimates the luck of the best draw, not the model's typical skill, and it drifts upward as you add seeds even when nothing has improved. It also cannot be reproduced by whoever ships the model, since they get one seed rather than the best of several. Report the central tendency and the spread instead.
Claiming a 0.3-point improvement from single runs with a 0.6-point spread is like timing a sprinter with a stopwatch you can only read to the nearest second and announcing a new record.
saying these in an interview costs you the question
- Accepts a gain smaller than the seed spread from single runs
- Compares the best run of each arm instead of paired differences
- Uses different seed lists for baseline and candidate
- Treats the observed range as if it were a standard deviation
- Tunes only the candidate and calls the difference architectural
- Ignores that best-checkpoint selection inflates each arm