A teammate says Adam beat SGD with momentum after tuning only Adam's rate. What do you challenge?
answer
- one variable, not two
- the untuned arm is a straw man
- both arms need a schedule and a sweep
- decay coefficient is not portable across rules
- three seeds before believing one point
basics
~10 sChallenge the tuning asymmetry first: an optimizer given a rate search cannot be compared with one whose rate was inherited. Then demand schedule parity, per-optimizer regularization tuning, equal epoch budgets, and multiple seeds.
solid answer
~40 sThe result as stated is uninterpretable, because the only thing varied properly was Adam. Momentum's best rate is typically far larger than Adam's, so an inherited rate is almost guaranteed to be wrong, and an untuned momentum run is a straw man. I would ask for an equal tuning budget per arm — the same number of trials searching each optimizer's own rate — plus a decay schedule for both, since a constant-rate momentum run loses by construction. Weight decay must be re-tuned per optimizer rather than copied, because its effective strength interacts with the step-size rule. Everything else stays fixed: batch size, epochs, augmentation, initialization, the early-stopping criterion, and the validation split. Finally I would ask how many seeds ran. A one-point validation difference from one seed each is noise, not a finding.
go deeper
Know that comparing two optimizers means changing only the optimizer, and that a learning rate good for one is usually wrong for the other.
Be able to list the controls: matched rate sweeps, a decay schedule on both arms, identical data pipeline and epoch budget, and a metric chosen before the runs start.
Demonstrate that you would push back on the result rather than adopt it, name the specific confounds you suspect, and propose a cheap rerun that would actually settle the question.
Own the standard: define what counts as a valid optimizer comparison for the team so that recipes are not chosen by anecdote, and decide when a measured gap is too small to be worth anyone's tuning time.
## Why the claim is empty as stated An optimizer comparison is a controlled experiment with one intended variable: the update rule. The report described here varied two things — the update rule and how carefully each arm's hyperparameters were chosen — and attributes the outcome to the first. That is the single most common defect in optimizer folklore, and it is why the literature is full of contradictory claims about which optimizer wins. The asymmetry has a specific shape here. Adam's step is normalized, so its useful learning rates cluster in a small range and a default often works. Momentum's step is proportional to the raw gradient, so its useful rate depends on the architecture, the loss scale, the batch size and the normalization layers, and it is typically one to two orders of magnitude larger than Adam's. A momentum run handed a rate borrowed from an adaptive recipe will crawl. Concluding "Adam is better" from that is like timing two runners after tying one of their shoelaces. ## The protocol you should ask for **Equal tuning budget per arm.** Fix a number of trials — say the same count for each optimizer — and let each search its own rate over its own sensible range. The claim being tested is "best achievable Adam versus best achievable momentum," so both arms must actually be searched. (How the search itself is organized is a model-selection question in its own right; here the point is simply that the budgets match.) **Schedule parity.** Both arms need a decay schedule run to completion. A constant-rate momentum run is not the practitioner's configuration, and it loses to almost anything. Report the final numbers, not a mid-run snapshot, because the ordering commonly changes during the decay phase. **Regularization tuned per optimizer, not copied.** Weight decay strength is not comparable across step-size rules: the amount of shrinkage a given coefficient produces depends on how the optimizer scales its updates. Copying a value from one arm to the other silently changes the regularization strength and can produce the entire reported difference. **Everything else held fixed.** Same data pipeline and augmentation, same batch size, same initialization scheme, same number of epochs, same early-stopping rule, same validation split, same metric. If wall-clock rather than epochs is the budget, say so explicitly, and note that per-step cost is nearly identical for the two rules, so epochs and wall-clock rank the arms the same way here. **Seeds and spread.** Deep training runs vary run to run. Three to five seeds per arm, reported as mean and spread, is the minimum for believing a gap of about a point; a single seed each cannot distinguish a real effect from initialization luck. If the spread within an arm is as large as the gap between arms, there is no result to report. **One metric decided in advance.** Choosing after the fact between top-1, top-5, loss and a calibration number is a way to find a winner in any experiment. Name the deciding metric before the runs start. ## What a good answer sounds like Do not turn this into a defence of momentum. The right posture is that you do not yet know which optimizer wins on this task, because the experiment as run could not tell you. Offer the cheap version of the fix: rerun with a small matched rate sweep per optimizer, both with the same decay schedule, three seeds each, and compare the final validation means. That is usually a day of compute and it converts an opinion into a measurement. Also be prepared to say when the answer does not matter. If the gap after a fair comparison is smaller than the seed spread, the honest conclusion is that the optimizers are equivalent on this task, and the choice should then be made on operational grounds — tuning effort, robustness across the other models the team trains, and how much the recipe is likely to be re-tuned when the architecture changes.
- How many seeds do you want before believing a one-point validation gap?At least three per arm, five if the compute allows, reported as mean and spread rather than best-of. If the within-arm spread across seeds is comparable to the between-arm gap, the correct conclusion is that the two optimizers are indistinguishable here. Reporting the best seed of each arm manufactures differences reliably.
- Why must the weight decay coefficient be re-tuned per optimizer rather than copied?Because the shrinkage a given coefficient actually applies depends on how the optimizer scales its updates, so the same number means different regularization strengths under the two rules. Copying it makes regularization strength a second uncontrolled variable, and it can account for the whole reported difference on its own.
- Does an equal epoch budget guarantee an equal compute budget?For the training runs themselves, near enough: the per-step cost of the two update rules differs only slightly, so equal epochs is close to equal wall-clock. The real budget asymmetry is in the tuning trials, not the runs, which is why the number of search trials per arm is the thing to equalize explicitly.
saying these in an interview costs you the question
- Accepts the comparison because the winning arm was tuned
- Uses one optimizer's learning rate for the other
- Leaves the momentum arm at a constant learning rate
- Copies the weight decay coefficient across both arms
- Reports a one-point gap from a single seed per arm
- Picks the deciding metric after seeing the results