Your team quadrupled the batch size, kept the epoch budget, and the run got worse — why?
answer
- epochs fix examples, not updates
- count the updates per epoch
- two causes, one rerun separates them
- compare curves against examples processed
basics
~20 sTwo effects are being confused. A four-fold batch at the same epoch count means four times fewer updates, and an unscaled learning rate leaves each update the same size, so the run underfits. Scale the rate, then judge.
solid answer
~50 sSplit the report into two independent causes before believing anything about batch size. First, the **step-budget** effect: a fixed epoch count with a four-fold batch means a quarter as many parameter updates. Second, the **rate** effect: an unscaled learning rate leaves each of those fewer updates the same size, so the run makes roughly a quarter of the progress per epoch and underfits. The diagnostic is to plot training loss against examples processed, not against steps, and rerun with the rate scaled by four for momentum SGD, ramped in at the start. If the scaled run tracks the baseline, the original result said nothing about batch size. If it still lags after a rate sweep, you are past the batch size at which extra parallelism buys steps, and the answer is more epochs or a smaller batch, not a bigger rate.
go deeper
Remember that one epoch is the dataset size divided by the batch size, so a bigger batch means fewer updates per epoch. Check whether the learning rate was scaled before blaming anything else.
Explain the two effects separately — fewer updates and unscaled step size — and state the arithmetic for each. Know that comparing curves against steps rather than examples flatters the large-batch run.
Design the isolating control: one rerun at the scaled, ramped rate plotted against examples processed, followed by a short sweep. Distinguish a spiky opening from a smoothly underfit curve before choosing a fix.
Set the norm that batch-size changes are never reported without the matching rate change and an honest compute axis, so the team does not accumulate folklore from uncontrolled comparisons.
### What was actually reported 'We went from batch 512 to batch 2,048, kept the same number of epochs, and the model came out worse.' Taken at face value this sounds like evidence about batch size. It usually is not. Three distinct things changed at once, and only one of them is about the batch. ### Effect 1 — the step budget collapsed Epochs fix the number of **examples** processed, not the number of **updates**. With `N` training examples and batch `B`, one epoch is `N / B` updates. Quadruple `B` at fixed epochs and you have taken a quarter as many optimizer steps. Even with a perfectly scaled learning rate, you have handed the optimizer four times fewer opportunities to react to what it has seen. Whether that costs you anything depends on where you are relative to the batch size at which extra examples per step stop translating into fewer steps needed — below it, four times fewer steps is fine because each is four times more informative; above it, it is a real loss. ### Effect 2 — the rate was probably not scaled This is the common case and it is pure arithmetic. If the rate stayed at its old value, then per epoch the run performs a quarter as many updates of the *same* size. Total distance travelled through weight space per epoch is roughly a quarter of what it was. The run does not diverge; it does not look broken; it produces a smooth, well-behaved, **underfit** curve. Teams then attribute the gap to the batch size when the correct statement is 'we accidentally cut the effective step size per epoch by four.' The fix under momentum SGD is the linear scaling rule: rate times four, brought up gradually at the start rather than applied at full size from step zero. ### Effect 3 — the opening may have been unstable If the rate *was* scaled but applied immediately, the first few hundred iterations may have spiked or diverged and the run limped along afterwards from a bad region. This shows up as a loss curve with a visible early excursion, and it is a different failure from a smoothly underfit curve. It is worth checking the first epoch's loss trace before drawing any conclusion, because the two failures have opposite fixes: one wants a larger rate applied more carefully, the other wants a larger rate at all. ### The diagnostic sequence 1. **Plot loss against examples processed, not against steps.** Against steps, the large-batch run always looks better per step and worse per unit of compute; against examples the comparison is honest. If the two curves overlap, there is no problem to solve. 2. **Rerun with the rate scaled** by the batch factor (full factor for momentum SGD, square root for an adaptive optimizer), ramped in over the opening. This one control usually resolves the whole report. 3. **Sweep the rate over a factor of two either side** of the scaled value. The rule sets the range; it does not replace tuning. 4. **Only then** ask whether the batch is too large in the sense that matters — that extra examples per step have stopped reducing the number of steps needed to hit the target loss. That is a property of the training run, not of the code, and the response is a longer epoch budget or a smaller batch, never a still-larger rate. ### What not to say Do not accept 'large batches train worse' as a conclusion from an uncontrolled comparison. Nearly every such report from a team that has not scaled the rate is explained by the step budget and the unscaled rate, and both are recoverable with one rerun. The claim only becomes interesting after the rate has been scaled, ramped and swept, and the epoch budget has been made honest — and by then you have usually converted a vague complaint into a concrete statement about how many updates the run actually needs. ### Why interviewers like this scenario It separates candidates who know the linear scaling rule as a slogan from candidates who can decompose a regression into independent causes and design the control that isolates each one. The strong answer names both effects, states which is more likely, and specifies the exact rerun that distinguishes them.
- What single control run would you launch first to settle this?The same configuration at batch 2,048 with the learning rate multiplied by the batch factor and brought up gradually over the opening, plotted against the batch-512 baseline with examples processed on the horizontal axis. If the curves overlap, the original report was a rate-scaling mistake. If the scaled run still trails after a short sweep around that rate, the batch is genuinely past the point where more examples per step buy fewer steps.
- How would you tell an unstable opening apart from a plain underfit run?Look at the loss trace over the first epoch at full resolution. An unstable opening shows a spike or a divergence and then a recovery from a worse region; an underfit run is smooth throughout and simply sits above the baseline everywhere. The fixes differ: the first wants the same scaled rate applied more gradually, the second wants a larger rate at all.
- The team proposes fixing it by raising the learning rate further. When is that wrong?Once the rate has already been scaled by the batch factor and swept, a further increase is not addressing the remaining gap. If extra examples per step have stopped reducing the number of steps to target, the run is short of updates, and no step size recovers updates you did not take. The remaining levers are more epochs, a smaller batch, or a different form of update such as a layer-wise adaptive scheme.
saying these in an interview costs you the question
- Concludes large batches train worse from an uncontrolled run
- Forgets that fixed epochs means fewer updates
- Compares the two runs against step count
- Keeps raising the rate after it is already scaled
- Changes batch size and epoch budget in the same experiment