skip to content

Why does the loss fall smoothly with full-batch gradient steps but rattle with stochastic ones?

level: seniorimportance: should knowfreq 62%

answer

  1. right on average, wrong each time
  2. variance falls with batch size
  3. descent holds only in expectation
  4. the true gradient vanishes, the noise does not
  5. hovering region, not a fixed point

basics

~20 s

A full-batch step uses the exact gradient, so with a small enough step size the loss decreases every iteration. A stochastic step uses a sampled gradient that is correct only on average, so individual steps can go uphill and the iterates settle into a noise ball rather than a point.

solid answer

~50 s

Full-batch descent evaluates the gradient over every example, so the direction is exact and, with a step below the stability threshold, the objective decreases monotonically — the curve is smooth by construction. A mini-batch gradient is the average over a random sample; if the sample is drawn uniformly it is an **unbiased** estimate of the full gradient, but it carries variance roughly proportional to `1 / b` for batch size b. So each step is a descent direction only in expectation, and any single step can raise the loss. Near the optimum this matters most: the true gradient goes to zero while the sampled one does not, so the iterates stop converging and hover in a region whose size grows with the step size and the gradient noise and shrinks with batch size. Part of the visible wiggle is also that each reported mini-batch loss is measured on different examples, not just that the parameters moved.

go deeper

for a junior

Recall that full-batch steps use all the data and give a smooth curve, while mini-batch steps use a random sample and give a noisy one that still trends downward.

for a middle

Explain unbiasedness of the sampled gradient, that its variance falls like one over the batch size, and why descent then holds only in expectation rather than at every step.

for a senior

Diagnose real runs: separate sampling noise from instability, recognise the noise floor at the end of training, and use a step-size cut as the test that tells the two apart.

for a principal

Own the compute trade-off — batch size, step size and hardware utilisation are one joint decision, and set the convention for what the team plots and how a plateau is adjudicated before more compute is spent.

## Two different objects called the gradient The training loss is an average over n examples, `F(theta) = (1/n) * sum_i f_i(theta)`. Full-batch gradient descent computes `grad F` exactly, using all n examples per step. Stochastic gradient descent computes the average gradient over a random subset of b examples and treats it as a stand-in for `grad F`. If the subset is drawn uniformly at random, that mini-batch gradient is an **unbiased estimator**: its expected value, over the draw of the batch, equals the full gradient. Unbiased does not mean accurate on any single draw. If per-example gradients have variance sigma squared around the full gradient, the average over b of them has variance about `sigma^2 / b`, so its standard deviation falls like `1 / sqrt(b)`. Quadrupling the batch halves the noise. ## Why full-batch is monotone and stochastic is not For a smooth objective with curvature bounded by L, a full-batch step with `eta < 2 / L` is guaranteed to decrease the objective at every iteration — the direction is exactly downhill and the step is short enough not to overshoot the descent. The loss curve is therefore smooth and monotone by construction; any bump in it is a bug or a step size past the threshold. A stochastic step points downhill only *in expectation*. On a given iteration the sampled direction can be tilted enough that the true objective goes up. Averaged over many iterations the drift is downward, which is why the curve trends down while individual points scatter around the trend. There is a second, purely cosmetic source of wiggle that candidates often miss: the number usually plotted is the loss of the current mini-batch, computed on different examples every step. Some batches are intrinsically harder than others, so the plotted series jitters even when the parameters barely move. Plotting a running average, or evaluating a fixed held-out set periodically, separates the two effects. ## The noise ball The important consequence is what happens at the end. Approaching the optimum, the full gradient shrinks to zero, so full-batch steps shrink to nothing and the iterates settle. The stochastic gradient does **not** shrink to zero: even at the exact minimiser of the average loss, individual examples still pull in different directions, so the sampled gradient has magnitude around `sigma / sqrt(b)`. With a constant step size the iterates keep taking steps of that size forever and end up wandering inside a region around the optimum rather than converging to it. The size of that region grows with the step size and with the gradient noise, and shrinks as the batch size grows. That gives a clean operational reading: if the smoothed loss curve has flattened, halve the step size. If the plateau drops to a new lower level, you were sitting at the noise floor. If it does not move, you are genuinely near an optimum of the objective you are optimising. ## Why anyone accepts the noise Cost accounting. One full-batch step costs n per-example gradient evaluations. For the same compute you can take `n / b` stochastic steps. Early in training the parameters are far from any optimum, and a rough estimate of the direction is more than good enough to make progress; many cheap approximate steps beat one exact step by a wide margin. Only near the end, where the true gradient is small compared with the noise, does the accuracy of the estimate start to matter, and by then most of the useful progress has been made. This is why stochastic steps dominate at scale even though full-batch steps have the stronger per-iteration guarantee. There is a second practical reason: a full-batch step requires touching the whole dataset before any parameter changes, which is impossible for data that does not fit in memory or arrives as a stream. A mini-batch step needs only b examples at a time. ## Diagnosing a run - Noisy but downward-trending curve, with a smoothed version falling: normal and healthy. - Noisy curve whose smoothed version is flat: at the noise floor or at an optimum — distinguish by cutting the step size. - Noise amplitude growing over time: the step size is too large for the local curvature; this is instability, not sampling noise. - Curve that looks smooth but stalls high: check that the batches are actually being resampled, and that the estimate is unbiased — sorting the data and reading it in order makes consecutive batches systematically unrepresentative and breaks the unbiasedness that the whole argument rests on. ## What to say in an interview Name the three pieces in order: the mini-batch gradient is unbiased but has variance about sigma squared over b; steps therefore descend only in expectation, so individual steps can go uphill; and near the optimum the noise does not vanish, so with a constant step the iterates hover in a region whose radius grows with the step size and shrinks with the batch size. Then close with the reason it is worth it — n over b times as many steps for the same compute.

  • Is a mini-batch gradient a biased estimate of the full-data gradient?
    Not if the batch is a uniform random sample — its expectation equals the full gradient exactly, which is what makes the method work at all. Bias creeps in when the sampling is not uniform: reading a sorted or grouped file in order, oversampling one class without correcting the weights, or reusing a fixed subset all produce batches that are systematically unrepresentative.
  • How does quadrupling the batch size change the noise in the gradient estimate?
    The variance of an average of b independent per-example gradients falls like 1 / b, so the standard deviation falls like 1 / sqrt(b). Quadrupling b halves the noise. Note the return is sublinear in compute: four times the work per step for a factor of two in noise, which is why very large batches stop paying for themselves.
  • The smoothed training loss has flattened. How do you tell a real optimum from the noise floor?
    Cut the step size by a factor of two to four and keep training. If the plateau drops to a new lower level, you were sitting in the noise ball, whose size scales with the step. If the level does not move, the optimiser is genuinely near a stationary point of the objective, and the remaining gap is a modelling or data problem rather than an optimisation one.

Polling fifty random voters instead of the whole electorate: the answer is right on average, but each fresh sample lands somewhere different, and no amount of repeating removes the wobble unless you enlarge the sample.

saying these in an interview costs you the question

  • Says the mini-batch gradient is biased toward the sampled examples
  • Thinks a noisy loss curve always means the run is broken
  • Believes a constant step size converges to a point under sampling noise
  • Claims larger batches reduce noise proportionally to batch size
  • Ignores that each plotted mini-batch loss is measured on different data

context