skip to content

Why is a mini-batch gradient an unbiased estimate of the full-batch gradient, and what shrinks its noise?

level: middleimportance: must knowfreq 72%

answer

  1. same average, different spread
  2. expectation over uniform sampling
  3. variance divides by the batch size
  4. four times the batch, half the noise
  5. compute grows linearly, noise falls by a root

basics

~20 s

A uniformly sampled mini-batch gradient has the same expected value as the full-batch gradient, so it is unbiased. Only its noise differs: the standard deviation of that noise falls roughly as one over the square root of the batch size.

solid answer

~50 s

The full-batch gradient is the average of the per-example gradients over the whole training set. A mini-batch gradient is the average over `B` examples drawn uniformly at random, and the average of a uniform sample has the same expectation as the average of the population — so the mini-batch gradient is an unbiased estimate of the full-batch one. What differs is spread. If a gradient coordinate has per-example variance `s^2` across the dataset, the mini-batch average has variance about `s^2 / B`, so its standard deviation goes as `s / sqrt(B)`. That is the whole scaling law: going from batch 32 to batch 512 is sixteen times the examples but only four times less noise, at sixteen times the compute per step. Bigger batches buy a cleaner direction with strongly diminishing returns, which is why nobody just sets the batch as large as memory allows.

code

python · 15 lines
python
import random, statistics

random.seed(0)
# one coordinate of the gradient, one value per training example
per_example = [random.gauss(1.0, 3.0) for _ in range(10000)]
full_batch = statistics.fmean(per_example)

def batch_means(B, trials=3000):
    return [statistics.fmean(random.sample(per_example, B)) for _ in range(trials)]

print('full-batch value:', round(full_batch, 4))
for B in (4, 16, 64, 256):
    means = batch_means(B)
    bias = statistics.fmean(means) - full_batch
    print('B =', B, 'bias:', round(bias, 4), 'noise sd:', round(statistics.pstdev(means), 4))

go deeper

for a junior

Be ready to say what a mini-batch gradient is: the average of the per-example gradients over the sampled examples, used because averaging over the whole set every step is too expensive.

for a middle

You are expected to state both halves out loud — same expectation as the full-batch gradient, standard deviation falling as one over the square root of the batch size — and to do the arithmetic on the spot for a concrete jump like 32 to 512.

for a senior

Show that you read batch size as a noise dial with a compute price, not as a memory setting. Say what the per-example variance depends on in the data you actually train on, and note that it shrinks as the model fits.

for a principal

Own the tradeoff framing: noise reduction is sublinear in compute while step cost is linear, so past some point extra batch buys almost nothing, and the residual noise is a training-dynamics choice your team makes deliberately rather than an artefact to minimise.

## The object being estimated Training minimises an average loss over the training set: `L(w) = (1/N) * sum_i L_i(w)` over `N` examples. Its gradient, the full-batch or exact gradient, is likewise an average: `g_full = (1/N) * sum_i g_i`, where `g_i` is the gradient of example `i`'s loss with respect to the parameters. Computing it costs a pass over all `N` examples, which is why almost nobody computes it during real training. A mini-batch step replaces that average with an average over a sample. Draw a set `S` of `B` indices uniformly at random and use `g_B = (1/B) * sum_(i in S) g_i`. ## Why it is unbiased Unbiased means the expected value of the estimator equals the quantity it estimates: `E[g_B] = g_full`. The argument is one line. Each index in the batch is drawn uniformly, so for any single slot the expected per-example gradient is `(1/N) * sum_i g_i = g_full`. Averaging `B` slots each with that expectation gives the same expectation back, because expectation is linear. Nothing about the model, the loss surface or the current parameters enters — unbiasedness is a property of the sampling scheme, not of the network. Two consequences people miss. First, unbiasedness is a statement about the average over many possible batches, not about any one batch: the batch you actually drew almost certainly points somewhere other than downhill on the full loss. Second, it holds at every point in parameter space, including points where the model is badly wrong, so it does not decay as training proceeds. Sampling without replacement inside a shuffled epoch changes the variance but not the mean: a batch taken from a uniformly shuffled epoch is still a uniform sample of the dataset, and its variance carries a finite-population correction factor of `(N - B) / (N - 1)`, which is negligible whenever `B` is a small fraction of `N`. ## Where the noise comes from and how it scales The noise is sampling variance — the spread of the per-example gradients around their mean. Write the per-example variance of one gradient coordinate as `s^2 = (1/N) * sum_i (g_i - g_full)^2`. The variance of the mini-batch average of `B` independent draws is `s^2 / B`, so the standard deviation is `s / sqrt(B)`. That square root is the fact to have on hand. Quadrupling the batch quarters the variance and halves the standard deviation. Sixteen times the batch cuts the noise by four. Compute per step, meanwhile, grows linearly in `B`. So noise reduction per unit of compute gets worse and worse as the batch grows: the first doubling is cheap relative to what it buys, the twentieth is not. A concrete version of this question shows up on whiteboards: a spoken-command classifier over 40 keywords is trained at batch 32 and then at batch 512. What changes? The expected gradient is identical at both settings — the model, the data and the loss are the same, only the sampler changed. The gradient noise standard deviation falls by a factor of four, because 512 / 32 = 16 and sqrt(16) = 4. Anyone who answers 16 has confused variance with standard deviation. ## What makes `s` itself large The variance is a property of the data and the current parameters, not something the optimizer sets. It is large when examples disagree: heterogeneous or multimodal classes, mislabelled rows, outliers with large-magnitude gradients, badly scaled input features. It is small when the per-example gradients agree — early in training when almost every example pushes the same way, or late when the model fits nearly everything and most per-example gradients are near zero. This last point matters: the noise is state-dependent and shrinks as the model fits, so the same batch size delivers different amounts of noise at different points in the run. ## Why you do not just remove the noise Given the choice, the instinct is to make the estimate as clean as affordable. That instinct is wrong twice over. It is wrong on economics, because of the square root. And it is wrong on dynamics, because the noise is doing work — a stochastic run explores directions the averaged gradient does not point in, which is what lets it leave flat stretches a deterministic run would sit on. The mini-batch gradient is not a cheap approximation you tolerate; it is a different algorithm with different behaviour, and the batch size is the dial that sets how different.

  • Does the unbiasedness argument survive sampling without replacement inside an epoch?
    Yes for any single batch. A batch taken from a uniformly shuffled epoch is still a uniform sample of the dataset, so its expectation is the full-batch gradient; only the variance changes, by a finite-population factor of `(N - B) / (N - 1)` that is negligible when the batch is a small fraction of the set. Batches within one epoch are not independent of each other, but that affects their joint behaviour, not the mean of any one of them.
  • If the noise falls as one over the square root of the batch, why is a bigger batch not always the better use of compute?
    Because cost per step is linear in the batch size and noise reduction is only a square root, so four times the compute buys two times less noise, and the ratio keeps worsening. You also spend that compute on fewer parameter updates per epoch. And the noise is not purely a cost — a run with too little of it can sit on flat stretches that a noisier run walks off.
  • What in the data makes the per-example gradient variance large in the first place?
    Disagreement between examples. Multimodal or heterogeneous classes, mislabelled rows, and outliers whose gradients are large in magnitude all widen the spread of per-example gradients around their mean, and unnormalised input features do it coordinate by coordinate. The variance is a property of the dataset and the current parameters — the optimizer does not choose it, it only chooses how many samples to average over.

Polling a city: asking 400 people instead of 100 does not change who the city favours on average, it only narrows the margin of error, and it narrows it by half, not by four.

saying these in an interview costs you the question

  • Calls the mini-batch gradient a biased approximation of the true gradient
  • Says doubling the batch halves the gradient noise
  • Answers 16 when the batch grows 16-fold and the question asks about standard deviation
  • Claims a larger batch changes the expected gradient direction
  • Treats gradient noise as pure cost with no benefit
  • Confuses how smooth the loss curve looks with whether the estimator is biased

context