With squared-error loss, why can one outlier target own almost the entire mini-batch gradient?
answer
- Differentiate the loss, do not eyeball it
- Gradient is linear in the residual
- Averaging scales all terms equally
- Absolute error's gradient magnitude is one
- Quadratic core, linear tail, capped slope
basics
~20 sSquared error's per-example gradient grows linearly with the residual, so an example whose error is 50 times larger contributes about 50 times the gradient. Batch averaging divides everyone equally, so it never dilutes that imbalance.
solid answer
~50 sFor one example, squared error is `L = (p - y)^2`, so the gradient flowing back into the prediction is `dL/dp = 2 * (p - y)` — linear in the residual and unbounded. The batch loss averages those terms, but averaging scales every example by `1/B`, so the relative shares are untouched: an example with a residual 500 units wide against neighbours at 1 unit supplies over 99% of the batch gradient. In a day-ahead electricity-demand model, a single meter outage logged as a 50x spike therefore makes one corrupted row decide the whole update; you see a loss spike, a large weight move, and degraded predictions everywhere else. Absolute error caps each example's gradient magnitude at 1, and Huber keeps the quadratic behaviour for small residuals while capping the tail gradient, which is what actually protects the step.
code
python · 20 lines# Share of a mini-batch gradient owned by one corrupted row.
batch = [(10.0, 11.0), (12.0, 11.5), (9.0, 10.0),
(11.0, 9.5), (500.0, 12.0)] # (target, prediction)
def d_squared(r):
return 2.0 * r
def d_absolute(r):
return 1.0 if r > 0 else -1.0
def d_huber(r, d=1.0):
return r if abs(r) <= d else d * (1.0 if r > 0 else -1.0)
for name, grad in (("squared ", d_squared),
("absolute", d_absolute),
("huber ", d_huber)):
per_example = [grad(p - y) for y, p in batch]
total = sum(abs(g) for g in per_example)
share = abs(per_example[-1]) / total
print(name, "outlier share of batch gradient: %.3f" % share)go deeper
Recall that squared error penalises big misses much harder than small ones, and that its derivative is proportional to the error rather than fixed. Be able to say why one very wrong row can wreck a training step.
Be ready to write dL/dp = 2 * (p - y) from memory, show that batch averaging scales all terms alike, and contrast that with absolute error's constant-magnitude gradient and Huber's capped tail slope.
Show the diagnosis: loss spikes, weight-norm jumps, and lingering damage through an optimizer's running averages. Explain when you fix the data instead of the loss, and what a capped gradient costs you if the extremes are real.
Own the tradeoff between robustness and fidelity. A robust loss quietly changes which statistic your model estimates and hides pipeline defects from monitoring, so decide deliberately whether extremes are corruption to remove or events the business is paying you to predict.
## The loss is quadratic; the gradient is linear A regression head emits a scalar prediction `p` for each example and squared error scores it as `L = (p - y)^2`, with `y` the target. What drives learning is not that number but its derivative with respect to the prediction: ``` dL/dp = 2 * (p - y) = 2 * r ``` where `r = p - y` is the residual. The loss grows quadratically in `r`, the gradient only linearly — and it is the gradient, not the loss value, that is backpropagated into the weights. Both facts matter: the quadratic loss is what makes the *reported* number explode, the linear-but-unbounded gradient is what makes the *update* explode. ## Why averaging does not help The batch objective is usually the mean, `L_batch = (1/B) * sum_i (p_i - y_i)^2`, so each example's contribution to the gradient is `(2/B) * r_i`. The factor `1/B` is applied identically to every example. It shrinks the overall step size; it does not change any example's *share* of the direction. That share is ``` share_i = |r_i| / sum_j |r_j| ``` for squared error. Take a day-ahead electricity-demand regression where residuals are normally around 1 unit, and one row is a meter outage recorded as roughly 50x the true load, giving a residual near 500. Four ordinary examples contribute |r| of about 1, 0.5, 1 and 1.5; the outage row contributes 488. Its share of the batch gradient is about 99%. Increasing the batch size to 256 does not fix this — it only means the corrupted row is diluted by more terms that are each still tiny next to it. ## What actually goes wrong during training The step taken is essentially "lower the prediction for the outage row", which is a direction the rest of the data never asked for. Consequences you can observe: - a visible loss spike at that step and a large jump in weight norm; - degraded predictions on the whole input region that shares features with the corrupted row, because a shared trunk moves with it; - lingering damage with optimizers that keep running averages of gradients or squared gradients: the outlier step poisons those accumulators, so several subsequent steps are distorted too, even on clean batches; - with a large learning rate, an outright divergence to non-finite values. ## How the alternatives behave Absolute error, `L = |p - y|`, has derivative `sign(p - y)`: magnitude exactly 1 for any nonzero residual. The outage row's share collapses to `1/B`. The price is that the gradient stays at full magnitude when the residual is nearly zero, so near convergence the parameters jitter at the scale of the learning rate rather than settling, and the derivative is undefined exactly at zero. Huber is the compromise: quadratic for small residuals, linear beyond a crossover `d`. ``` L = 0.5 * r^2 if |r| <= d L = d * (|r| - 0.5 * d) otherwise dL/dp = r if |r| <= d dL/dp = d * sign(r) otherwise ``` Inside the core it behaves like squared error, so the gradient shrinks smoothly to zero as the fit improves and convergence is clean. Outside, the gradient magnitude is *capped at* `d`, no matter how absurd the residual is. That cap — not the smaller loss value — is what stops one row from owning the update. The two pieces are chosen so the loss and its derivative agree at `|r| = d`, which is why Huber is differentiable everywhere including the crossover. ## The judgment that goes with it A capped gradient is a change to what you are estimating, not just to numerical hygiene: squared error's optimum is the conditional mean, absolute error's is the conditional median, and Huber's sits between them. So reaching for a robust loss is only free when the giant residuals are corruption. If they are real, rare events you are accountable for predicting, capping their gradient means systematically under-predicting them. Diagnose the source first: a sensor fault deserves a data fix, a mis-scaled target deserves a scale fix, and a genuinely heavy-tailed target may deserve a loss that models its spread rather than one that ignores it. One more practical note: the same mechanism is why an unnormalised target hurts. If targets live on the order of thousands, the *typical* residual is already large, so every gradient is large and the learning rate must be shrunk to compensate — the outlier case is just the extreme version of the same scale sensitivity.
- Does raising the batch size reduce the outlier's influence on the update?Not its share. Averaging multiplies every per-example gradient by the same `1/B`, so the corrupted row still supplies the same fraction of the direction; a bigger batch only means more small terms sit alongside it. What a larger batch does change is how often a batch contains such a row, and it shrinks the overall step magnitude, which softens the visible spike without fixing the cause.
- Why does capping the gradient magnitude help more than capping the loss value?The loss value is only a report; the optimizer never sees it. Weights move by the gradient, so a loss that is smaller for huge residuals but still has a huge derivative changes nothing. Huber works because its derivative saturates at the crossover value: past that point every extreme example pushes with exactly the same, bounded force.
- How would you tell a corrupted target from a genuine extreme before changing the loss?Look at the largest residuals as records, not numbers: check whether they cluster on one sensor, one timestamp range, or one ingestion path, and whether the value is physically possible. A meter reading 50x the plant's capacity is corruption. A demand spike that recurs on cold snaps and appears in independent sources is signal, and a robust loss would teach the model to miss it.
One person shouting in a room of quiet speakers. Taking the average of the room does not make the shout quieter — the average is still mostly whatever the shouter said.
saying these in an interview costs you the question
- Says squared error's gradient is constant in the residual
- Thinks averaging over the batch cancels an outlier's influence
- Confuses the quadratic loss value with a quadratic gradient
- Reaches for a robust loss without checking whether the spike is real
- Claims a larger batch size alone solves outlier-dominated updates