skip to content

Which zero-denominator bugs make a loss return NaN on only some batches?

level: middleimportance: should knowfreq 50%

answer

  1. the denominator is decided by the data
  2. an empty mask means a zero count
  3. zero over zero, not zero
  4. epsilon inside the square root
  5. zero times NaN is still NaN

basics

~20 s

Denominators the data decides: a mean over a mask that is empty for that batch, a normalization by a channel of zero variance, an average over an element with no neighbours. Each is zero over zero, which gives NaN.

solid answer

~50 s

The pattern is a denominator the data decides. A masked mean divides the sum by the number of valid elements, and one batch where a filter leaves zero valid elements gives `0/0`. A per-sample normalization divides by a standard deviation, and a constant channel — a stuck sensor, a padded region — makes it exactly zero, with a zero numerator too, so again `0/0`. An average over a neighbour set is NaN for an element that has no neighbours. It is intermittent precisely because it depends on batch composition, which is why smoke tests miss it. The guards: clamp the count to at least one, and add epsilon *inside* the square root of the variance, not after it, because the square root's derivative is already infinite at zero. And note that multiplying the bad output by the mask does not help — zero times NaN is NaN.

go deeper

for a junior

Remember that zero divided by zero gives NaN, not zero, and that a count or a variance computed from the batch can legitimately be zero even when the code reads correctly.

for a middle

Be ready to explain why the failure is intermittent — it depends on what is in the batch — and where the guard belongs: clamp the count to at least one, and put epsilon inside the square root of the variance.

for a senior

Demonstrate that masking an output does not remove a NaN, and that your response is a degenerate-batch test asserting finiteness plus logged per-batch denominators, not a scattering of epsilons.

for a principal

Argue for where these guards live: one shared safe-divide and safe-normalize used everywhere, plus data contracts that turn an empty group into a loud validation failure rather than a silently reweighted objective.

## The shape of the bug A fixed denominator cannot surprise you. A denominator computed from the batch can, and that is the entire family. Three instances cover most real occurrences. **The empty mask.** A masked mean is written as the sum of the elementwise product of values and mask, divided by the sum of the mask. That denominator is a count. Anything upstream that can filter a batch down to nothing — a length filter, a confidence threshold, a validity flag, a padding-only sequence, a class that happens to be absent — makes the count zero. The numerator is zero too, because every term was masked out, so the result is `0/0`, which is NaN rather than the zero a reader might expect. **The zero variance.** A per-sample or per-channel normalization subtracts a mean and divides by a standard deviation. If the channel is constant across whatever axis you reduce over — a sensor stuck at its floor, an all-padding region, a one-hot feature that is absent in this batch, an image patch of uniform colour — the variance is exactly zero. Numerator and denominator are both zero, and you get NaN again. **The empty neighbourhood.** When a model averages over a set of related elements — the neighbours of a node in a batched graph, the members of a group in a grouped pooling — an element with an empty set divides a zero sum by a zero count. Batching makes this worse, because a sampled subgraph or a filtered batch can strip the last neighbour from an element that had several in the full data. ## Why it is intermittent, and why tests miss it All three depend on what happens to be in the batch. With a shuffled loader, an all-empty mask or a fully constant channel might occur once in tens of thousands of batches. Overfitting to a tiny fixed batch — the standard first smoke test — deliberately uses a benign, hand-picked batch and will never surface it. So the failure looks random, appears at a large step count, and appears to correlate with nothing. The fix for the *testing* gap is a deliberately degenerate test case: a batch whose mask is entirely false, a feature column with a single repeated value, an element with no neighbours. Assert that the loss and every gradient are finite. That test is cheap and it is the only thing that makes the bug reproducible on demand. ## Guarding correctly The naive guard is to add a small epsilon to every denominator. It works often enough to be tempting and it hides which denominator was zero, so prefer a guard that says what it is defending. For a count, clamp it: divide by the maximum of the count and one. When the count is zero the numerator is zero, so the term contributes zero, which is usually the intended meaning of "no valid elements here". Be explicit about whether an empty batch should contribute zero to the loss or should be an error — silently contributing zero changes the effective batch size and quietly reweights your objective. For a variance, the epsilon must go **inside** the square root: divide by the square root of (variance + epsilon), not by (square root of variance) + epsilon. The reason is the backward pass. The derivative of the square root is one over twice the square root, which is infinite at zero. Adding epsilon after the square root makes the forward value finite while the gradient through the square root is still infinite, so the loss prints a normal number and the gradients are NaN — one of the most confusing versions of this bug. Putting epsilon inside means the square root is never evaluated at zero and both directions stay finite. ## The masking trap A very common attempt at a fix is to compute the unsafe quantity for every element and then zero out the invalid ones by multiplying by the mask. This does not work, because zero times NaN is NaN and zero times infinity is NaN. The NaN survives the multiplication, survives the sum that follows, and propagates into every parameter. The correct pattern is to make the *input* safe before the unsafe operation runs: select a harmless constant wherever the mask is false, then apply the division or logarithm to the already-safe values, and only then mask the result if you still need to. "Compute then hide" fails; "select then compute" works. The same reasoning applies to the backward pass — a branch you never used still contributes its gradient if its output entered the sum, and if that output was NaN, so is the gradient. ## How to spot it during diagnosis When a run is NaN on a rare batch, print the denominators. Log the minimum valid-element count and the minimum variance per batch alongside the loss. If the failing step shows a count of zero or a variance of zero, the diagnosis is done in one line, and you have a metric you can keep watching afterwards.

  • Why does adding epsilon after the square root of the variance still leave a NaN gradient?
    Because the square root is still being evaluated at zero. Its derivative is one over twice the square root, which is infinite there, so the backward pass produces a non-finite value even though the forward value looks fine after epsilon is added. Putting epsilon inside — the square root of (variance + epsilon) — keeps the argument strictly positive, so both the value and the derivative stay finite.
  • You multiply the offending term by the mask so invalid entries are zeroed. Why is the loss still NaN?
    Zero times NaN is NaN, and zero times infinity is NaN. Masking the output cannot remove a NaN that has already been produced; it flows into the sum and into every gradient. Fix it by making the input safe first — substitute a harmless constant wherever the mask is false, then apply the division or the logarithm — so the unsafe value is never computed at all.
  • Why did this survive every test you had?
    Because it depends on batch composition. Unit tests use hand-picked benign batches, and the overfit-one-batch smoke test uses the same batch every time, so an entirely empty mask or a perfectly constant channel never appears. Add an explicit degenerate-case test — an all-false mask, a constant feature, an element with no neighbours — that asserts the loss and gradients are finite.
  • If the count is zero, should the term contribute zero to the loss or raise an error?
    It depends on whether an empty group is a legitimate state of the data or a pipeline defect. If legitimate, contribute zero and be explicit that the effective number of contributing terms shrank, since averaging over batches then quietly reweights those samples. If it is a defect, fail loudly at data validation rather than absorbing it in the loss, where it becomes invisible.

saying these in an interview costs you the question

  • Sprinkles epsilon everywhere without finding the zero denominator
  • Believes masking the output removes an existing NaN
  • Expects a mean over zero elements to return zero
  • Adds epsilon after the square root instead of inside it
  • Assumes a NaN can only come from too large a learning rate
  • Never logs the minimum count or variance per batch

context