Your training loss becomes NaN by epoch three of a gradient-descent fit — how do you diagnose it?
answer
- log per update, not per epoch
- does it grow before it breaks?
- geometric growth with sign flips
- cut the learning rate tenfold
- log of zero is infinite
basics
~20 sLog the loss per update, not per epoch. A loss climbing geometrically with weights flipping sign means the step size is too large: cut the learning rate tenfold. A NaN on the very first update points to bad inputs or a broken loss.
solid answer
~50 sLog the loss per update over the first few dozen steps, not just per epoch. Divergence has a fingerprint: the loss rises geometrically, the weights alternate in sign and grow, and the run hits `inf` then `NaN`. That is a step size too large — a learning rate of 1.0 can reach `NaN` in three epochs on data where 0.01 converges cleanly, with everything else identical. Cutting the learning rate tenfold and rerunning settles it cheaply. If instead the loss is `NaN` on the very first update with no growth phase, look at the data for missing or infinite values, and at the loss code: log loss is infinite when a positive example is predicted at exactly probability 0, so clip predicted probabilities away from 0 and 1. And check the update sign — `w <- w + lr * g` ascends and produces a smoothly rising loss.
go deeper
Recognise that a rising training loss means something is wrong rather than that training needs longer, and know that lowering the learning rate is the first thing to try.
Describe the mechanism: an oversized step overshoots the minimum and lands further out each time, so the weights oscillate with growing magnitude until the loss overflows. Be able to name log-of-zero as a separate source of infinity.
Show a triage order and the evidence at each stage — per-update logging first, then a tenfold learning-rate cut as a decisive experiment, then inputs, then the loss code and the sign of the update. Say what distinguishes each cause.
Own the practice that makes this cheap for everyone: fits log per-update loss early in the run, guard against non-finite loss by failing loudly rather than continuing, and record the step size that worked so the next run does not rediscover it.
## What a diverging fit looks like A gradient-descent fit that is diverging has a signature, and the first job is to confirm you are seeing it rather than guessing. Log the training loss for the first few dozen *updates*, not just per epoch. Divergence looks like: - the loss **increases** from update to update rather than decreasing; - the increase is roughly geometric — each step multiplies the error, so the loss goes 0.7, 1.4, 6, 90, 3e4, then overflows; - the weights alternate in sign and grow in magnitude, because each step overshoots the minimum and lands further out on the opposite side; - the run reaches `inf` and then `NaN` (typically `inf - inf` or `0 * inf` inside the loss) within a handful of epochs. That pattern is the fingerprint of a step size too large for the problem as posed. On a 50-million-row ad file with a raw money-valued feature, a learning rate of 1.0 can send the loss to `NaN` inside three epochs, while 0.01 on exactly the same data and code converges cleanly. Same model, same loss, same rows — the only difference is how far each update moves. ## The triage, in order 1. **Print the per-update loss for the first 50 updates.** If it climbs monotonically, you have divergence, not a data problem. If it is `NaN` on update *one*, before any weight has moved far, it is not the step size — it is the data or the loss code. 2. **Cut the learning rate by a factor of 10 and rerun.** This is the cheapest possible experiment and it decides the question. If 0.1 diverges and 0.01 descends, you are done diagnosing. Repeat the cut once more if needed. 3. **Check the inputs for `NaN` or `inf` before they reach the model.** A single missing value in a feature or a target poisons the average gradient, and then every weight becomes `NaN` in one update — note the tell: the loss is `NaN` *immediately* and stays there, rather than growing first. 4. **Check the loss implementation for a logarithm of zero.** With log loss, `-[y*log(p) + (1-y)*log(1-p)]` is infinite when a positive example gets predicted probability exactly 0, or a negative example gets exactly 1. That happens when the linear score saturates the sigmoid at the limits of floating-point precision — which itself is often *downstream* of weights that have already grown too large. The usual guard is to clip `p` into `[eps, 1-eps]`, or to compute the loss directly from the score in a numerically stable form rather than from the rounded probability. 5. **Check the sign of the update.** `w <- w - lr * g` descends; `w <- w + lr * g` ascends. A flipped sign produces a loss that rises smoothly and relentlessly from the very first update — one of the most common bugs in hand-written fitting code, and indistinguishable from a too-large step until you look at the code. ## Distinguishing the three causes | Symptom | Most likely cause | |---|---| | Loss grows geometrically over the first updates, weights flip sign | Step size too large | | Loss `NaN` on the first update, no growth phase | `NaN`/`inf` in the inputs or targets | | Loss rises smoothly and steadily from step one, no oscillation | Sign error in the update, or maximising instead of minimising | | Loss descends for many epochs, then a single spike to `NaN` | One extreme row or an overflow in the loss at a saturated prediction | ## What not to do Do not respond to a rising loss by training longer: extra epochs on a diverging run only reach `NaN` more thoroughly. Do not immediately reach for a fancier optimiser — a fit that diverges at every step size has a bug, and a fit that converges at a smaller step size never needed one. And do not silently replace `NaN` losses with zeros to "keep training"; that hides the failure and produces weights nobody can defend. ## Recovering the run Once the cause is the step size, the fixes are ordered by how much you have to change: reduce the learning rate; if the loss is still fragile, clip the gradient's norm so a single extreme row cannot produce a huge update; and check whether one or two outlier rows dominate the gradient, since squared error weights a large residual quadratically. Record the learning rate that worked, because it is the number the next person running the fit will need.
- How do you tell a step size that is too large from a NaN in the input data?By whether there is a growth phase. A too-large step produces several updates of visibly increasing loss and oscillating weights before overflow. A `NaN` or `inf` in a feature or target poisons the averaged gradient instantly, so every weight becomes `NaN` on the first update and the loss is `NaN` from the start with no climb. Logging per update, not per epoch, makes the difference obvious.
- Why can log loss produce an infinite value even when the data is clean?Log loss is `-[y*log(p) + (1-y)*log(1-p)]`. If a positive example is assigned predicted probability exactly 0, or a negative one exactly 1, the logarithm is infinite. Large weights push the linear score far enough that the sigmoid rounds to 0 or 1 in floating point. Clipping `p` into `[eps, 1-eps]`, or computing the loss from the raw score in a stable form, prevents it.
- The loss descends for 40 epochs and then spikes to NaN in one step. What do you suspect?Not a globally wrong step size — that would have failed at the start. Suspect a single extreme row producing a huge gradient in one batch, or a prediction that saturated far enough for the loss to overflow. Inspect the batch that preceded the spike, cap the gradient norm so one row cannot dominate an update, and check for outliers, which squared error punishes quadratically.
saying these in an interview costs you the question
- Adds more epochs when the loss is already rising
- Assumes NaN always means missing values in the input
- Reaches for a different optimiser before checking the step size
- Replaces NaN losses with zero to keep training
- Never inspects the loss between epoch boundaries