When is gradient clipping the wrong fix for a run whose gradients keep exploding?
answer
- it bounds the step, not the reason
- the run survives, so nobody looks
- one bad row, every few hundred steps
- a breaker that trips constantly
- log the batch behind each clip event
basics
~20 sWhenever the huge gradient has a findable cause. Clipping bounds the step but leaves the cause running, and because the run now survives, nobody investigates. It is a seatbelt for rare tail events, not a substitute for finding what produces them.
solid answer
~60 sClipping caps the length of an update; it does nothing about why the update wanted to be that long. The usual causes are concrete and fixable: a corrupted training row with an absurd target, a learning rate too large for the architecture, a loss that takes a logarithm or a reciprocal of something approaching zero, unscaled inputs or targets. Once clipping is on, each of those still fires — it just no longer wrecks the run, so no one looks. Consider one bad row producing a gradient with a norm near 1e6 every few hundred steps: clipped, it becomes a full-length step in a direction determined entirely by garbage, applied silently forever. My rule is to keep clipping on as insurance against genuine tail events, but treat a clip that fires often as a defect report: record the step and batch identifiers whenever the pre-clip norm is far above the median, replay those batches, and fix what you find. Clipping that hides an investigation is technical debt with a stability story attached.
go deeper
Know that gradient clipping limits how big a step can be but does not repair whatever made the gradient huge, so a stable-looking curve is not by itself evidence that the problem is gone.
Be able to list the usual causes behind an exploding gradient — a corrupted example, too large a learning rate, a numerically fragile loss, unscaled targets — and say which fix addresses each.
Demonstrate the investigation: instrument clip events with step and batch identifiers, replay the offending batches, and act on what you find rather than accepting a run that merely survives.
Own the tradeoff and the policy: when a bounded step is legitimate insurance against an intrinsic heavy tail, when frequent clipping should block a result, and whether a fixed threshold belongs in a shared training configuration at all.
## The trap Gradient clipping is unusually satisfying to apply. A run that was diverging stops diverging, the loss curve becomes well behaved, and the change is one line. That is exactly why it deserves suspicion: it converts a loud failure into a quiet one. The mechanism is a bound on the size of a step, and a bound on the size of a step says nothing about the correctness of the quantity being bounded. ## Causes that clipping does not remove **A corrupted training row.** Suppose one example in the training data has a target off by six orders of magnitude — a units error, a sentinel value that was never filtered, a mislabeled record. Every time that row appears in a batch, the loss for that batch is enormous and the gradient norm reaches something like 1e6. Without clipping, the run visibly blows up and someone goes looking. With clipping, the update is rescaled down to the threshold — but it is still a *full-length* step, and its direction is still whatever that one absurd row wanted. Every few hundred steps the model takes a maximum-size step in a garbage direction, and no metric complains. The model ends up measurably worse than it should be, with nothing in the training logs pointing at why. **A learning rate above what the architecture tolerates.** Here the large gradients are a consequence of the parameters being repeatedly overshot into bad regions. Clipping damps the symptom while leaving the run bouncing around a region it should not be in. The right fix is the learning rate — or the initialization scale, or the normalization that would have kept activations in range. **A numerically fragile loss.** A logarithm of a probability that reaches zero, a division by a quantity approaching zero, a square root at the origin: these produce genuinely unbounded derivatives. Clipping bounds them, so training continues, but the model is now being trained through a region where the gradient it receives is a truncated version of an infinity. The fix belongs in the loss — a floor on the argument, a numerically stable formulation, a reparameterization. **Unscaled inputs or targets.** If one input feature or the target lives on a scale thousands of times larger than the rest, the resulting gradients are correspondingly large by construction. That is a preprocessing bug wearing a stability costume. ## Why it is not free even when it works Global-norm clipping preserves the direction of an individual minibatch gradient, but it is a nonlinear function of that gradient, so the *average* clipped gradient is not a rescaled copy of the true full-data gradient. When it fires on a small tail of steps the bias is negligible. When it fires routinely, the process being run is no longer stochastic gradient descent on the stated loss — it is descent on something adjacent, and the difference is systematic rather than random. Anyone claiming an optimization result on a heavily clipped run should be asked about that. ## The judgment call The defensible position is not "never clip" — rare tail events genuinely happen, and a run destroyed at hour nine of ten by one pathological batch is an expensive way to be principled. It is: 1. **Keep clipping on as insurance,** calibrated to the tail so it almost never binds. 2. **Instrument the clip as a signal, not just a guard.** Whenever the pre-clip norm exceeds a high multiple of its running median, record the step index and the identifiers of the examples in that batch. That log is the attribution trail. 3. **Replay the recorded batches.** Look for outlier targets, sentinel values, degenerate lengths, duplicated records. A handful of rows is usually the whole story, and removing them beats clipping around them. 4. **Treat "clipping fires constantly" as an unresolved defect,** never as a solved problem. Frequent clipping means the tail assumption is false, and the reason has not been found. 5. **Resist the shared default.** A threshold written into a team-wide training configuration propagates a number that was calibrated for one model to every model that inherits it, and it makes stability look like a settled question that nobody needs to measure again. The short version to say out loud: clipping is a bound on damage, not an explanation. It belongs in a run for the same reason a circuit breaker belongs in a building — and a breaker that trips every few minutes is a report about the wiring, not a working solution.
- The clip fires on 40 percent of steps and the loss is still falling. Do you ship that run?No, or at least not without an explanation. At that rate the threshold is binding on the body of the distribution, not the tail: most updates now have the same length, and the average update is a biased version of the loss's gradient. A falling loss shows the run is not broken, not that it is right. I would find whether the driver is a data tail, the learning rate or the loss formulation before treating the result as trustworthy.
- How do you find the specific batch behind a rare clip event?Instrument the clip itself. Whenever the pre-clip norm exceeds a high multiple of its running median, record the step index and the identifiers of the examples in that batch. Then replay those batches offline and inspect them for outlier targets, sentinel values, degenerate sequence lengths or duplicated records. Rare spikes are usually a small, findable set of rows, and the attribution log is what turns a statistic back into examples.
- Is there a case where clipping genuinely is the fix rather than a patch?Yes, when the tail is intrinsic rather than a defect. Some losses and some data distributions produce legitimately heavy-tailed per-batch gradients with nothing wrong in any individual example. There the occasional huge step is real, not corrupted, and bounding it is the correct engineering choice. The distinguishing evidence is that you looked at the responsible batches and they were ordinary.
A circuit breaker is the right thing to install and the wrong thing to rely on: one that trips every few minutes is telling you about the wiring, not solving it.
saying these in an interview costs you the question
- Treats clipping as the standard cure for exploding gradients
- Says stability is proven because the loss stopped diverging
- Never asks what produced the enormous gradient
- Assumes a clipped bad batch is a harmless batch
- Ignores that heavy clipping biases the average update