skip to content

Your logistic fit's loss drops 1e-9 per epoch while the weight norm doubles — what is happening?

level: seniorimportance: should knowfreq 38%

answer

  1. two symptoms that contradict each other
  2. not slow learning, no bottom
  3. check absolute loss and training accuracy
  4. hunt for a zero cell in a cross-tab
  5. rare one-hot level with three rows

basics

~20 s

Those two symptoms together are the signature of separated data, not slow learning. The loss is creeping toward zero along a direction with no finite optimum, so the weights grow without bound and extra training will not help.

solid answer

~50 s

A loss still falling, however slightly, while the weight norm keeps doubling says the optimiser has found a direction that is downhill forever — the classic picture of complete or quasi-complete separation. Slow learning looks different: the gradient shrinks, the weights settle, and the loss plateaus at a clearly non-zero value. Here the loss is heading to zero and training accuracy is already perfect. To confirm it, check the absolute loss level and the training accuracy, then find the culprit: cross-tabulate each categorical level against the label looking for cells with zero rows of one class — a one-hot column for a city with three rows in it is the usual suspect. The fixes are on the data or estimator side: pool rare levels, drop an artefact column, gather rows for the empty cell, or fit with a penalty instead of plain maximum likelihood. Raising the iteration cap is not a fix.

go deeper

for a junior

Recall that a tiny loss improvement is not the same as convergence, and that growing weights alongside near-zero loss is a warning sign rather than a good result. Check training accuracy before celebrating.

for a middle

Explain why the two symptoms together rule out benign explanations, and name the cheap confirmations: absolute loss level, training accuracy, and a refit with more epochs returning a larger coefficient.

for a senior

Demonstrate the full diagnosis on a real dataset: sort weights by magnitude, cross-tabulate levels against the label for empty cells, check the feature-to-row ratio, then choose between pooling, dropping, collecting data, or a penalised fit — and justify the choice.

for a principal

Own the guardrails: convergence criteria that check the weight norm and not only the loss delta, level-count minimums before encoding, and a standing rule about unpenalised maximum likelihood on wide or sparse data so this never reaches production undiagnosed.

## Read the two numbers together Neither symptom alone tells you much. A loss falling by 1e-9 an epoch could be a fit that has essentially converged. A growing weight norm could be a model legitimately sharpening its decision boundary early in training. Together they are diagnostic, because they are contradictory under every benign explanation: a converged fit does not keep doubling its weights, and a fit still making real progress does not improve by 1e-9. The explanation that reconciles them is that the optimiser is walking out along a ray on which the loss decreases forever and reaches zero only in the limit. That is separation: some feature or combination splits the classes with no overlap, so scaling the weights up always improves the fit and nothing ever pulls back. ## Distinguishing it from the alternatives | Symptom | Separation | Genuinely slow learning | Converged | |---|---|---|---| | Loss level | creeping toward ~0 | plateaued well above 0 | stable | | Weight norm | growing without bound | roughly stable | stable | | Gradient magnitude | shrinking, but never to zero | small | ~0 | | Training accuracy | 100% | below ceiling | at its ceiling | | Refit with more epochs | larger coefficients | slightly better loss | identical result | The cheapest single check is the **absolute loss value plus training accuracy**. A loss near zero with 100% training accuracy on data that is not trivially easy means separation until proven otherwise. The second cheapest is refitting with double the epochs: a converged model reproduces itself; a separated one hands back a bigger coefficient. ## Finding the culprit 1. **Sort the weights by magnitude.** The separating column is usually the largest by an order of magnitude, and it is the same one growing each epoch. 2. **Cross-tabulate each categorical level against the label.** You are looking for a cell with zero rows of one class. A one-hot column for a city with 3 rows in it, all positive, is the archetype — the coefficient for that level grows every epoch because nothing contradicts it. 3. **Check level counts before anything else.** Any level with a handful of rows is a candidate; the fewer the rows, the likelier the empty cell. 4. **For continuous features, sort and look at the ends.** If one class occupies a contiguous tail of a feature with no interleaving, that feature separates. 5. **Check the feature-to-row ratio.** With as many features as rows, separation happens by dimension counting alone and there may be no single culprit to find. If the pathology is confined to one subgroup — a handful of positive cases all sharing an indicator — you have the quasi-complete case, which is easier to overlook because most coefficients look perfectly normal. ## Fixing it - **Pool rare levels** into an "other" bucket. The empty cell disappears, both labels appear within every level, and the fit gains a finite optimum. This is usually the first thing to try and it often costs nothing in accuracy. - **Drop the column** if it is an artefact — an identifier, a near-duplicate of another feature, or something that cannot be present at prediction time. - **Collect more rows** for the empty cell when the level genuinely matters and you can afford to wait. - **Fit with a penalty** rather than plain maximum likelihood, which keeps the coefficient finite by construction. - **Stop deliberately** if you must ship as-is: fix a maximum weight norm or an iteration count and record it, so that at least the arbitrary stopping point is documented rather than accidental. ## What does not work Raising the iteration cap, tightening the convergence tolerance, or switching optimiser are all attempts to converge to something that is not there. They change how far along the ray you travel, nothing else. Standardising the features helps conditioning but does not create a finite optimum. And reporting the run as "converged" because the loss delta fell below a threshold is the worst outcome — the threshold was crossed by an asymptote, not by an optimum. ## Why it matters beyond the training run While the fit is separated, three things are unreliable. The coefficient on the separating feature is an artefact of when you stopped. Its standard error is meaningless, because the likelihood is nearly flat along the ray, so any inference or significance claim about that variable is void. And the predicted probabilities are pinned near 0 and 1, so anything downstream that consumes them as probabilities — a threshold tuned on expected value, a risk score, an expected-cost calculation — is receiving overconfident garbage even where the ranking happens to be fine. That is why this is worth diagnosing rather than tolerating: the model can look excellent by accuracy and be unusable by every other measure.

  • How would you tell this apart from a fit that is simply learning slowly?
    Look at the loss level and the weight norm together. Slow learning plateaus at a loss clearly above zero with the weights settling and accuracy short of its ceiling. Separation creeps toward zero loss with the weights growing without bound and training accuracy already perfect. Refitting with more epochs settles it: slow learning improves a little, separation returns a larger coefficient.
  • You pooled the rare city levels into 'other' and the fit converged. What changed?
    The empty cell disappeared. Each rare level previously contained rows of only one class, so its indicator could be scaled up without cost. Once pooled, the combined level contains both labels, some rows are necessarily misclassified, and the loss now increases if the weight grows too far — which is exactly what creates a finite optimum.
  • The stakeholder wants to ship it anyway because training accuracy is 100%. What do you say?
    That 100% is a symptom, not an achievement — the model has found a column that reproduces the label on this sample, and its confidence is an artefact of where training stopped. Show the held-out numbers, show that refitting changes the coefficient, and offer the pooled or penalised version, which will look worse on training accuracy and behave better in production.

saying these in an interview costs you the question

  • Concludes the model has converged because the loss delta is tiny
  • Blames the optimiser and switches to a different one
  • Raises the iteration cap and calls it fixed
  • Says a local minimum trapped the fit
  • Reports the runaway coefficient as a significant effect
  • Never checks the absolute loss level or training accuracy

context