Your training loss sits flat at the majority-class prior for 20 epochs — what do you check?
answer
- what value is it stuck at
- compute the entropy of the class split
- constant output, not slow descent
- output bias carries the prior
- are gradients alive or zero
basics
~20 sCheck whether the plateau equals the entropy of the class base rate - that means the model emits the prior for every input. Then confirm outputs are constant, and inspect output-bias initialization, saturated units and step size before killing the run.
solid answer
~50 sA flat trace has two very different causes, and the plateau's height tells you which. Compute the entropy of the label distribution: for a 90/10 binary split, predicting the base rate gives about 0.325 nats. If the loss is pinned there, the network has found the constant solution and is not yet using its inputs — that is a different failure from a step size that is merely too small, which shows as a slow but continuing descent. Confirm it by measuring the spread of predictions across a batch; near-zero spread means constant output. Then check the usual causes: an output bias initialized far from the prior's log-odds, saturated or dead activations killing the gradient, unnormalized inputs, a rate too small to escape within your budget, or a target with little signal. Prior plateaus often break on their own, so check whether gradient norms are non-trivial before killing the run.
go deeper
Recognise that a flat loss is not automatically a converged model, and know that a classifier's easiest solution is to predict the class base rate for every input.
Explain how to compute the loss of the constant-prediction solution from the label distribution, and how to confirm constant output by looking at prediction spread rather than at loss.
Show the ordered debugging path — plateau value, prediction spread, per-layer gradient norms, then one change at a time — and judge from liveness signals whether the run deserves more epochs.
Own the stopping policy and the instrumentation that makes it decidable: what liveness signals every run logs by default, and how much compute the team will spend on a stalled run before it is cut.
## What the plateau value tells you The constant-prediction solution is the first thing any classifier finds, because it is easy: ignore the input, output the label base rate. Under cross-entropy, the loss of that solution is exactly the entropy of the label distribution. For a binary split with majority fraction `p` it is `-(p*ln(p) + (1-p)*ln(1-p))`. At a 90/10 split that is about 0.325 nats; at 50/50 it is `ln(2)`, about 0.693. So the first action on a flat trace is arithmetic, not intuition: compute that number and compare it with the plateau. A plateau sitting at it is the constant solution. A plateau meaningfully below it means the model has learned something and then stalled — a different diagnosis. A plateau above it means the model is worse than the trivial baseline, which usually points at a broken loss, mismatched labels, or a rate that is destroying progress. ## Confirming constant output The cheap confirmation is to look at predictions rather than at loss: take one batch and measure the spread of the model's outputs. If the predicted probabilities are nearly identical across very different inputs, the network is not using its inputs at all. That single check separates the prior plateau from every other flat curve, and it is what an interviewer wants to hear before any hypothesis list. ## The usual causes, in the order worth checking **Output layer initialization.** With a randomly initialized output bias, the network starts far from the prior and has to spend its first epochs simply moving the bias to the base rate. On a heavily imbalanced problem that walk can take a long time, and it looks exactly like a stall. Initializing the final bias to the log-odds of the class prior starts training at the plateau instead of on the way to it, so the epochs are spent learning features. **Dead or saturated units.** If activations sit in a flat region — all negative pre-activations through a rectifier, or extreme inputs through a squashing nonlinearity — gradients into earlier layers are near zero and only the output bias can move. Inspect activation statistics per layer and gradient norms per layer; a run where the last layer has healthy gradients and everything earlier is near zero is diagnostic. **Input scaling.** Unnormalized inputs, a feature with a huge scale, or an accidentally constant feature block will pin the early layers. A pipeline bug that hands the model the same tensor every batch produces a textbook prior plateau. **Step size and schedule.** A rate too small to escape in your budget, or a warmup so long that the informative phase never starts, both look like stalls. The distinguishing sign is that the loss is descending, just imperceptibly on the plotted scale — re-plot on a log axis or print four decimals before concluding it is flat. **The target.** Sometimes the labels really are close to unpredictable from the given inputs, and the prior is the best available answer. Check whether a simple baseline on the same features does better than chance at all before blaming the network. ## Why waiting is sometimes right Prior plateaus frequently break on their own. The run spends many epochs on the constant solution, then the loss drops sharply and keeps falling — a classic on small, hard datasets such as a few hundred surgical video clips. The mechanism is that the useful features take time to emerge from noise; until then, the constant solution is a stationary-ish region rather than a true minimum, and the gradient is small but not zero. So the operational question is not only what to change but whether the run is alive. Signs that it is: non-trivial and slowly changing gradient norms, drifting weights, prediction spread that is small but growing. Signs that it is dead: gradient norms at or near zero into every layer but the last, no change in any weight statistic across epochs, prediction spread identically zero. ## The interview register A weak answer restarts with a bigger model. A strong answer does four things in a fixed order: computes the entropy of the label distribution and compares it with the plateau; checks the prediction spread; checks per-layer gradient norms and activation statistics; and only then changes one variable — output bias initialization first, because it is nearly free, then input normalization, then the step size. Throughout, it keeps a fixed rule for how long the run gets before it is killed, so the decision does not depend on how long you happened to be watching.
- How exactly does initializing the output bias to the class log-odds help?It starts the model at the plateau instead of on the road to it. With a random bias the first epochs are spent moving the output toward the base rate, which produces a long flat stretch that looks like a stall and delays every feature-learning gradient. Setting the final bias to `log(p/(1-p))` makes the constant solution the starting point, so early gradients go into the features.
- What distinguishes a prior plateau from a loss that is descending too slowly to see?The plateau's value and the prediction spread. A prior plateau sits at the entropy of the label distribution with near-identical outputs across a batch; a slow descent sits wherever it sits and shows real output variation. Printing the loss at four decimals or plotting on a log axis usually settles it in one epoch.
- You have a fixed compute budget. How do you decide when to kill a plateaued run?Decide the rule before the run: a maximum number of epochs at the prior, plus liveness checks. If gradient norms are non-trivial and prediction spread is growing, the run is working and the budget should be spent. If gradients into every layer but the output are effectively zero across several epochs, waiting cannot help, and the cheap fixes — bias initialization, input normalization, a larger step — should be tried in a fresh short run.
A student who answers the most common option to every question on a test: their score is fixed, respectable, and completely unrelated to the questions in front of them.
saying these in an interview costs you the question
- Kills every flat run without checking the plateau value
- Assumes a flat loss always means the learning rate is too low
- Never inspects whether predictions are constant across a batch
- Calls a prior plateau converged training
- Reaches for a bigger model before checking input normalization