skip to content

Train and Validation Traces

Loss against epochs has readable shapes: a divergent spike from too high a rate, a flat crawl from too low, a plateau, and a validation trace that turns up while training loss keeps falling.

on this pageshow

questions

4

How do you tell a too-high learning rate from a too-low one by the shape of the training loss curve?

level: juniorimportance: must knowfreq 82%

answer

  1. shape, not height
  2. band versus slope
  3. steps too big overshoot
  4. still falling at the last epoch

basics

~20 s

A learning rate that is too high drops the loss fast, then leaves it oscillating inside a band or diverging; one that is too low gives a smooth, near-linear descent still falling at the last epoch.

solid answer

~50 s

Read the trend of the epoch-mean training loss, not the per-step jitter. Too high: the loss falls fast for a few epochs and then bounces inside a band without descending further, or shows upward spikes, or blows up. The band exists because each update overshoots along the sharpest directions and because gradient noise is scaled by the step size, so the loss settles at a floor proportional to the rate. Too low: the loss descends smoothly and almost linearly, the band is narrow, and it is clearly still falling at the last epoch — nothing is broken, you have simply underfit the budget. The cheap confirmation is a short sweep: run a few hundred steps at rates separated by factors of about three and keep the largest one whose loss still descends smoothly, then decay it as training proceeds.

go deeper

for a junior

Be ready to name the two shapes on sight: bouncing inside a band or blowing up means the rate is too high, a slow straight descent that has not finished means it is too low.

for a middle

Explain why the shapes happen — overshoot along sharp curvature directions, and a noise floor that scales with the step size — and describe a rate sweep that picks a value in a few hundred steps.

for a senior

Show diagnostic discipline: check the epoch mean rather than the jitter, change one knob per run, and recognise when a run that was stable becomes unstable later because the local curvature grew.

for a principal

Own the policy: what the default schedule is across the team's models, how rate and batch size are re-tuned together when either changes, and how much compute the org spends on rate search versus on more training.

## What the plot actually is A training trace puts training loss on the vertical axis and epochs (or optimizer steps) on the horizontal axis. It says nothing about generalization — that requires the validation trace — but it is the fastest read on whether the optimizer is being asked to take steps of a sensible size. ## Why the step size shows up as a shape Gradient descent updates parameters as `w <- w - lr * g`. Approximate the loss near the current point by a quadratic bowl whose curvature along some direction is `h`. Along that direction the update multiplies the distance to the minimum by `(1 - lr * h)`. If `lr < 1/h` you creep in; if `lr` sits between `1/h` and `2/h` you overshoot but still contract, giving a zig-zag; if `lr > 2/h` the distance grows and that direction diverges. A real network has a spread of curvatures, so a single rate is comfortable for the flat directions and marginal for the sharpest ones — which is exactly why the trace usually shows partial progress plus bouncing rather than a clean blow-up. Stochastic gradients add a second effect. Each mini-batch gradient is a noisy estimate, and the noise enters the update multiplied by the rate. Once the descent has taken you near a basin, the run does not stop; it random-walks in a region whose loss floor grows with the learning rate and shrinks with batch size. That floor is the band you see. ## The too-high signature - A fast initial drop, then an oscillation inside a band whose floor does not fall for many epochs. - Occasional upward spikes: single batches whose curvature is locally sharper than average. - In the worst case, the loss rises monotonically or becomes very large. The useful test is not the width of the jitter but whether the epoch mean is still trending down. Per-step noise averaged over a whole epoch mostly cancels; a genuine rate problem survives that averaging. ## The too-low signature - A smooth, near-linear or gently curved descent, tight around its trend line. - Still visibly falling at the final epoch. - Validation loss tracking it closely, because the model has not had enough effective optimization to memorize anything. This run is not failing. It is under-trained for the budget you gave it, and the fix is either a larger rate or more epochs. On a retail demand-forecasting series a 60-epoch trace that descends painfully but consistently is telling you the budget, not the model, is the binding constraint. ## The staircase When you decay the rate mid-run on a run that had settled into a band, the loss drops sharply and then flattens into a new, lower band. That step is not the model escaping a local minimum; it is the noise floor falling with the step size. The staircase is the healthiest common shape in deep learning traces. When a decay stops buying a visible step, you are at the limit of the model, the data, or the remaining optimization signal. ## What to change, in order 1. Confirm the trend on the epoch mean, and confirm you are looking at training loss, not a mixed plot. 2. Sweep the rate over a handful of values separated by factors of about three for a few hundred steps each; take the largest rate whose curve descends smoothly rather than the one with the single lowest point. 3. Add a short warmup if the very first steps spike, then a decay schedule so the band floor falls as you approach a basin. 4. Change one thing at a time. Changing the rate and the architecture in the same run makes the next trace uninterpretable. ## Confounders worth naming in an interview - Batch size and learning rate interact: the same rate can be stable at a large batch and unstable at a small one, because the noise term scales with the ratio. - A rate that was fine early can become too high later as the local curvature grows; that shows up as a run that was smooth for many epochs and then starts bouncing. - Spikes that land at the same point in every epoch usually mean a data ordering or a particular batch, not the rate. - Normalization layers change the effective curvature, so a rate imported from a different architecture is a guess, not a default.

  • After you decay the learning rate mid-run, the loss drops sharply and then flattens again. Is that healthy?
    Yes — that is the expected staircase. The run had settled into a band whose floor is set by gradient noise scaled by the step size, so shrinking the step drops the floor and the loss with it. Judge the schedule by whether each successive decay still buys a visible step; once a decay produces no drop, further decays are just freezing the current solution.
  • How do you separate a genuinely unstable learning rate from ordinary mini-batch noise?
    Compare per-step loss with the epoch mean. Mini-batch noise largely averages out over an epoch, so a noisy per-step curve with a steadily falling epoch mean is fine. If the epoch mean itself refuses to descend or trends upward, the step size is the suspect. Confirm by halving the rate for two epochs: if the band floor drops, the rate was the constraint.

Walking down a narrow valley: strides that are too long bounce you off the opposite wall forever, while tiny shuffling steps go straight down but never reach the bottom before dark.

saying these in an interview costs you the question

  • Treats per-step jitter as evidence of divergence
  • Assumes the smallest learning rate is always the safe choice
  • Never checks whether the loss is still falling at the final epoch
  • Changes learning rate and architecture in the same run
  • Reads generalization off the training curve alone

context

open as a page

Why can a network's validation loss sit below its training loss during the first few epochs?

level: middleimportance: should knowfreq 58%

basics

~20 s

Validation loss below training loss is usually a measurement artifact: training loss is measured on a network handicapped by dropout and augmentation and averaged across the epoch, while validation is measured at the epoch's end on the full network.

open as a page

Your training loss sits flat at the majority-class prior for 20 epochs — what do you check?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Check whether the plateau equals the entropy of the class base rate - that means the model emits the prior for every input. Then confirm outputs are constant, and inspect output-bias initialization, saturated units and step size before killing the run.

open as a page

Validation cross-entropy is rising while validation accuracy still improves — what explains it?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Cross-entropy scores confidence, accuracy scores only the decision. As training continues the model becomes more confident everywhere, so the examples it still gets wrong contribute a large and growing penalty while newly-correct ones add almost nothing.

open as a page