A run trains cleanly but validation is far below expectation — how do you decide: data bug or model bug?
answer
- cheapest checks first, model last
- score one checkpoint twice
- run training data through the eval path
- decode a fully transformed batch
- data bugs impose a floor knobs cannot move
basics
~20 sWork outward from the cheapest check. Score the same checkpoint twice to test whether the evaluation path is deterministic, then score training data through that same path, then inspect a fully transformed batch. Only after those come model-side knobs.
solid answer
~50 sOrder the checks by cost, and keep the model untouched until the data and the evaluation path are cleared. First, score one checkpoint twice: if the number moves, the evaluation path is stochastic — regularization such as dropout still sampling masks, or random crops and flips left in the evaluation transform, which typically costs a few points and looks exactly like a generalization gap. Second, score a slice of training data through the same evaluation code; if that is also bad, the harness or the weights are at fault rather than generalization. Third, dump a batch after every transform and look at it: label histogram, value ranges, and whether an augmentation contradicts the task — a horizontal flip on a left-versus-right classifier destroys the label it is meant to preserve. The heuristic separating the families: data bugs impose a floor no model change moves.
go deeper
Know that a bad validation number is not automatically the model's fault, and be able to name two cheap checks: score the same checkpoint twice, and look at an actual batch after the transforms rather than only at the curve.
Explain why an evaluation pass must be deterministic and what dropout does differently in each mode. Be ready to describe scoring training data through the evaluation path and what each of the two outcomes rules out.
Demonstrate an ordered triage under time pressure and the reasoning that ends it: data bugs impose a floor that model knobs cannot move, so two very different configurations landing on the same number redirects you upstream.
Own the standard, not the incident. Decide which of these checks run automatically on every training job, what they cost in time and false alarms, and how a decoded sample and a determinism assertion become part of the definition of a valid run.
## Framing the question The run finished. No exception, no NaN, a loss curve that looks reasonable — and a validation number far below what the task should support. The bug lives in exactly one of three places, and they cost very different amounts to check: 1. **The evaluation path** — the model is fine, the number is measured wrong. 2. **The data that reaches the loss** — the model is fine, it is being taught the wrong thing. 3. **The model or the optimization** — architecture, capacity, learning rate, loss weighting. Engineers reach for (3) first because it is the interesting one. It is also the most expensive to explore and the least often responsible. Work the list in order. ## Step 1: is the measurement stable? Load one saved checkpoint and score the validation set twice, changing nothing. If the two numbers differ, your evaluation path is stochastic and you cannot trust any comparison you have made all week. The usual culprit is regularization left in its training behaviour. Dropout during training zeroes a random subset of units and rescales the survivors so the expected activation is preserved; at evaluation the units are all kept and nothing is rescaled, which computes an approximation to the average over that ensemble of thinned subnetworks. Leave it sampling at evaluation and you are scoring one random thinned subnetwork instead — a few points of accuracy worse, and different every run. This looks *precisely* like a generalization gap: validation below training from the very first epoch, by a stable-ish margin. It is not a generalization gap. It is a measurement bug. The same step catches random crops, flips or colour jitter left in the evaluation transform, and any metric accumulated in a nondeterministic order with numerically sensitive reduction. ## Step 2: separate the harness from generalization Now score a slice of the **training** data through the **evaluation** code path — the same transforms, the same batching, the same metric function. Two outcomes, two different worlds: - Training data scores well through the evaluation path, held-out data scores badly: the harness works and you have a genuine generalization or distribution problem. Now questions about the validation set's labels, its distribution, and about overfitting become legitimate. - Training data also scores badly: the model cannot even reproduce what it was trained on when read through this path. The bug is in the path or in what was loaded — not in generalization at all. This single comparison eliminates the largest class of wrong hypotheses. ## Step 3: look at what actually reaches the loss Most teams never do this, and it is the highest-yield step. Take a batch **after every transform in the pipeline**, invert the encoding, and inspect it as a human would: the images, the text, the numbers, and the labels beside them. What to look for: - **Label histogram.** One or two classes per batch says the order is wrong; a distribution unlike the dataset's says the sampling is wrong. - **Value ranges.** Inputs normalized twice, or not at all, or clipped by an unexpected cast. - **Label-destroying augmentation.** An augmentation is an assertion that the label is invariant to a transformation. When the task is *defined* by the symmetry you are augmenting over, the assertion is false. A horizontal flip on a left-versus-right lane-change classifier turns a left example into a right example while keeping the old label, so the training set contains directly contradictory pairs; the achievable loss floors well above zero and accuracy on the affected distinction sits near chance no matter what you do to the model. The same trap catches vertical flips on digit or arrow recognition and aggressive crops that remove the object being labelled. - **Shapes and dtypes** at the boundary into the loss. ## Step 4: only now, the model With the measurement stable, the harness cleared and the batch inspected, model-side hypotheses are worth spending on: capacity, learning rate and schedule, initialization, loss weighting across terms, the objective itself. ## The discriminating heuristic Data bugs impose a **floor** the model cannot cross, because the information needed is absent or contradictory. Model bugs respond to model knobs. So the sharpest single signal is this: run two variants that *should* behave differently — a much larger model, a very different learning rate — and see whether the result moves. If a ten-times-larger model and a ten-times-smaller learning rate both land on the same disappointing number, stop touching the model. Something upstream is capping you, and no architecture will lift it. A related smell is a validation metric that is suspiciously *stable* across configurations that should have separated. Identical results are rarely a coincidence; they usually mean the thing you changed never reached the computation. ## Make it standing practice Every check above is cheap enough to automate: an assertion that evaluation is deterministic given a fixed checkpoint, an assertion on batch label diversity, a stored decoded sample from the first batch of every run, saved next to the metrics. Teams that write these down stop rediscovering the same three bugs, and the checklist is what turns a two-day investigation into a ten-minute one.
- You score the same checkpoint twice and get two different numbers. What does that tell you?That the evaluation path is stochastic, so every comparison you have drawn from it is unreliable. The usual causes are regularization still sampling — dropout left in its training behaviour — or random crops and flips left in the evaluation transform. Switch the model to its evaluation behaviour, make the evaluation transforms deterministic, and expect a few points to come back immediately.
- Why does leaving dropout active at evaluation cost accuracy rather than just adding noise?Because a single sampled mask is one thinned subnetwork, not the ensemble the deterministic pass approximates. During training, units are dropped at random and the survivors are rescaled so expected activations are preserved; evaluating with all units and no rescaling approximates the average over those subnetworks. Sampling one instead is both noisier and systematically worse, typically by a few points.
- A left-versus-right lane-change classifier plateaus near chance. Which augmentation would you suspect first?A horizontal flip. It mirrors the scene so a left-change example now shows a right change, while the label is kept unchanged, which puts directly contradictory pairs into the training set and floors the achievable loss. Either remove the flip for this task or flip the label together with the image.
- What is the fastest signal that no model change will help?Run two variants that should clearly differ — a much larger model, a much smaller learning rate — and compare. If they converge to the same disappointing number, the ceiling is upstream in the data or the measurement, not in the model. Suspiciously identical results across configurations usually mean the change never reached the computation at all.
saying these in an interview costs you the question
- Tunes the learning rate before ever inspecting a batch
- Never scores the same checkpoint twice for determinism
- Assumes a train/validation gap always means overfitting
- Blames model capacity without decoding a post-transform sample
- Treats every augmentation as automatically label-preserving