Training Dynamics and Debugging
You will learn the debugging playbook for training runs: pick initialization correctly, read a loss curve, overfit a single batch first, and diagnose NaNs, dead ReLUs, and silent data bugs. 'Your loss is not decreasing — walk me through it' is a staple senior DL interview scenario.
on this pageshowhide
explore
- Starting the Run Right13 questions
- Weight Initialization Scales3 questions
- Learning-Rate Range Test3 questions
- Overfitting One Batch3 questions
- Seeding and Run Variance4 questions
- Loss Curve Reading11 questions
- Train and Validation Traces4 questions
- Double Descent3 questions
- Flat and Sharp Minima4 questions
- Diagnosing a Broken Run14 questions
- NaN and Infinite Losses3 questions
- Dead and Saturated Units3 questions
- Data Versus Model Bugs3 questions
- Gradient Attribution Maps5 questions
questions
page 2 of 2Two training runs with identical seeds and identical data order diverge by step 4000 - why?
basics
~10 sFloating-point addition is not associative, and parallel kernels sum partial results in an order that varies between launches. Last-bit gradient differences feed back through thousands of updates until the two runs are visibly apart.
Validation cross-entropy is rising while validation accuracy still improves — what explains it?
basics
~20 sCross-entropy scores confidence, accuracy scores only the decision. As training continues the model becomes more confident everywhere, so the examples it still gets wrong contribute a large and growing penalty while newly-correct ones add almost nothing.
In a 200-block residual network, why zero-initialize each block's final scale parameter?
basics
~20 sIt makes every residual branch output exactly zero at step zero, so the stack starts as the identity map and the signal keeps the scale it entered with instead of growing block by block. The branches then switch on gradually.
How do you decide whether 60% dead units in one layer of a network is a bug or an acceptable cost?
basics
~20 sJudge it on trend and consequence, not the level. A fraction still growing means an active pathology; a stable one is a capacity tax to weigh against the validation metric. Attribute cause by changing one knob and re-measuring.
Your team cites a double descent plot to argue for dropping regularization and just scaling up. What do you say?
basics
~10 sA double descent plot does not show that scale replaces regularization: those sweeps are run under-regularized on noisy labels, and tuning the penalty at each capacity can flatten the peak away.
Rescaling a ReLU network leaves its function identical but inflates measured sharpness — what does that break?
basics
~10 sIt breaks sharpness as a standalone explanation of generalization. Scaling one layer up and the next down computes the same function with the same test error, yet raw curvature measures change freely.
When is a learning-rate range test the wrong tool for choosing the rate you ship?
basics
~20 sA range test measures one thing: the largest rate that keeps training loss falling over a couple of hundred steps. It is silent on generalization, on stability thousands of steps later, and on recipes whose good rate is already known.
How would you harden a multi-day training loop against an occasional non-finite loss?
basics
~20 sCheck that the loss and the gradient norm are finite before the update is applied, not after, because the optimizer's running buffers keep a NaN forever. Skip and log the step, persist the batch, and alarm on the skip rate.
showing 31–38 of 38