skip to content

Training Dynamics and Debugging

You will learn the debugging playbook for training runs: pick initialization correctly, read a loss curve, overfit a single batch first, and diagnose NaNs, dead ReLUs, and silent data bugs. 'Your loss is not decreasing — walk me through it' is a staple senior DL interview scenario.

on this pageshow

explore

questions

page 2 of 2

Two training runs with identical seeds and identical data order diverge by step 4000 - why?

level: seniorimportance: nice to knowfreq 38%

basics

~10 s

Floating-point addition is not associative, and parallel kernels sum partial results in an order that varies between launches. Last-bit gradient differences feed back through thousands of updates until the two runs are visibly apart.

open as a page

Validation cross-entropy is rising while validation accuracy still improves — what explains it?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Cross-entropy scores confidence, accuracy scores only the decision. As training continues the model becomes more confident everywhere, so the examples it still gets wrong contribute a large and growing penalty while newly-correct ones add almost nothing.

open as a page

In a 200-block residual network, why zero-initialize each block's final scale parameter?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

It makes every residual branch output exactly zero at step zero, so the stack starts as the identity map and the signal keeps the scale it entered with instead of growing block by block. The branches then switch on gradually.

open as a page

How do you decide whether 60% dead units in one layer of a network is a bug or an acceptable cost?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Judge it on trend and consequence, not the level. A fraction still growing means an active pathology; a stable one is a capacity tax to weigh against the validation metric. Attribute cause by changing one knob and re-measuring.

open as a page

Your team cites a double descent plot to argue for dropping regularization and just scaling up. What do you say?

level: principalimportance: nice to knowfreq 20%

basics

~10 s

A double descent plot does not show that scale replaces regularization: those sweeps are run under-regularized on noisy labels, and tuning the penalty at each capacity can flatten the peak away.

open as a page

Rescaling a ReLU network leaves its function identical but inflates measured sharpness — what does that break?

level: principalimportance: nice to knowfreq 18%

basics

~10 s

It breaks sharpness as a standalone explanation of generalization. Scaling one layer up and the next down computes the same function with the same test error, yet raw curvature measures change freely.

open as a page

When is a learning-rate range test the wrong tool for choosing the rate you ship?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

A range test measures one thing: the largest rate that keeps training loss falling over a couple of hundred steps. It is silent on generalization, on stability thousands of steps later, and on recipes whose good rate is already known.

open as a page

How would you harden a multi-day training loop against an occasional non-finite loss?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Check that the loss and the gradient norm are finite before the update is applied, not after, because the optimizer's running buffers keep a NaN forever. Skip and log the step, persist the batch, and alarm on the skip rate.

open as a page

showing 31–38 of 38