How do you decide whether 60% dead units in one layer of a network is a bug or an acceptable cost?
answer
- slope first, level second
- a lost fraction is lost width
- compare against a deliberately narrower run
- change one knob, re-measure the same probe batch
- recovery after the change names the cause
basics
~20 sJudge it on trend and consequence, not the level. A fraction still growing means an active pathology; a stable one is a capacity tax to weigh against the validation metric. Attribute cause by changing one knob and re-measuring.
solid answer
~50 sThree questions decide it. **Is it growing?** A dead fraction still climbing at step 5000 means something in the run is actively killing units and will keep doing so; a fraction flat since early training is a fixed cost. **Does it cost anything?** Sixty percent dead is a layer running at 40% of its configured width — compare the validation metric against a deliberately narrower run, and if the narrow run matches, you were over-provisioned rather than broken. **What caused it?** Attribute by intervening: cut the step size by ten and re-measure the same probe batch. If the dead fraction recovers from 60% to a few percent, the optimizer step size was the cause and the layer's design was not, which is exactly the distinction a dashboard alone cannot make. Only then choose the knob — step size and warmup, bias initialization, input and feature scaling, or a wider layer.
go deeper
You are not expected to make this call, but know that a dead unit costs the layer real width and that the number should be watched over time rather than read once. Ask what the fraction was at the start of training.
Be ready to explain what 60% dead means concretely — a layer running at 40% of its configured width — and to describe a controlled experiment that changes one knob and re-measures the same probe batch.
Show the attribution discipline: single-variable interventions, a step-0 baseline reading, and a check of whether the lost capacity actually shows up in the validation metric before anything is retuned.
Own the policy. Decide when a unit-health number blocks a run, insist the team compares against a deliberately narrower baseline before spending time reviving units, and standardise the measurement so profiles are comparable across runs.
## Why this is a judgment call rather than a threshold A dead-unit dashboard produces a number, and numbers invite policies like "investigate above 10%". That policy is wrong in both directions: it will send people chasing a harmless 15% in an over-wide layer, and it will stay quiet while a 5% fraction doubles every thousand steps. What matters is trend, consequence and cause, in that order. ## Trend: is the process still running? A fraction that rose to 60% in the first few hundred steps and has not moved since is the residue of a transient — something happened early, capacity was lost, and the run has been stable ever since. A fraction still climbing is qualitatively different: the mechanism is still operating, so the layer will keep narrowing for the rest of training and any decision you make now is about a moving target. Plot the fraction against step count for every layer and read the slope, not the level. The same rule applies to the saturated share: a gating layer whose pinned fraction creeps up over training is losing its ability to discriminate, however good the loss looks today. ## Consequence: what did the capacity cost you? Sixty percent dead in a layer means you are training something close to a layer 60% narrower, with the surviving units in whatever positions the deaths left them. That is not automatically bad — large layers are routinely over-provisioned, and the effective width may still exceed what the task needs. The honest test is empirical. Run the same configuration with that layer configured narrower on purpose — at roughly the surviving width — and compare validation metrics. If the deliberately narrow run matches, the deaths cost you compute and nothing else, and the right response may be to shrink the layer rather than resurrect the units. If the narrow run is clearly worse, the deaths are costing accuracy and are worth fixing. This reframes the dashboard number as a question about the model you actually wanted, which is the only framing that can justify spending team time on it. ## Cause: intervene, then re-measure Attribution is where most of the value sits, and it comes from experiment rather than inspection. Change exactly one thing, rerun the same probe batch, and compare the same statistic: - **Cut the step size by ten.** If layer 7's dead fraction falls from 60% to around 4%, the size of the optimizer's steps was driving the deaths. That single comparison separates cause from symptom: it rules out the layer's width, the activation family and the data as the primary driver, because none of those changed. Without the intervention you would have had a correlation and a guess. - **Change the initialization of that layer's biases.** If a large negative bias was holding pre-activations under zero from step one, the fraction will be wrong from the very first reading, before any optimizer step has been taken — which is itself diagnostic. Always log the statistic at step 0. - **Rescale or renormalize the inputs to the block.** If unnormalized features are pushing pre-activations far from the useful range, standardizing them moves the histogram back on its own. - **Add a warmup.** If deaths are concentrated in the first few hundred steps and never recover, the run is being damaged before the optimizer has settled, and a gentler opening may cost nothing elsewhere. The same intervene-and-remeasure discipline handles saturation. A layer whose sigmoid outputs sit above 0.99 for every input from the first reading is almost certainly a bias or input-scale problem rather than something training produced; if instead the pinned fraction grows over training, the network has learned to be confident and you should ask whether that confidence is warranted before you fight it. ## Deciding, and making the decision reusable By the end you should be able to say something like: the fraction stabilised at 60% by step 300, an intentionally narrower run loses two points of validation accuracy, and a tenfold smaller step size keeps the fraction under 5% with no loss in final metric — therefore this is a step-size problem and we are changing the schedule. That is a defensible engineering statement; "the dashboard was red so I switched something" is not. The organisational half of this is keeping the number comparable. Fix the probe batch and the measurement mode across the team, log at step 0 as well as during training, and record the statistic alongside the run configuration so profiles from different runs can be laid over each other. A telemetry number nobody can compare across runs generates debate rather than decisions, and the cost of collecting it is then pure overhead.
- What does it tell you when the dead fraction is already high at step zero?That no optimizer step caused it. Before any update, the only inputs to the statistic are the initialization and the data scale, so a high reading at step 0 points at bias or weight initialization, or at unnormalized inputs pushing pre-activations far from the useful range. It rules out step size entirely, which is why the step-0 reading is worth logging separately.
- When would you shrink the layer rather than try to revive the dead units?When a deliberately narrower run matches the validation metric of the wide one. That evidence says the extra width was never earning its keep, so the deaths cost only compute and memory. Shrinking then buys you a smaller, faster model and removes the telemetry alarm honestly, instead of tuning to make a number look better with no metric change behind it.
- How do you keep dead-fraction numbers comparable across runs and people?Fix the probe batch, the measurement mode and the layer indexing as a shared convention, log at step 0 as well as periodically during training, and store the statistic next to the run configuration. Without that, two engineers report different numbers for the same network and the discussion turns into an argument about measurement rather than about the model.
saying these in an interview costs you the question
- Sets a fixed percentage threshold and treats it as universal
- Acts on the level while ignoring whether it is growing
- Changes several knobs at once, then cannot attribute the fix
- Never checks whether the lost capacity affects the metric
- Assumes any dead unit is wasted and must be revived
- Compares fractions measured on different random batches