skip to content

Dead and Saturated Units

Spotting units that output zero for every input, sigmoid and tanh layers pinned at their asymptotes, and layers whose gradient norms have collapsed, from per-layer statistics rather than guesswork.

on this pageshow

questions

3

Which per-layer statistics reveal dead or saturated units in a neural network's training run?

level: middleimportance: must knowfreq 58%

answer

  1. one scalar per layer, logged over time
  2. reduce over the batch, then over units
  3. zero for every example, not for one
  4. bounded outputs past 0.99 in magnitude
  5. half the values zero per example is normal

basics

~20 s

Per layer, log the fraction of units that output zero for every example in a fixed probe batch, plus the fraction of tanh or sigmoid outputs past 0.99 in magnitude. Both are per-unit across the batch, not per example.

solid answer

~50 s

Log two per-layer numbers on a fixed probe batch. First, the **dead fraction**: the share of units whose output is exactly zero for *every* example in the batch. A unit that is zero on some examples is just sparse, not dead — half the units of a healthy rectified layer are typically zero on any single example. Second, the **saturated fraction** for tanh or sigmoid layers: the share of outputs whose magnitude sits past about 0.99, or within a hair of 0 or 1 for a sigmoid. Take both with the network in evaluation behaviour so that dropout masking is not counted as death, keep the probe batch fixed across steps so successive readings are comparable, and plot them per layer over time — the trend matters more than any single reading. Back them with a full per-layer activation histogram, since a mean hides a bimodal pile-up at the asymptotes.

go deeper

for a junior

Be ready to say what you would actually log: one number per layer, computed on the same probe batch each time, and an activation histogram next to it. Knowing that a unit stuck at zero stops learning is enough at this level.

for a middle

Explain the reduction order precisely — maximum over the batch per unit, then share of units — and why the per-value zero share is sparsity rather than death. Name the saturation counterpart and the 0.99 style threshold.

for a senior

Show that you make the measurement trustworthy in a real run: evaluation behaviour so dropout is not counted, a fixed probe batch, per-layer breakdown, cheap cadence, and reading the trend rather than a single snapshot.

for a principal

Own the question of what this telemetry is for. Decide which statistics every run in the team logs by default, what threshold or trend opens an investigation, and how to keep the numbers comparable across runs so they can be argued about.

## The two failures these statistics are looking for A unit is **dead** when its output is zero for every input in your data distribution: it contributes nothing to the forward pass, and nothing flows back through it, so it never changes again. A unit is **saturated** when a bounded activation — tanh, sigmoid — is pinned against an asymptote (±1 for tanh, 0 or 1 for a sigmoid) for essentially every input. The local slope there is nearly flat, so the unit is nearly frozen too and its output carries almost no information about the input. Both are invisible in the loss curve until they have cost you enough capacity to matter, which is why they are instrumented rather than inferred. ## Measure per unit, over a batch — not per activation This is where most people get it wrong. If you take a rectified layer's activation tensor and compute "what share of the numbers are zero", you will get something near 50% on a perfectly healthy layer at initialization, because roughly half the pre-activations are negative for any given example. That number is **sparsity**, and it is normal and even useful. The dead fraction is a different reduction. For each unit, take the maximum of its output over the whole probe batch; the unit is dead only if that maximum is zero (or, more robustly, below a small epsilon). Then the dead fraction is the share of units in the layer for which that holds. Concretely: pushing a 4096-example probe batch through a network and finding that 60% of layer 7's units emit exactly zero on all 4096 examples is a genuine alarm; finding that 60% of the individual activation values are zero is not. The saturation statistic mirrors it. For a tanh layer, log the share of outputs with |h| > 0.99; for a sigmoid gate, the share outside [0.01, 0.99]. Here you often want the per-value share as well as a per-unit version, because a gate that is pinned for every input is a different pathology from a gate that is pinned for the easy 90% of inputs. A recurrent stack whose hidden-state histogram piles against ±1 at every time step is telling you the state has stopped encoding gradations at all. ## Make the reading trustworthy A few practicalities separate a useful dashboard from a misleading one. - **Evaluation behaviour.** Measure with dropout inactive and normalization layers in their inference mode, or read the activations *before* the dropout mask. Otherwise you are measuring the mask. - **A fixed probe batch.** Reuse the same held-out examples every time you log, so that a change in the number reflects a change in the network rather than a change in the data. - **Per layer, and layer-by-layer.** One number for the whole network averages away the fact that the pathology is concentrated in one block. The shape of the profile across depth is the diagnostic. - **Cheap cadence.** Every few hundred steps is plenty; this is a forward pass on one batch, not a second training loop. - **Histograms, not just scalars.** The scalar tells you something is wrong; the histogram tells you what shape of wrong — a spike at zero, a two-humped pile at ±1, or a distribution drifting steadily wider as training goes on. ## Reading the numbers There is no universal threshold, but there are useful reference points. A few percent of permanently-zero units in a wide layer is ordinary attrition and rarely worth chasing. Tens of percent means you have effectively trained a much narrower layer than the one you configured. A dead fraction that is *growing* over training is more alarming than a larger one that has been stable since step 100 — growth means an active process is still killing units, whereas a stable fraction is a fixed capacity tax you can measure against the validation metric. For saturation, a small pinned share is expected in any confident model — a gate that has learned to be open should sit near 1. What you are looking for is a layer where nearly all units and nearly all inputs land at the asymptote, especially early in training, since that means the layer is not yet distinguishing inputs at all. A classic instance: a gating unit whose bias was initialized several units positive outputs above 0.99 for every input, which shows up as a single spike in the output histogram and nowhere else. ## Pair it with gradient statistics Activation statistics tell you which units are inert; per-layer gradient statistics tell you whether the learning signal reaching a block is proportionate to its size. Logged together, they usually localise the problem to one block within a couple of readings, which is the point of the instrumentation: turn "training is worse than it should be" into "layer 7 has lost 60% of its capacity since step 400".

  • Why is the share of zero activation values a poor proxy for the dead fraction?
    Because a healthy rectified layer is roughly half zeros on any single example — that is sparsity, not death. The dead fraction reduces over the batch first: a unit counts only if its maximum output across every probe example is zero. Confusing the two makes a normal layer look catastrophic and hides a layer where a few units are genuinely stuck.
  • How does dropout change what an activation histogram shows?
    Dropout zeroes a random subset of activations during training, so a histogram captured in training mode has a spike at zero that has nothing to do with unit health. Measure with dropout inactive, or read the activations before the mask is applied. The same care applies to normalization layers, whose training-time and inference-time behaviour differ.
  • What makes a growing dead fraction more urgent than a large but stable one?
    Growth means something in the current run is still killing units, so the damage compounds and the layer keeps narrowing. A fraction that has been flat since early training is a fixed capacity tax: you can evaluate it directly against the validation metric and decide whether it costs anything. Trend, not level, is what should trigger intervention.

A dead unit is a light switch wired shut: it is not merely off today, it fails to turn on for anything you show it. Sparsity is a switch that is simply off for this particular room.

saying these in an interview costs you the question

  • Calls a unit dead because it output zero on one example
  • Expects a healthy rectified layer to contain almost no zeros
  • Reads only the mean activation, hiding a bimodal histogram
  • Measures activations with dropout still active
  • Reports one histogram for the whole network, not per layer
  • Uses a different random batch for each logged reading

context

open as a page

Per-layer gradient norms logged every 50 steps put layer 1 four orders below layer 12 — what do you check?

level: seniorimportance: should knowfreq 46%

basics

~20 s

First make the comparison fair: raw norms grow with parameter count, so switch to a per-parameter root-mean-square or the update-to-weight ratio. Then check whether layer 1's weights actually move, and read the activation statistics in between.

open as a page

How do you decide whether 60% dead units in one layer of a network is a bug or an acceptable cost?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Judge it on trend and consequence, not the level. A fraction still growing means an active pathology; a stable one is a capacity tax to weigh against the validation metric. Attribute cause by changing one knob and re-measuring.

open as a page