skip to content

Why does elementwise dropout barely regularize a convolutional feature map?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the mask removes less than you think
  2. neighbours hold the same evidence
  3. overlapping receptive fields, near-duplicate activations
  4. change the unit of dropping, not the rate
  5. drop whole feature maps instead

basics

~20 s

Neighbouring positions in a feature map are strongly correlated because their receptive fields overlap, so zeroing scattered individual activations leaves the same evidence present next door. Channel-wise dropout, which zeroes whole feature maps, removes a detector outright and actually regularizes.

solid answer

~40 s

A convolutional layer applies one filter across the whole input, so adjacent activations in a feature map come from overlapping receptive fields and carry nearly the same evidence. Zeroing a scattered 30 percent of positions removes almost no information: the next position over still fires, and the following layer's own receptive field pools over the survivors. The perturbation is real but weak, so you pay the noise and the slower convergence without buying much regularisation. Channel-wise, or spatial, dropout instead zeroes an entire feature map at once, so a whole learned detector is absent for that step and the network cannot lean on it. On a 1D convolutional ECG-arrhythmia classifier, where adjacent timesteps are near-duplicates within a beat, this is the difference between dropout doing nothing measurable and dropout mattering.

go deeper

for a junior

Know that dropout is not one-size-fits-all: convolutional layers have variants that zero whole feature maps rather than individual positions. The reason is that neighbouring positions in a map carry nearly the same information.

for a middle

Explain the correlation argument precisely: overlapping receptive fields make adjacent activations near-duplicates, so an elementwise mask removes information the next layer can recover. Then name channel-wise dropping as the coarser unit that fixes it.

for a senior

Show judgment about granularity. Say where in a stack the redundancy is worst, why raising the elementwise rate is the wrong lever, and what the coarser mask costs in gradient variance on a narrow layer.

for a principal

Own the general principle across architectures: choose the unit of dropping so what you removed cannot be reconstructed, and keep the mask consistent along whatever axis the model accumulates information over.

## The correlation problem A convolutional layer slides one filter over its input, so the value at position `i` of a feature map and the value at position `i + 1` are computed by the same weights from receptive fields that overlap almost completely. In a 1D convolution over an ECG signal with a kernel spanning, say, 15 samples and a stride of 1, two adjacent outputs share 14 of their 15 inputs. Their activations are close to duplicates. Dropout's regularising force comes from making a piece of information genuinely unavailable, so that whatever depends on it has to survive without it. When the masked value is nearly reproduced by its neighbour, that force evaporates. Zero position `i` and the evidence is still sitting at `i - 1` and `i + 1`; the next layer, whose own receptive field spans many positions, aggregates over the survivors and barely notices. You have injected noise — which slows convergence and makes the validation curve bumpier — without removing anything the network relied on. The same argument is why the effect is worst in early layers with large spatial extent and heavy overlap, and less pronounced deep in the stack where activations have been repeatedly downsampled and are less redundant. ## Changing the unit that gets dropped The fix is not a different rate but a different *unit of dropping*. Channel-wise dropout, usually called spatial dropout, samples one Bernoulli variable per channel rather than per element, and zeroes that channel's entire feature map together. Because a channel is one learned filter's response, dropping it removes a whole detector for that step: whatever the network learned to do with the output of the eleventh filter must be doable without it. That is a real, non-recoverable perturbation, and it is what makes co-adaptation between filters costly. Practical notes that follow from the coarser unit. The same inverted scaling applies — survivors are divided by the keep probability — but the mask has one entry per channel instead of one per activation, so the noise is far coarser and the gradient variance per step is higher. Rates are usually kept lower than the classic setting for wide fully connected layers, because dropping half the channels of a 64-channel layer removes a large fraction of the model's representational width in one step. And with few channels the coarse mask can be too blunt to use at all. ## A coarser unit still: stochastic depth The same idea taken one level up drops entire residual blocks. In a 110-block residual stack, stochastic depth assigns each block a survival probability and, when a block is dropped for a training step, the block's residual branch is skipped entirely and the block computes the identity — the input passes through the skip connection unchanged. The network stays valid because the skip path always exists; you have simply trained a shallower network for that step. Two details matter. Survival probability is typically decayed linearly with depth, so early blocks that everything downstream depends on are kept almost always while late blocks are dropped often; this reflects that late blocks make smaller, more specialised refinements. And at evaluation all blocks are used, with each block's residual branch weighted by its survival probability so that the expected contribution matches training. Beyond regularisation this shortens the average backward path, which helps very deep stacks train, and it cuts training time because dropped blocks are never computed. ## The recurrent case Recurrent layers show the same lesson from a different angle. If you resample an independent mask at every time step of a sequence, the noise applied to the hidden state compounds step after step and the state's ability to carry information over long spans is destroyed — the model degrades rather than regularises. Sharing one mask across all time steps of a given sequence, and drawing a fresh one for the next sequence, keeps the perturbation consistent along the sequence: the same units are absent throughout, so the recurrence can still transport what survives. On a part-of-speech tagger, per-step resampling erases exactly the long-range context the recurrence exists to provide. ## The unifying rule Drop at a granularity coarse enough that what you removed cannot be reconstructed from what remains, and consistent along whatever axis the model uses to accumulate information. Elementwise masks are the right choice for a fully connected layer, where units are genuinely separate features; they are the wrong choice wherever the architecture has deliberately built in redundancy or recurrence.

  • In stochastic depth, what does a residual block compute on a step where it is dropped?
    The identity. The residual branch is skipped entirely and the input flows through the skip connection unchanged, so the network stays valid and simply behaves as a shallower stack for that step. Survival probability usually decays linearly with depth, keeping early blocks almost always and dropping late ones often, and at evaluation every block runs with its residual branch weighted by its survival probability.
  • Why do you share one dropout mask across all time steps of a sequence in a recurrent layer?
    Because resampling per step compounds the noise along the recurrence and erases the state. With a fresh mask at every step, whichever units carry long-range information are randomly interrupted over and over, so nothing survives many steps. Holding one mask for the whole sequence, and redrawing per sequence, keeps the same units absent throughout, so the recurrence can still transport what remains.
  • Would you use channel-wise dropout on a layer with only 16 channels?
    Cautiously, and at a low rate. A per-channel mask on 16 channels is a very coarse perturbation: at rate 0.5 you would routinely delete half the layer's representational width in one step, and the step-to-step gradient variance becomes large enough to destabilise training. Narrow layers are usually better served by other regularisers, or by putting the dropout on the wider layers further along.

saying these in an interview costs you the question

  • Suggests fixing weak regularisation by raising the elementwise rate
  • Thinks a feature map's neighbouring activations are independent
  • Believes a dropped residual block outputs zeros rather than the identity
  • Applies a fresh per-step mask inside a recurrent layer
  • Confuses dropping whole channels with dropping whole training examples

context