skip to content

What does inverted dropout divide surviving activations by during training?

level: middleimportance: must knowfreq 62%

answer

  1. the two modes must agree on average
  2. a constant has to live somewhere
  3. put it on the training side
  4. keep probability, not the drop rate

basics

~20 s

It divides every surviving activation by the keep probability, one minus the drop rate. That makes the layer's expected output equal its unmasked value, so evaluation runs the plain forward pass with no mask and no rescaling.

solid answer

~50 s

Inverted dropout multiplies each activation by a Bernoulli mask that is 1 with probability `keep = 1 - p`, then divides the survivors by `keep`. Per unit the expectation is `keep * (a / keep) + (1 - keep) * 0 = a`, so the layer's expected output during training matches what the unmasked layer would produce. On a 512-unit layer at drop rate 0.5, about 256 units survive and each is doubled, so the expected total is preserved. The point of putting the constant on the training side is that the evaluation path is then identical to a network with no dropout at all — no mask, no scale factor. The older alternative scaled weights by `keep` at test time instead, which meant the deployed graph depended on the rate you happened to train with.

code

python · 19 lines
python
import random
random.seed(0)

p = 0.5                 # drop rate
keep = 1.0 - p          # keep probability
h = [1.0] * 512         # one hidden layer's activations

# training forward pass: Bernoulli mask, then divide survivors by keep
mask = [1.0 if random.random() < keep else 0.0 for _ in h]
train_out = [a * m / keep for a, m in zip(h, mask)]

# evaluation forward pass: no mask, no rescaling
eval_out = list(h)

print(sum(train_out) / len(h))   # ~1.0, matches evaluation
print(sum(eval_out) / len(h))    # 1.0

# what it would be if you forgot the divide
print(sum(a * m for a, m in zip(h, mask)) / len(h))   # ~0.5

go deeper

for a junior

Remember the direction: survivors are divided by the keep probability during training, and evaluation is an ordinary forward pass. Getting the divisor the wrong way round is the mistake to rehearse away.

for a middle

Derive it on the spot. Expectation over the Bernoulli mask gives keep times a-over-keep, equals a, so the two modes agree on average and the evaluation graph needs no correction. Be able to work the 512-unit case at rate 0.5.

for a senior

Volunteer the limitation before you are pushed: only the mean is preserved, while the second moment is inflated by one over the keep probability. Connect that to why placement next to scale-sensitive layers matters.

for a principal

Frame it as an interface decision. Putting the constant on the training side keeps the served model free of regularisation residue, which is what lets rates be tuned or scheduled per layer without any change to the serving path.

## The arithmetic Let `a` be one unit's activation, `p` the drop rate, and `keep = 1 - p` the keep probability. Inverted dropout computes, at training time, ``` m ~ Bernoulli(keep) out = a * m / keep ``` and at evaluation time simply ``` out = a ``` Take the expectation over the mask: `E[out] = keep * (a / keep) + (1 - keep) * 0 = a`. The training-time layer therefore agrees *in expectation* with the evaluation-time layer, which is why no adjustment is needed when you switch modes. Without the division the expectation would be `keep * a` — at a drop rate of 0.5 every downstream layer would suddenly see inputs twice as large at evaluation as the ones it was fitted against. ## The worked case Take a 512-unit hidden layer at drop rate 0.5, and suppose for clarity every activation equals 1. Each unit survives with probability 0.5 and, when it survives, is divided by 0.5, i.e. doubled to 2. The expected number of survivors is 256, each contributing 2, so the expected sum is 512 — the same as the unmasked layer's sum of 512. The match is only in expectation. The realised sum fluctuates: each masked-and-scaled unit has variance `(1 - keep) / keep`, which is 1 at `keep = 0.5`, so the sum over 512 units has variance 512 and a standard deviation of about 23, roughly 4 to 5 percent of the mean. Wide layers average this noise down; narrow layers do not, which is one reason high rates on small layers destabilise training. ## Why the expectation is matched but the variance is not The second moment tells a different story. `E[out^2] = keep * (a / keep)^2 = a^2 / keep`, so the per-unit second moment is inflated by a factor `1 / keep` relative to the unmasked activation, and the variance contributed by the mask is `a^2 * (1 - keep) / keep`. At evaluation that extra variance vanishes entirely, because there is no mask left. This is the single most important consequence of inverted scaling, and it is what makes dropout interact badly with any downstream layer whose behaviour depends on the *scale* of its inputs rather than only their mean. Matching the first moment is enough for a plain linear layer; it is not enough for a layer that estimates and stores the variance of what it sees. ## Inverted versus the original scheme The original formulation left training untouched — mask, no rescale — and multiplied the weights by `keep` at test time so that each unit's expected input matched training. Both schemes make the two modes agree; they differ in where the constant lives. Inverted scaling is preferred because it keeps the inference-time computation exactly equal to that of an ordinary network. Three practical consequences follow. You can use different rates in different layers without carrying a table of per-layer scale factors into deployment. You can anneal or schedule the rate during training and the evaluation path never changes. And a model exported for serving has no dropout-shaped residue in it at all, so nothing about serving depends on how the model was regularised. ## Common mistakes Dividing by the *drop* rate instead of the keep probability inverts the correction and blows up the activations at high rates. Multiplying by `keep` during training rather than dividing doubles the shrinkage instead of undoing it. Applying the scale to dropped units as well is a no-op, since they are already zero. And scaling the *loss* rather than the activations does not restore anything, because the mismatch lives in the forward pass of every downstream layer, not in the objective's magnitude. ## What to say in an interview State the rule (`divide survivors by 1 - p`), give the one-line expectation calculation that justifies it, name the payoff (an untouched evaluation path), and mention that only the mean is preserved — the variance is inflated by `1 / keep`. That last sentence is what separates a memorised rule from an understood one.

  • The original formulation scaled at test time instead. What did it multiply by, and why was that abandoned?
    It multiplied the weights by the keep probability at test time, leaving training unscaled. Both schemes make the modes agree, but that one bakes the training rate into the deployed model: per-layer rates need a table of scale factors, a scheduled rate changes the inference path, and an exported model carries a correction that exists only because of how it was regularised. Inverted scaling moves the constant into training and leaves inference ordinary.
  • Does inverted scaling make the training-time and evaluation-time activations identical?
    No — it matches only the first moment. Each unit's second moment is inflated by one over the keep probability, so the training-time activations carry mask variance that disappears at evaluation. The means agree, the distributions do not, and any downstream computation that depends on the scale or spread of its inputs rather than their average will notice.
  • What happens if you divide by the drop rate rather than the keep probability?
    You apply the wrong correction, and at low drop rates it is catastrophic. At a drop rate of 0.1 the correct divisor is 0.9, giving survivors a modest boost of about 1.11; dividing by 0.1 multiplies them by 10 instead. The expected layer output becomes nine times too large, activations saturate or overflow, and the loss usually diverges within a few steps.

saying these in an interview costs you the question

  • Divides by the drop rate instead of the keep probability
  • Says the mask is still sampled during evaluation
  • Thinks inverted scaling makes training and evaluation distributions identical
  • Cannot state the one-line expectation that justifies the divisor
  • Multiplies by the keep probability during training rather than dividing

context