skip to content

Why would you raise the epsilon in Adam's denominator from a tiny default to a much larger value?

level: seniorimportance: should knowfreq 41%

answer

  1. look at the denominator, not the numerator
  2. what if the second moment goes tiny
  3. noise divided by its own magnitude
  4. it caps the maximum possible step
  5. a large epsilon makes steps gradient-proportional

basics

~20 s

Epsilon puts a floor under Adam's denominator, capping how large a step a tiny second moment can produce. Raise it when gradients have become genuinely small and the optimizer is amplifying noise into full-size, jittery updates.

solid answer

~50 s

Epsilon is usually described as preventing division by zero, but its real job is to bound the step. The update is `lr * m_hat / (sqrt(v_hat) + eps)`, so when a parameter's root second moment falls far below epsilon the step becomes roughly `lr * m_hat / eps` — proportional to the gradient again, rather than normalized to about one. That is the behaviour you want when gradients are genuinely tiny, for example late in training a multivariate time-series forecaster that has already fit most of the signal: the residual gradients are mostly noise, and dividing noise by its own root-mean-square manufactures a full-size step out of nothing. Raising epsilon from around 1e-8 to around 1e-4 damps that jitter without touching parameters whose gradients are still large compared with epsilon. The cost is that you give up adaptivity exactly where the gradient really is small but real.

go deeper

for a junior

Know that a small constant is added to the denominator of the adaptive step so that a near-zero second moment cannot produce a division by zero or an enormous update.

for a middle

Explain the three regimes by comparing the root second moment with epsilon, and show that once epsilon dominates the denominator the update becomes proportional to the smoothed gradient again.

for a senior

Diagnose the symptom before reaching for the knob: identify that jitter is coming from noise-scale gradients being normalized up, and be able to argue why raising epsilon and not lowering the learning rate is the right response.

for a principal

Own the scale dependence. Epsilon is the one term that breaks Adam's invariance to loss scaling, so decide as a matter of policy how a team records and revalidates it when loss reductions, batch sizes or output heads change.

## What epsilon actually does Adam's per-parameter step is ``` step = lr * m_hat / (sqrt(v_hat) + eps) ``` The stated reason for `eps` is to avoid dividing by zero when a parameter has seen no gradient. That reason is real but minor. The operational reason is that `eps` sets a **floor on the denominator**, and therefore a **ceiling on the step**: no matter how small the second moment becomes, the step cannot exceed `lr * |m_hat| / eps`. With `lr = 1e-3` and `eps = 1e-8` that ceiling is `1e5 * |m_hat|` — effectively no ceiling at all. With `eps = 1e-4` it is `10 * |m_hat|` — a real constraint. ## The regime where it matters Think of a parameter's root second moment `sqrt(v_hat)` as its recent typical gradient magnitude, and compare it with `eps`. - **`sqrt(v_hat) >> eps`.** The epsilon is irrelevant; the update is the normalized ratio, close to plus or minus one for a consistent gradient. This is the regime most parameters live in for most of training. - **`sqrt(v_hat) ~ eps`.** The epsilon starts to shrink the step, damping this parameter relative to the others. - **`sqrt(v_hat) << eps`.** The update degenerates to `lr * m_hat / eps`, which is simply a momentum step with an effective learning rate of `lr / eps`. Adaptivity is switched off for that parameter and it moves in proportion to its smoothed gradient again. So epsilon is not a numerical nicety; it is the knob that decides *at what gradient scale Adam stops normalizing*. ## The failure it fixes Consider a multivariate time-series forecaster in the late phase of training. The model has captured the systematic structure, and what is left in the gradient is mostly minibatch noise around zero — genuinely tiny in magnitude for many parameters. Adam's normalization does not know the difference between a small consistent signal and small noise: it divides by the same root-mean-square either way. Weights start jittering at close to a full learning rate driven by nothing, validation loss becomes ragged, and each evaluation lands somewhere slightly different. Raising `eps` from around `1e-8` to around `1e-4` puts a floor under the denominator that those noise-scale gradients cannot get under. Parameters whose gradients are still meaningfully large are essentially unaffected — for them `sqrt(v_hat)` still dominates the sum — while the noise-scale parameters quiet down. The run settles. ## Why this is not the same as lowering the learning rate This is the discriminating question, and it is worth having the answer ready. Lowering `lr` scales **every** parameter's step by the same factor, including the ones making real progress. Raising `eps` shrinks **only** the steps of parameters whose second moment has fallen near or below the new epsilon. It is a selective damping keyed on gradient scale, not a global slowdown. If the diagnosis is "the whole model is overshooting", lower the rate; if the diagnosis is "parameters whose gradients have gone to noise are still moving at full speed", raise epsilon. ## The costs and the caveats **Epsilon is not scale-free.** It is compared against a gradient magnitude, so a good value depends on the scale of your gradients: the loss reduction used, the output head, the batch size, the parametrisation. A value carried over from a different model is a guess, not a default. If you rescale the loss, the sensible epsilon rescales with it — this is exactly the point where Adam's otherwise clean invariance to loss scaling breaks. **You lose real adaptivity.** The whole reason to use Adam is that a parameter with a small but consistent gradient still travels. Raise epsilon far enough and you have re-created the problem you were avoiding: layers with genuinely small gradients crawl again. This is why the change is a targeted response to an observed pathology, not a default worth raising "just in case". **It is a single global knob for a per-parameter problem.** Every parameter shares the same epsilon, but only some of them are in the noise regime. Anything you gain by damping the noisy ones you also apply to any legitimately small-gradient parameter. **Placement varies between formulations.** Some write the denominator as `sqrt(v_hat) + eps` and others as `sqrt(v_hat + eps)`; the two put epsilon on different scales, so a value tuned under one formulation does not carry over to the other. When you inherit an epsilon from someone else's recipe, check which form it was tuned against. The short version for an interview: epsilon caps the largest step a vanishing second moment can produce, and you raise it when tiny gradients are being amplified into meaningless full-size updates — accepting, in exchange, that genuinely small gradients will now move more slowly too.

  • How does raising epsilon differ from simply lowering the learning rate?
    Lowering the learning rate scales every parameter's step by the same factor, including the ones still making progress. Raising epsilon only shrinks the steps of parameters whose root second moment has fallen near or below it, leaving large-gradient parameters essentially untouched. One is a global slowdown, the other is selective damping keyed on gradient scale.
  • Is a good epsilon value transferable between models?
    Not reliably. Epsilon is compared against a gradient magnitude, so the right value depends on the loss scale, the output head, the reduction used over the batch, and the parametrisation. Rescale the loss and the sensible epsilon rescales with it. Treat an inherited epsilon as a starting guess and confirm it against the actual size of your root second moments.
  • What is the largest step Adam can take for one parameter?
    Roughly `lr * |m_hat| / eps`, reached when the root second moment is negligible against epsilon. With a tiny epsilon that ceiling is enormous, which is exactly why a parameter whose gradients have collapsed to noise can still move at full speed. Raising epsilon lowers the ceiling in a way no learning-rate change can imitate.

saying these in an interview costs you the question

  • Says epsilon exists only to prevent division by zero
  • Thinks a larger epsilon produces larger updates
  • Claims raising epsilon is equivalent to lowering the learning rate
  • Treats epsilon as dimensionless and portable across losses
  • Raises epsilon as a default rather than in response to a diagnosis

context