skip to content

Why is SiLU's dip below zero acceptable when most activations are monotone?

level: middleimportance: should knowfreq 41%

answer

  1. two factors fight on the negative side
  2. gate decays faster than the input grows
  3. there is a minimum, then a climb back
  4. about -0.28 near x = -1.28
  5. bounded below, derivative still continuous

basics

~20 s

Monotonicity is not required for training or for function approximation. SiLU dips to about -0.278 near x = -1.278 and then rises; its derivative is continuous and its output is bounded below, so gradient descent handles it fine.

solid answer

~40 s

SiLU is `x * sigmoid(x)`. Moving left from the origin the input shrinks the product while the gate is still fairly open, so the output falls to a minimum of about -0.278 near `x = -1.278`; further left the gate closes faster than the input grows, and the output climbs back toward 0. That makes the function non-monotone, with a negative derivative for inputs below about -1.278. Nothing in backpropagation requires monotone units - it requires a well-defined, bounded derivative, which SiLU has everywhere. The dip also gives two useful properties: the unit is bounded below, so a wildly negative pre-activation cannot produce a wildly negative output, and a mildly negative pre-activation still carries signal forward instead of being erased. GELU is non-monotone in the same way, with a shallower minimum near -0.17.

code

python · 15 lines
python
import math

def silu(x):
    return x / (1.0 + math.exp(-x))

def gelu(x):
    return x * 0.5 * (1.0 + math.erf(x / math.sqrt(2.0)))

grid = [-4.0 + 0.001 * i for i in range(8001)]
xs = min(grid, key=silu)
xg = min(grid, key=gelu)
print("SiLU min %.4f at x = %.3f" % (silu(xs), xs))
print("GELU min %.4f at x = %.3f" % (gelu(xg), xg))
for x in (-2.0, -0.5, 0.5, 2.0):
    print("x=%5.1f  silu=%7.4f  gelu=%7.4f" % (x, silu(x), gelu(x)))

go deeper

for a junior

Know that SiLU and GELU are not monotone: their output falls a little below zero for mildly negative inputs and then comes back toward zero. Being able to sketch that shape is enough at this level.

for a middle

Explain the mechanics: a linearly growing negative input against an exponentially closing gate produces a minimum, roughly -0.28 at -1.28 for SiLU. State the derivative's sign on either side and why continuity is what optimisation actually needs.

for a senior

Demonstrate calibrated judgment. Say plainly that theory permits the dip and that the measured benefit on a given architecture is usually small, so it belongs in an ablation rather than in an argument.

for a principal

Frame it as a research-versus-engineering call: how much of the team's time should go into activation-shape exploration when the effect size is small and architecture-dependent, compared with data, scale, or training-recipe work.

## Where the dip comes from SiLU is `f(x) = x * sigmoid(x)`, a product of two factors that pull in opposite directions on the negative side. As `x` moves left of the origin, the first factor becomes more negative while the second - the gate - shrinks toward 0. Near the origin the input's growth in magnitude wins and the product falls. Far to the left the gate's decay wins, because a logistic gate decays roughly exponentially while the input only grows linearly, so the product turns back toward 0. Between the two regimes there is a minimum. Numerically the minimum is about **-0.2785 at x = -1.278**. On either side the function is closer to 0: it is about -0.238 at -2, and about -0.142 at -3. GELU has the same shape for the same reason, with a shallower minimum of about **-0.17 near x = -0.752**, because a Gaussian CDF closes faster than a logistic curve. So neither smooth gated unit is monotone, and that is a property of the construction, not an accident of one of them. ## Is non-monotonicity a problem? It is worth being precise about what monotonicity is and is not needed for. **Not needed for expressive power.** The classical universal-approximation arguments need a nonlinearity that is not a polynomial, plus enough width. Monotonicity is not part of the requirement. Non-monotone activations approximate functions perfectly well. **Not needed for backpropagation.** Backpropagation multiplies derivatives along the chain. It needs the derivative to exist and to be numerically sane; it does not care about its sign. SiLU's derivative is `s(x) * (1 + x * (1 - s(x)))` with `s = sigmoid`, and it is continuous everywhere, equal to 0.5 at the origin, and negative for `x` below about -1.278. A negative local derivative simply means an increase in the pre-activation there decreases the output - a legitimate local response, the same thing that happens on the falling side of any bump in a learned function. **What it does cost.** The unit is not invertible: two different pre-activations can give the same output, one on each side of the minimum. That matters if you were relying on a per-unit monotone response for interpretation - for example, arguing "a larger pre-activation always means a stronger downstream signal from this unit". With SiLU that argument is false in the dip region. It also means loss surfaces can be slightly more textured than with a monotone unit; in practice this has not stopped anyone. ## What the dip buys 1. **A bounded negative side.** `f(x) >= -0.2785` for all `x`. However negative a pre-activation drifts during training, the unit cannot emit an arbitrarily large negative value; the gate clamps it. That is a mild self-limiting property that a linear-on-the-negative-side unit does not have. 2. **Preserved small negatives.** A pre-activation of -0.5 leaves as about -0.19 rather than being erased. The next layer can still distinguish "slightly negative" from "very negative", because those map to different outputs on the rising part of the curve. 3. **A non-zero gradient on the negative side.** Because the function is not flat for negative inputs, units sitting in the negative region still receive gradient and can move, rather than being permanently stuck at a constant output. ## Where the shape came from Swish - the same function as SiLU - was found by an automated search over compositions of simple primitives, and the shape the search kept returning was self-gated and non-monotone. That is worth mentioning in an interview because it reframes the design: the dip is not a hand-crafted idea someone argued for from first principles, it is a shape that repeatedly came out well and was then explained afterwards. ## How to talk about it A good answer separates three claims. *Descriptive*: SiLU dips to about -0.28 near -1.28 and GELU dips to about -0.17 near -0.75. *Theoretical*: nothing in approximation theory or gradient-based optimisation requires monotone activations, only a well-behaved derivative. *Empirical and honest*: the accuracy effect of the dip on any given model is usually small and architecture-dependent, so treat "the dip helps" as a hypothesis you would test on your own model rather than a law. Candidates who assert a large guaranteed gain are overclaiming just as badly as candidates who assert the dip breaks training.

  • Is GELU monotone, or does it dip too?
    It dips too, for the same structural reason: the input grows in magnitude while the Gaussian CDF gate closes. Its minimum is about -0.17 near `x = -0.752`, shallower and closer to the origin than SiLU's, because a Gaussian CDF closes faster than a logistic curve. Neither smooth gated unit is monotone.
  • Over which inputs is SiLU's derivative negative?
    For inputs below about -1.278, the location of the minimum. Left of it the function is falling as `x` increases toward the minimum, so the derivative there is negative; right of it, the derivative is positive. The derivative formula is `s(x) * (1 + x * (1 - s(x)))` with `s = sigmoid`, and it is continuous through the whole range.
  • Does the dip mean gradients vanish for very negative inputs?
    Not exactly. The output tends to 0 for very negative inputs and so does the derivative, so a unit driven far negative does contribute little gradient - but it approaches that limit smoothly and is never identically flat, so the unit can still drift back. That is different from a region with a derivative that is exactly zero.

saying these in an interview costs you the question

  • Claims non-monotone activations break backpropagation
  • Says monotonicity is required for universal approximation
  • Thinks SiLU is unbounded below on the negative side
  • Believes the derivative is undefined at the minimum
  • Asserts the dip guarantees a large accuracy gain

context