Why is SiLU's dip below zero acceptable when most activations are monotone?
answer
- two factors fight on the negative side
- gate decays faster than the input grows
- there is a minimum, then a climb back
- about -0.28 near x = -1.28
- bounded below, derivative still continuous
basics
~20 sMonotonicity is not required for training or for function approximation. SiLU dips to about -0.278 near x = -1.278 and then rises; its derivative is continuous and its output is bounded below, so gradient descent handles it fine.
solid answer
~40 sSiLU is `x * sigmoid(x)`. Moving left from the origin the input shrinks the product while the gate is still fairly open, so the output falls to a minimum of about -0.278 near `x = -1.278`; further left the gate closes faster than the input grows, and the output climbs back toward 0. That makes the function non-monotone, with a negative derivative for inputs below about -1.278. Nothing in backpropagation requires monotone units - it requires a well-defined, bounded derivative, which SiLU has everywhere. The dip also gives two useful properties: the unit is bounded below, so a wildly negative pre-activation cannot produce a wildly negative output, and a mildly negative pre-activation still carries signal forward instead of being erased. GELU is non-monotone in the same way, with a shallower minimum near -0.17.
code
python · 15 linesimport math
def silu(x):
return x / (1.0 + math.exp(-x))
def gelu(x):
return x * 0.5 * (1.0 + math.erf(x / math.sqrt(2.0)))
grid = [-4.0 + 0.001 * i for i in range(8001)]
xs = min(grid, key=silu)
xg = min(grid, key=gelu)
print("SiLU min %.4f at x = %.3f" % (silu(xs), xs))
print("GELU min %.4f at x = %.3f" % (gelu(xg), xg))
for x in (-2.0, -0.5, 0.5, 2.0):
print("x=%5.1f silu=%7.4f gelu=%7.4f" % (x, silu(x), gelu(x)))go deeper
Know that SiLU and GELU are not monotone: their output falls a little below zero for mildly negative inputs and then comes back toward zero. Being able to sketch that shape is enough at this level.
Explain the mechanics: a linearly growing negative input against an exponentially closing gate produces a minimum, roughly -0.28 at -1.28 for SiLU. State the derivative's sign on either side and why continuity is what optimisation actually needs.
Demonstrate calibrated judgment. Say plainly that theory permits the dip and that the measured benefit on a given architecture is usually small, so it belongs in an ablation rather than in an argument.
Frame it as a research-versus-engineering call: how much of the team's time should go into activation-shape exploration when the effect size is small and architecture-dependent, compared with data, scale, or training-recipe work.
## Where the dip comes from SiLU is `f(x) = x * sigmoid(x)`, a product of two factors that pull in opposite directions on the negative side. As `x` moves left of the origin, the first factor becomes more negative while the second - the gate - shrinks toward 0. Near the origin the input's growth in magnitude wins and the product falls. Far to the left the gate's decay wins, because a logistic gate decays roughly exponentially while the input only grows linearly, so the product turns back toward 0. Between the two regimes there is a minimum. Numerically the minimum is about **-0.2785 at x = -1.278**. On either side the function is closer to 0: it is about -0.238 at -2, and about -0.142 at -3. GELU has the same shape for the same reason, with a shallower minimum of about **-0.17 near x = -0.752**, because a Gaussian CDF closes faster than a logistic curve. So neither smooth gated unit is monotone, and that is a property of the construction, not an accident of one of them. ## Is non-monotonicity a problem? It is worth being precise about what monotonicity is and is not needed for. **Not needed for expressive power.** The classical universal-approximation arguments need a nonlinearity that is not a polynomial, plus enough width. Monotonicity is not part of the requirement. Non-monotone activations approximate functions perfectly well. **Not needed for backpropagation.** Backpropagation multiplies derivatives along the chain. It needs the derivative to exist and to be numerically sane; it does not care about its sign. SiLU's derivative is `s(x) * (1 + x * (1 - s(x)))` with `s = sigmoid`, and it is continuous everywhere, equal to 0.5 at the origin, and negative for `x` below about -1.278. A negative local derivative simply means an increase in the pre-activation there decreases the output - a legitimate local response, the same thing that happens on the falling side of any bump in a learned function. **What it does cost.** The unit is not invertible: two different pre-activations can give the same output, one on each side of the minimum. That matters if you were relying on a per-unit monotone response for interpretation - for example, arguing "a larger pre-activation always means a stronger downstream signal from this unit". With SiLU that argument is false in the dip region. It also means loss surfaces can be slightly more textured than with a monotone unit; in practice this has not stopped anyone. ## What the dip buys 1. **A bounded negative side.** `f(x) >= -0.2785` for all `x`. However negative a pre-activation drifts during training, the unit cannot emit an arbitrarily large negative value; the gate clamps it. That is a mild self-limiting property that a linear-on-the-negative-side unit does not have. 2. **Preserved small negatives.** A pre-activation of -0.5 leaves as about -0.19 rather than being erased. The next layer can still distinguish "slightly negative" from "very negative", because those map to different outputs on the rising part of the curve. 3. **A non-zero gradient on the negative side.** Because the function is not flat for negative inputs, units sitting in the negative region still receive gradient and can move, rather than being permanently stuck at a constant output. ## Where the shape came from Swish - the same function as SiLU - was found by an automated search over compositions of simple primitives, and the shape the search kept returning was self-gated and non-monotone. That is worth mentioning in an interview because it reframes the design: the dip is not a hand-crafted idea someone argued for from first principles, it is a shape that repeatedly came out well and was then explained afterwards. ## How to talk about it A good answer separates three claims. *Descriptive*: SiLU dips to about -0.28 near -1.28 and GELU dips to about -0.17 near -0.75. *Theoretical*: nothing in approximation theory or gradient-based optimisation requires monotone activations, only a well-behaved derivative. *Empirical and honest*: the accuracy effect of the dip on any given model is usually small and architecture-dependent, so treat "the dip helps" as a hypothesis you would test on your own model rather than a law. Candidates who assert a large guaranteed gain are overclaiming just as badly as candidates who assert the dip breaks training.
- Is GELU monotone, or does it dip too?It dips too, for the same structural reason: the input grows in magnitude while the Gaussian CDF gate closes. Its minimum is about -0.17 near `x = -0.752`, shallower and closer to the origin than SiLU's, because a Gaussian CDF closes faster than a logistic curve. Neither smooth gated unit is monotone.
- Over which inputs is SiLU's derivative negative?For inputs below about -1.278, the location of the minimum. Left of it the function is falling as `x` increases toward the minimum, so the derivative there is negative; right of it, the derivative is positive. The derivative formula is `s(x) * (1 + x * (1 - s(x)))` with `s = sigmoid`, and it is continuous through the whole range.
- Does the dip mean gradients vanish for very negative inputs?Not exactly. The output tends to 0 for very negative inputs and so does the derivative, so a unit driven far negative does contribute little gradient - but it approaches that limit smoothly and is never identically flat, so the unit can still drift back. That is different from a region with a derivative that is exactly zero.
saying these in an interview costs you the question
- Claims non-monotone activations break backpropagation
- Says monotonicity is required for universal approximation
- Thinks SiLU is unbounded below on the negative side
- Believes the derivative is undefined at the minimum
- Asserts the dip guarantees a large accuracy gain