skip to content

Sigmoid and Tanh

Sigmoid and tanh squash their input into a bounded range, so the derivative collapses toward zero far from the origin. Interviewers ask because that saturation is why deep stacks of them stalled.

on this pageshow

questions

3

What happens to a sigmoid hidden unit's gradient when its input is far from zero?

level: juniorimportance: must knowfreq 70%

answer

  1. the curve goes flat at the ends
  2. slope is the transmission coefficient
  3. the peak sits at the middle of the range
  4. output times one minus output
  5. 0.25 at best, near nothing in the tails

basics

~20 s

A sigmoid unit's gradient collapses toward zero. The curve flattens at both extremes, so its derivative - at most 0.25, at the origin - becomes almost nothing in the tails, and the layers below stop learning.

solid answer

~50 s

The sigmoid squashes any real input into (0, 1) and flattens out at both ends, so far from zero the output barely responds to the input. That flatness is exactly the derivative: `sigma'(x) = sigma(x) * (1 - sigma(x))`, which peaks at 0.25 when the output is 0.5 and falls off fast - 0.0099 at an output of 0.99, about 0.001 at 0.999. Backpropagation multiplies the incoming error signal by that factor at the unit, so a saturated unit attenuates it by two or three orders of magnitude, and everything beneath it barely moves. Even an unsaturated sigmoid attenuates by at least 4x, since 0.25 is its best case. Tanh has the same S-shape and the same tail problem; its peak derivative is 1.0 rather than 0.25, so it only starts from a better place.

go deeper

for a junior

Be ready to sketch the S-curve, say that it flattens at both ends, and state that a flat region means a near-zero slope and therefore almost no learning signal. Knowing the 0.25 peak by heart is a cheap win.

for a middle

You are expected to produce the derivative from the output - output times one minus output - and turn it into numbers on the spot, then explain that backpropagation multiplies by that factor at the unit, so the layers beneath it are the ones that stall.

for a senior

Show that you have seen the symptom: a loss that plateaus early, early-layer parameters that barely move while the top of the network still trains, and outputs pinned near an extreme. Name the pre-activation scale as the lever you actually pull.

for a principal

Own the framing that a bounded, fast-decaying derivative is a structural property, not a tuning problem. Argue about where in a design a saturating unit is worth its cost because its bounded output is the point, and where it is simply a trainability tax.

## The curve The logistic sigmoid maps any real number into the open interval (0, 1): `sigma(x) = 1 / (1 + exp(-x))` It is monotone and S-shaped: it passes through 0.5 at the origin, bends toward 0 as the input goes very negative, and toward 1 as it goes very positive. Tanh is the same shape with a different range: it maps into (-1, 1) and passes through 0 at the origin. **Saturation** is the name for a unit sitting out in one of those flat tails. Its output has effectively stopped responding to its input: push the input from 8 to 12 and a sigmoid moves from 0.99966 to 0.99999. ## Flatness is the derivative Backpropagation, at each unit, multiplies the error signal arriving from above by the derivative of that unit's activation function evaluated at the unit's own pre-activation (the weighted sum before the nonlinearity). So the slope of the curve *is* the transmission coefficient of the unit. The sigmoid has a convenient closed form for that slope, expressible in terms of its own output: `sigma'(x) = sigma(x) * (1 - sigma(x))` Read the numbers straight off it: - output 0.5 (input 0): `0.5 * 0.5 = 0.25` - the largest value the sigmoid derivative ever takes - output 0.9: `0.9 * 0.1 = 0.09` - output 0.99: `0.99 * 0.01 = 0.0099` - output 0.999: `0.999 * 0.001 = 0.000999` Tanh's derivative is `1 - tanh(x)^2`, which peaks at 1.0 at the origin and likewise decays toward 0 in the tails. Both functions are strongest in the middle and useless at the edges; they differ in how strong the middle is. ## Two separate costs A saturated unit hurts in two directions at once, and strong candidates name both. **Backward**: the tiny derivative shrinks the error signal passing through the unit, so the weights feeding it barely update and, because that shrunken signal continues downward, the layers below it barely update either. This is why the symptom usually shows up as *early layers frozen while the last layer still learns*. **Forward**: a unit whose output is 0.999 for every example in the dataset is a constant. Whatever information its inputs carried has been destroyed before it reaches the next layer. Even if you could somehow restore the gradient, the layer above has nothing to work with. ## Small, not zero A precise answer says *near* zero, not zero. The sigmoid derivative is strictly positive everywhere, so a saturated unit still receives a gradient and can in principle crawl back toward the responsive region. In practice, once the pre-activation is in the tens, the escape rate is so slow that the run finishes first. Treat it as recoverable in theory and stuck in practice - and note that this is a property of the *unit's operating point*, not of its weights being zero. ## How far is far? Useful landmarks for the sigmoid, all from `sigma(x)*(1-sigma(x))`: - `|x| = 2`: derivative about 0.105, roughly 40 percent of peak - still healthy - `|x| = 5`: about 0.0066, roughly one thirty-eighth of peak - `|x| = 10`: about 0.000045, four orders of magnitude below peak So saturation is not a cliff; it is a fast, smooth decay that becomes fatal somewhere around a pre-activation magnitude of five to ten. ## Where the large inputs come from The pre-activation is a weighted sum of the previous layer's outputs plus a bias. It grows with the scale of the incoming features, the scale of the weights, and the number of terms being summed. Unstandardized features with magnitudes in the tens will saturate a sigmoid layer immediately, before a single update has happened. This is why keeping pre-activations near the origin - where the curve actually has slope - is a first-class design concern rather than a detail. ## Why this shaped modern practice Every consequence above follows from one fact: a squashing nonlinearity has a bounded, quickly decaying derivative. That single fact is the reason saturating units were displaced from hidden layers in deep networks in favour of nonlinearities that do not flatten on the positive side, and the reason the ones that remain are used where their bounded output is the point rather than an accident.

  • Does tanh escape this problem?
    No. Tanh is the same S-shape with range (-1, 1), and its derivative `1 - tanh(x)^2` decays to nearly zero in exactly the same way once the pre-activation is large. What tanh buys you is the starting point: its peak derivative is 1.0 rather than 0.25, so a unit operating near the origin passes the signal through undiminished instead of quartering it. Saturation still ends the same way.
  • How large does a sigmoid unit's input have to be before you would call it saturated?
    There is no hard threshold, but the arithmetic gives landmarks. At a pre-activation of 2 the derivative is about 0.105, still workable. At 5 it is about 0.0066, roughly one thirty-eighth of the 0.25 peak. At 10 it is about 0.000045. Anything past magnitude five is functionally saturated for training purposes.
  • If the gradients are tiny, can you not simply raise the learning rate?
    Not safely. The attenuation is uneven - saturated units shrink their signal by three orders of magnitude while healthy ones do not - so a learning rate large enough to move the stuck part will destabilise the part that was training fine. The fix is to move the units back into the responsive region by controlling the scale of the pre-activations, not to compensate downstream.

A sigmoid unit in its tail is a dimmer switch already turned all the way up. Pushing the knob harder changes the brightness by nothing, so nothing downstream can tell you how hard you pushed.

saying these in an interview costs you the question

  • Says the gradient is exactly zero in the tails
  • Thinks the sigmoid's derivative peaks at 1
  • Confuses a saturated output with the unit's weights being zero
  • Claims a larger learning rate cures saturation
  • Mentions only the lost gradient and not the lost information forward
  • Believes saturation only matters at the output layer

context

open as a page

Why is tanh usually preferred over sigmoid for a network's hidden layers?

level: middleimportance: must knowfreq 62%

basics

~20 s

Tanh is zero-centred and steeper. Its outputs span (-1, 1) rather than (0, 1), so the next layer's weight gradients are not all forced to share a sign, and its peak derivative is 1.0 against sigmoid's 0.25.

open as a page

Your sigmoid hidden layer outputs 0.999 for every training example - what caused it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The pre-activations are far too large and positive, because the input features were fed in unscaled. Decibel-scale loudness values in the tens, summed over many channels, push every unit deep into the sigmoid's flat tail.

open as a page