skip to content

Why is tanh usually preferred over sigmoid for a network's hidden layers?

level: middleimportance: must knowfreq 62%

answer

  1. compare two things, not one
  2. one argument is about slope
  3. the other is about the sign of the outputs
  4. 1.0 against 0.25 at the origin
  5. all-positive inputs lock a weight row's signs

basics

~20 s

Tanh is zero-centred and steeper. Its outputs span (-1, 1) rather than (0, 1), so the next layer's weight gradients are not all forced to share a sign, and its peak derivative is 1.0 against sigmoid's 0.25.

solid answer

~50 s

Two reasons, and a good answer gives both. First, the derivative: tanh peaks at 1.0 at the origin while the sigmoid peaks at 0.25, so near initialization a tanh unit passes the error signal through roughly intact whereas a sigmoid quarters it - stack ten sigmoid layers and the activation derivatives alone contribute a factor of at most `0.25^10`, about 1e-6. Second, the output range: sigmoid outputs are strictly positive, and since the gradient of a weight is the downstream unit's error term times the incoming activation, every weight feeding one unit gets a same-signed gradient for a given example. The weight vector can then only move into one orthant at a time, so it zig-zags toward optima that need some weights up and others down. Tanh's roughly zero-centred outputs remove that constraint. Neither fixes saturation - tanh still flattens in its tails.

go deeper

for a junior

Know the ranges - (0, 1) for sigmoid, (-1, 1) for tanh - and be able to say that tanh is centred on zero and has a steeper middle. That much already separates you from a candidate who has only seen the names.

for a middle

This is your tier. Give both arguments cleanly: the 1.0 versus 0.25 peak derivative and what it costs per layer, and the fact that all-positive inputs force every weight feeding one unit to receive a same-signed gradient for a given example.

for a senior

Add the caveats an interviewer is fishing for: batching softens the sign bias, zero-centring depends on the pre-activations being symmetric, and tanh saturates just as hard in its tails, so it is a better starting point rather than a fix for depth.

for a principal

Be ready to argue where the choice actually sits in a design. The two functions are affine reparameterizations of each other, so the win is in conditioning, not capacity - which reframes the question as one about initialization scale, normalization and depth budget rather than about picking a curve.

## Two independent arguments The preference for tanh over the logistic sigmoid in hidden layers rests on two separate properties. They are often conflated in interviews; keeping them apart is most of the answer. ## Argument one: the size of the derivative Backpropagation multiplies the error signal by each unit's activation derivative as it passes through. The sigmoid derivative is `sigma(x) * (1 - sigma(x))`, whose largest possible value is `0.5 * 0.5 = 0.25`, attained at the origin. Tanh's derivative is `1 - tanh(x)^2`, whose largest value is 1.0, also at the origin. So in the best case - a unit operating right in the middle of its range, which is where small initial weights put it - a sigmoid layer shrinks the signal by a factor of four while a tanh layer passes it through unchanged. Near the origin tanh is approximately the identity function, `tanh(x) ~= x` for small x, which is exactly the behaviour you want a fresh network to have. Compound that over depth and the difference is not subtle. From the activation derivatives alone, ten stacked sigmoid layers contribute a factor of at most `0.25^10`, roughly 1e-6, before anything else has had a chance to help or hurt. Ten tanh layers contribute at most 1.0. ## Argument two: the sign of the outputs The sigmoid's range is (0, 1) - strictly positive, with a mean around 0.5 for typical inputs. Tanh's range is (-1, 1), centred on 0. Why that matters comes from the shape of a weight gradient. For a weight `w_ij` connecting incoming activation `h_j` to downstream unit `i`, the gradient is `dL/dw_ij = delta_i * h_j` where `delta_i` is the error term at unit `i`. Now fix one downstream unit `i` and one example. `delta_i` is a single number with a single sign. If every `h_j` is positive - which is guaranteed when the layer below is sigmoid - then every one of that unit's incoming weight gradients carries the sign of `delta_i`. The whole row moves up together or down together. Geometrically, the update for that weight row is confined to a single orthant per example. If the loss is minimised at a point that needs one weight larger and another smaller, the optimizer cannot walk there directly; it has to alternate, producing a zig-zag staircase of small steps down a narrow valley instead of a straight descent. Zero-centred activations, which produce a mix of positive and negative `h_j`, lift that restriction. **The honest caveat**: the argument is per-example. In mini-batch training, the batch gradient averages over examples whose `delta_i` may have different signs, so the sign bias is softened rather than eliminated. It bites hardest when the errors across a batch agree, which is common early in training. And tanh's outputs are only zero-centred if the pre-activations are roughly symmetric about zero - a tanh layer whose inputs are all strongly positive is no better centred than a sigmoid one. ## What tanh does not fix Tanh is not a solution to vanishing gradients through depth, only a better starting point. It has the same flat tails: once a pre-activation reaches magnitude five or so, `1 - tanh(x)^2` is tiny and the unit is as saturated as a sigmoid would be. A plain twenty-layer tanh stack over sixty-channel accelerometer windows will happily show first-layer gradients around 1e-9 while the last layer trains normally - the top of the network learns, the bottom is frozen, and the loss plateaus above where it should. Depth needs structural help beyond swapping the squashing function. ## The identity that catches people out The two functions are related exactly: `tanh(x) = 2 * sigma(2x) - 1` A tanh unit is an affine rescaling of a sigmoid unit with a doubled input. That means the two choices are representationally equivalent: a network can express identical functions either way, by absorbing the factor of 2 into the following layer's weights, the -1 into its bias, and the input doubling into the preceding weights. So why does the choice matter at all? Because gradient descent is not invariant to reparameterization. With the same initial weight scale, the same learning rate and the same weight decay, the two parameterizations trace different trajectories, and the tanh one is the better-conditioned of the two. This is a useful thing to be able to say out loud: the benefit is an optimization benefit, not an expressiveness one, and rescaling the outputs after the fact does not recover it. ## The short version to say in a room Steeper where it matters, symmetric about zero, and still saturating - so tanh beats sigmoid in hidden layers without being a cure for depth.

  • Since `tanh(x) = 2 * sigma(2x) - 1`, why does swapping one for the other change anything?
    Representationally it does not: a network can absorb the factor of 2 into the next layer's weights, the -1 into its bias, and the input doubling into the previous weights, so both express the same functions. What differs is optimization. Gradient descent is not invariant to reparameterization, so at a fixed initial weight scale, learning rate and weight decay the two versions follow different trajectories - and the tanh parameterization is the better-conditioned one.
  • Does the same-signed gradient argument survive mini-batch training?
    Only partly. The constraint is per-example: for one example and one downstream unit, all incoming weight gradients share the sign of that unit's error term. Averaging over a batch mixes examples whose error terms differ in sign, so the batch gradient can point in mixed directions. The bias is softened, not removed, and it is worst when the errors across the batch agree - which is typical early in training.
  • Does switching to tanh fix vanishing gradients in a deep stack?
    No. Tanh has the same flat tails, so any unit with a pre-activation past magnitude five is saturated and passes almost nothing through. The 1.0 peak only helps units operating near the origin, and depth pushes them away from it. A plain twenty-layer tanh network can show first-layer gradients around 1e-9 while its last layer trains fine.

If every step you take must move all of your coordinates in the same direction, then reaching a spot that needs one coordinate up and another down takes an alternating staircase of small steps rather than one straight walk.

saying these in an interview costs you the question

  • Says tanh has a fundamentally different shape from sigmoid
  • Claims tanh does not saturate
  • Gives only the output range and never mentions the derivative
  • States tanh eliminates vanishing gradients in deep networks
  • Thinks zero-centred guarantees the outputs average exactly zero
  • Says sigmoid's problem is that its output is bounded

context