Why does the logistic function s(x) = 1/(1 + e^-x) have derivative s(x)(1 - s(x))?
answer
- quotient rule on one over something
- derivative of e^-x carries a minus sign
- denominator gets squared
- factor out 1/(1 + e^-x) twice
- the leftover factor equals 1 - s
basics
~20 sThe quotient rule turns 1/(1 + e^-x) into e^-x/(1 + e^-x)^2. That expression splits into 1/(1 + e^-x) times e^-x/(1 + e^-x), and the second factor is exactly 1 minus the first, giving s times (1 - s).
solid answer
~40 sWrite `s = u/v` with `u = 1` and `v = 1 + e^-x`. The quotient rule says `(u/v)' = (u'v - u*v')/v^2`. Here `u' = 0` and `v' = -e^-x`, so `s'(x) = (0 - 1*(-e^-x))/(1 + e^-x)^2 = e^-x/(1 + e^-x)^2`. Now factor: `e^-x/(1 + e^-x)^2 = [1/(1 + e^-x)] * [e^-x/(1 + e^-x)]`. The first bracket is `s(x)`, and the second is `1 - s(x)` because `1 - 1/(1 + e^-x) = e^-x/(1 + e^-x)`. Hence `s' = s(1 - s)`. The identity is exact, not an approximation, and it is useful because the slope is recovered from the already-computed output value. It also shows the slope peaks at `1/4` when `s = 1/2`, and collapses toward 0 as the output saturates near 0 or 1.
go deeper
Recall the shape and the fact that the derivative can be written in terms of the output itself. Knowing the sign of the derivative of e^-x is the piece most often dropped.
Carry out the quotient-rule computation on the spot and then perform the factoring step that reveals s(1 - s). Both halves are expected, not just the final formula.
Read the consequences off the formula: strictly positive slope, a maximum of one quarter at the centre, and saturation that makes the curve nearly flat at both extremes.
Weigh where a saturating transform is the right modelling choice at all, and be ready to justify picking a bounded squashing function over an unbounded one for a given quantity.
### The function The logistic (sigmoid) function is ``` s(x) = 1 / (1 + e^-x) ``` It is defined for every real `x`, takes values strictly between 0 and 1, equals `1/2` at `x = 0`, and increases monotonically from a lower asymptote of 0 to an upper asymptote of 1. It is the standard way to squash an unbounded real number into a probability-shaped quantity. ### Ingredient: the derivative of e^-x The exponential is the function that is its own derivative: `d/dx e^x = e^x`. For `e^-x`, note `e^-x = 1/e^x`, and apply the quotient rule to that: with `u = 1`, `v = e^x`, we get `(0*e^x - 1*e^x)/(e^x)^2 = -e^x/e^(2x) = -1/e^x = -e^-x`. So `d/dx e^-x = -e^-x`: same magnitude, opposite sign. Getting this sign wrong is the single most common error in the derivation. ### The quotient rule, stated precisely ``` (u/v)' = (u'*v - u*v') / v^2 ``` Two details trip people up: the numerator has a **minus** sign (unlike the product rule, which has a plus), and the order matters — it is `u'v` first, then subtract `uv'`. The denominator is squared. ### The derivation Take `u = 1` and `v = 1 + e^-x`, so `u' = 0` and `v' = -e^-x`. ``` s'(x) = (0*(1 + e^-x) - 1*(-e^-x)) / (1 + e^-x)^2 = e^-x / (1 + e^-x)^2 ``` That is already a complete answer. The elegant part is recognising it as a product of two copies of the function itself: ``` e^-x / (1 + e^-x)^2 = [ 1/(1 + e^-x) ] * [ e^-x/(1 + e^-x) ] ``` The first bracket is `s(x)` by definition. For the second, compute ``` 1 - s(x) = 1 - 1/(1 + e^-x) = (1 + e^-x - 1)/(1 + e^-x) = e^-x/(1 + e^-x) ``` So the second bracket is exactly `1 - s(x)`, and therefore ``` s'(x) = s(x) * (1 - s(x)) ``` This is an identity, valid for every real `x`, not an approximation. ### What the formula tells you **The slope is always positive.** Both `s` and `1 - s` lie strictly in `(0, 1)`, so the product is strictly positive: the logistic function is strictly increasing everywhere and has no flat spot, no maximum and no minimum. **The slope is capped at 1/4.** Regard `s(1-s)` as a function of the value `p = s` on the interval `(0, 1)`. The expression `p(1-p)` is a downward parabola peaking at `p = 1/2`, where it equals `1/4`. Since `s(0) = 1/2`, the steepest point of the curve is at the origin with slope exactly `0.25`. **The slope collapses at the extremes.** For large positive `x`, `s` is close to 1 and `1 - s` is close to 0, so the product is tiny; for large negative `x` the roles swap and the product is tiny again. This is the *saturation* behaviour: far from the centre the curve is nearly flat, so a change in the input barely moves the output. **The slope is symmetric.** A useful identity is `s(-x) = 1 - s(x)`, from which `s'(-x) = s(-x)(1 - s(-x)) = (1 - s(x))*s(x) = s'(x)`. The derivative curve is a symmetric bump centred at the origin. ### Self-referential derivatives elsewhere The logistic is not the only function whose derivative is expressible in its own value. The hyperbolic tangent satisfies `d/dx tanh(x) = 1 - tanh(x)^2`, and the two are related by `tanh(x) = 2*s(2x) - 1`. The exponential is the extreme case: `d/dx e^x = e^x`. Whenever a derivative can be written in terms of the function's own output, evaluating the slope costs almost nothing once the value is known, which is why these forms are quoted so often. ### The inverse relationship Solving `p = 1/(1 + e^-x)` for `x` gives `x = ln(p/(1 - p))`, the log-odds or logit. The quantity `p(1-p)` reappearing as the derivative is not a coincidence: it is the same product that measures the spread of a two-outcome random quantity with probability `p`, which is why the logistic curve is flattest where the outcome is nearly certain and steepest where it is a coin flip. ### Interview traps The common errors are: writing `d/dx e^-x = e^-x` and losing the minus sign; using the product rule sign pattern inside the quotient rule; forgetting to square the denominator; and claiming the slope can approach 1. Being able to state the maximum slope of `1/4` immediately is a good signal that the formula is understood rather than memorised.
- What is the maximum value of the logistic derivative and where does it occur?It is `1/4`, attained at `x = 0`. Treating `s(1-s)` as `p(1-p)` for `p` in `(0, 1)` gives a downward parabola peaking at `p = 1/2`, and the logistic function takes the value `1/2` exactly at the origin. So the curve is at its steepest in the middle, with slope `0.25`.
- What does the formula s(1 - s) say about the logistic curve far from the origin?It saturates. For large positive input `s` approaches 1, so `1 - s` approaches 0 and the product collapses; for large negative input the same happens with the roles swapped. The slope becomes vanishingly small, so large changes in the input barely move the output near either asymptote.
- What is the derivative of tanh, and how does it compare?`d/dx tanh(x) = 1 - tanh(x)^2`, another derivative written purely in terms of the function's own output. The two are related by `tanh(x) = 2*s(2x) - 1`. Its peak slope is 1 at the origin, four times the logistic peak of `1/4`, and it also saturates at both ends.
Think of an S-shaped adoption curve. Growth is fastest when about half the population has adopted and there are both plenty of adopters to spread the word and plenty of non-adopters left; at either extreme one of those two pools is empty and progress nearly stops. The derivative s(1-s) is exactly that product of the two pools.
saying these in an interview costs you the question
- States that the derivative of e^-x is e^-x
- Writes the quotient rule numerator as u'v + uv'
- Forgets to square the denominator in the quotient rule
- Claims the logistic slope can approach 1
- Treats s(1 - s) as an approximation rather than an exact identity