skip to content

In logistic regression, what does the sigmoid do to the linear score w*x + b?

level: juniorimportance: must knowfreq 82%

answer

  1. unbounded score, bounded output
  2. S-curve, never touches the ends
  3. score of zero maps to one half
  4. inverse of the log-odds transform
  5. slope peaks at 0.25, flat in the tails

basics

~20 s

The sigmoid squashes any linear score into the open range 0 to 1, monotonically: large positive scores approach 1, large negative scores approach 0, and a score of 0 maps to 0.5. The result is read as the event probability.

solid answer

~50 s

The linear part of the model produces an unbounded real number, `z = w*x + b`, which cannot be handed to anyone as a probability. The sigmoid `p = 1 / (1 + exp(-z))` maps that whole real line onto the open interval (0, 1): `z = 0` gives 0.5, `z = 2` gives about 0.88, `z = -2` about 0.12, and the ends are approached but never reached. It is strictly increasing, so it preserves the ordering the score already encoded - a patient with a higher risk score keeps the higher 30-day readmission probability. It is smooth, so the model stays differentiable end to end, and its inverse is the log-odds `log(p / (1 - p)) = z`, which is why a unit change in the score is a constant change in log-odds rather than in probability.

go deeper

for a junior

Be ready to draw the S-curve and say what it maps: any real score into the open range 0 to 1, approaching but never reaching the ends, with a score of 0 landing on 0.5.

for a middle

Explain the mechanics: the mapping is monotone so it preserves score order, its inverse is the log-odds so the model is linear in log-odds and not in probability, and its slope peaks at 0.25 and flattens in both tails.

for a senior

Show you know what the flat tails cost. Out where scores are extreme the probability barely responds, learning signals shrink, and effect sizes quoted in probability terms stop being comparable across rows. Be able to work a numeric example on the spot.

for a principal

Own the question of whether the number should be consumed as a probability at all. Squashing guarantees the range, not that the value matches reality, so decide who validates that before a clinician or planner is allowed to act on it.

## The problem the sigmoid solves A linear model computes a score by weighting the inputs and adding an intercept: ``` z = w1*x1 + w2*x2 + ... + wk*xk + b ``` That number is unbounded - it can be -14.2 or +31.7. If you are predicting whether a patient will be readmitted within 30 days of discharge, a discharge planner cannot act on "31.7". They need a number between 0 and 1 that behaves like a probability. The sigmoid (also called the logistic function) is the squashing step that produces it. ## The function ``` sigmoid(z) = 1 / (1 + exp(-z)) ``` Properties worth being able to recite: - **Range is the open interval (0, 1).** As `z` grows the output approaches 1; as `z` falls it approaches 0. Neither end is ever reached, so the model never claims literal certainty. - **sigmoid(0) = 0.5.** A zero score means the model has no net evidence either way. - **Strictly increasing.** The mapping is monotone, so the ranking of rows by score and the ranking by probability are identical. Anything that depends only on ordering is unchanged by the squashing. - **Symmetric:** `sigmoid(-z) = 1 - sigmoid(z)`. Flipping the sign of the score flips the predicted probability around 0.5. - **Inverse is the log-odds:** if `p = sigmoid(z)` then `log(p / (1 - p)) = z`. The linear model is therefore linear in log-odds, not in probability. - **Derivative:** `sigmoid'(z) = sigmoid(z) * (1 - sigmoid(z))`. It peaks at 0.25 when `z = 0` and decays towards zero in both tails - a fact that matters a great deal once you start training. ## Reading the numbers Some anchor values, useful in an interview because they let you sanity-check any claim on the spot: ``` z: -4 -2 -1 0 1 2 4 p: 0.018 0.119 0.269 0.500 0.731 0.881 0.982 ``` Notice the compression. Moving the score from 0 to 1 moves the probability by 23 points; moving it from 3 to 4 moves it by about 3 points. The same additive change in score has a large effect near the middle and almost none in the tails. That is exactly what "linear in log-odds" means in practice, and it is why you cannot describe a fitted weight as "adds 8% to the probability" - the effect on probability depends on where the row already sits. For the readmission model: a patient whose covariates give `z = 1.2` gets a predicted 30-day readmission probability of about 0.77. Add a comorbidity worth +0.4 in score and you get about 0.83, a 6-point rise. Apply the same +0.4 to a low-risk patient at `z = -3.0` and the probability moves from 0.047 to 0.068 - two points. Same weight, very different probability effect. ## Why not something simpler Why not clip the raw score into [0, 1]? Clipping is not differentiable at the two kinks and is exactly flat outside the range, so any row that lands outside gives no learning signal at all and there is no principled reason to prefer the resulting numbers as probabilities. The sigmoid, by contrast, is the function that turns a linear model into a Bernoulli probability model: it is the exact inverse of the log-odds transform, which means the fitted model has a likelihood you can write down and maximise. The training objective that does that maximising is cross-entropy, which is the natural companion question to this one. ## The saturation caveat The flat tails are not free. Once `z` is large in magnitude, the output barely responds to further change in `z`, and the derivative `p * (1 - p)` becomes tiny: at `p = 0.999` it is about 0.001, four hundred times smaller than its peak. Any training rule whose update is multiplied by that derivative will crawl in the tails. This is the reason logistic regression is not trained with squared error, and it is worth being able to say so. ## Common mistakes - Claiming the sigmoid outputs 0 or 1 exactly for extreme scores. It approaches them asymptotically. - Treating the output as a calibrated probability by construction. The sigmoid guarantees only the *range* and the monotone shape; whether the number matches observed frequencies is a separate property of the fitted model. - Interpreting a weight as a change in probability. It is a change in log-odds.

  • Why not simply clip the linear score into the range 0 to 1 instead of passing it through a sigmoid?
    Clipping is flat outside the range, so rows that land there produce no gradient and never learn, and it has two non-differentiable kinks. It also has no probabilistic justification: the sigmoid is the exact inverse of the log-odds transform, so the model it defines has a Bernoulli likelihood you can write down and maximise. Clipped scores are just truncated numbers.
  • What does a one-unit increase in the linear score do to the predicted probability?
    It adds exactly one to the log-odds, which is a constant effect; the effect on the probability itself is not constant. Near a score of 0 a unit increase moves the probability by roughly 23 points; out at a score of 3 it moves it by about 3 points. That is why you report a weight as an effect on log-odds, never as a fixed percentage-point change.
  • Why does the sigmoid never output exactly 0 or 1?
    Because `exp(-z)` is strictly positive for every finite `z`, so `1 / (1 + exp(-z))` is strictly between 0 and 1. Only an infinite score would reach the ends. This matters during training, since the loss involves the logarithm of the predicted probability and an exact 0 or 1 would be undefined; in floating point, extreme scores can still round to the endpoints, which is why the two steps are usually computed together rather than separately.

The linear score is a raw tally of evidence that can run to any size; the sigmoid is the dial that maps that tally onto a gauge marked 0 to 1, with most of the dial's travel spent near the middle and the ends squeezed almost flat.

saying these in an interview costs you the question

  • Says the sigmoid outputs exactly 0 or 1 for large scores
  • Describes a weight as adding a fixed number of percentage points
  • Claims the sigmoid makes the output a trustworthy probability by itself
  • Thinks squashing changes the ranking of rows by score
  • Cannot state sigmoid(0) = 0.5 without derivation

context