skip to content

Sigmoid Classifiers

Passing a linear score through the sigmoid turns it into a probability, trained by cross-entropy rather than squared error. 'Explain logistic regression' is the canonical whiteboard prompt.

on this pageshow

explore

questions

10

In logistic regression, what does the sigmoid do to the linear score w*x + b?

level: juniorimportance: must knowfreq 82%

answer

  1. unbounded score, bounded output
  2. S-curve, never touches the ends
  3. score of zero maps to one half
  4. inverse of the log-odds transform
  5. slope peaks at 0.25, flat in the tails

basics

~20 s

The sigmoid squashes any linear score into the open range 0 to 1, monotonically: large positive scores approach 1, large negative scores approach 0, and a score of 0 maps to 0.5. The result is read as the event probability.

solid answer

~50 s

The linear part of the model produces an unbounded real number, `z = w*x + b`, which cannot be handed to anyone as a probability. The sigmoid `p = 1 / (1 + exp(-z))` maps that whole real line onto the open interval (0, 1): `z = 0` gives 0.5, `z = 2` gives about 0.88, `z = -2` about 0.12, and the ends are approached but never reached. It is strictly increasing, so it preserves the ordering the score already encoded - a patient with a higher risk score keeps the higher 30-day readmission probability. It is smooth, so the model stays differentiable end to end, and its inverse is the log-odds `log(p / (1 - p)) = z`, which is why a unit change in the score is a constant change in log-odds rather than in probability.

go deeper

for a junior

Be ready to draw the S-curve and say what it maps: any real score into the open range 0 to 1, approaching but never reaching the ends, with a score of 0 landing on 0.5.

for a middle

Explain the mechanics: the mapping is monotone so it preserves score order, its inverse is the log-odds so the model is linear in log-odds and not in probability, and its slope peaks at 0.25 and flattens in both tails.

for a senior

Show you know what the flat tails cost. Out where scores are extreme the probability barely responds, learning signals shrink, and effect sizes quoted in probability terms stop being comparable across rows. Be able to work a numeric example on the spot.

for a principal

Own the question of whether the number should be consumed as a probability at all. Squashing guarantees the range, not that the value matches reality, so decide who validates that before a clinician or planner is allowed to act on it.

## The problem the sigmoid solves A linear model computes a score by weighting the inputs and adding an intercept: ``` z = w1*x1 + w2*x2 + ... + wk*xk + b ``` That number is unbounded - it can be -14.2 or +31.7. If you are predicting whether a patient will be readmitted within 30 days of discharge, a discharge planner cannot act on "31.7". They need a number between 0 and 1 that behaves like a probability. The sigmoid (also called the logistic function) is the squashing step that produces it. ## The function ``` sigmoid(z) = 1 / (1 + exp(-z)) ``` Properties worth being able to recite: - **Range is the open interval (0, 1).** As `z` grows the output approaches 1; as `z` falls it approaches 0. Neither end is ever reached, so the model never claims literal certainty. - **sigmoid(0) = 0.5.** A zero score means the model has no net evidence either way. - **Strictly increasing.** The mapping is monotone, so the ranking of rows by score and the ranking by probability are identical. Anything that depends only on ordering is unchanged by the squashing. - **Symmetric:** `sigmoid(-z) = 1 - sigmoid(z)`. Flipping the sign of the score flips the predicted probability around 0.5. - **Inverse is the log-odds:** if `p = sigmoid(z)` then `log(p / (1 - p)) = z`. The linear model is therefore linear in log-odds, not in probability. - **Derivative:** `sigmoid'(z) = sigmoid(z) * (1 - sigmoid(z))`. It peaks at 0.25 when `z = 0` and decays towards zero in both tails - a fact that matters a great deal once you start training. ## Reading the numbers Some anchor values, useful in an interview because they let you sanity-check any claim on the spot: ``` z: -4 -2 -1 0 1 2 4 p: 0.018 0.119 0.269 0.500 0.731 0.881 0.982 ``` Notice the compression. Moving the score from 0 to 1 moves the probability by 23 points; moving it from 3 to 4 moves it by about 3 points. The same additive change in score has a large effect near the middle and almost none in the tails. That is exactly what "linear in log-odds" means in practice, and it is why you cannot describe a fitted weight as "adds 8% to the probability" - the effect on probability depends on where the row already sits. For the readmission model: a patient whose covariates give `z = 1.2` gets a predicted 30-day readmission probability of about 0.77. Add a comorbidity worth +0.4 in score and you get about 0.83, a 6-point rise. Apply the same +0.4 to a low-risk patient at `z = -3.0` and the probability moves from 0.047 to 0.068 - two points. Same weight, very different probability effect. ## Why not something simpler Why not clip the raw score into [0, 1]? Clipping is not differentiable at the two kinks and is exactly flat outside the range, so any row that lands outside gives no learning signal at all and there is no principled reason to prefer the resulting numbers as probabilities. The sigmoid, by contrast, is the function that turns a linear model into a Bernoulli probability model: it is the exact inverse of the log-odds transform, which means the fitted model has a likelihood you can write down and maximise. The training objective that does that maximising is cross-entropy, which is the natural companion question to this one. ## The saturation caveat The flat tails are not free. Once `z` is large in magnitude, the output barely responds to further change in `z`, and the derivative `p * (1 - p)` becomes tiny: at `p = 0.999` it is about 0.001, four hundred times smaller than its peak. Any training rule whose update is multiplied by that derivative will crawl in the tails. This is the reason logistic regression is not trained with squared error, and it is worth being able to say so. ## Common mistakes - Claiming the sigmoid outputs 0 or 1 exactly for extreme scores. It approaches them asymptotically. - Treating the output as a calibrated probability by construction. The sigmoid guarantees only the *range* and the monotone shape; whether the number matches observed frequencies is a separate property of the fitted model. - Interpreting a weight as a change in probability. It is a change in log-odds.

  • Why not simply clip the linear score into the range 0 to 1 instead of passing it through a sigmoid?
    Clipping is flat outside the range, so rows that land there produce no gradient and never learn, and it has two non-differentiable kinks. It also has no probabilistic justification: the sigmoid is the exact inverse of the log-odds transform, so the model it defines has a Bernoulli likelihood you can write down and maximise. Clipped scores are just truncated numbers.
  • What does a one-unit increase in the linear score do to the predicted probability?
    It adds exactly one to the log-odds, which is a constant effect; the effect on the probability itself is not constant. Near a score of 0 a unit increase moves the probability by roughly 23 points; out at a score of 3 it moves it by about 3 points. That is why you report a weight as an effect on log-odds, never as a fixed percentage-point change.
  • Why does the sigmoid never output exactly 0 or 1?
    Because `exp(-z)` is strictly positive for every finite `z`, so `1 / (1 + exp(-z))` is strictly between 0 and 1. Only an infinite score would reach the ends. This matters during training, since the loss involves the logarithm of the predicted probability and an exact 0 or 1 would be undefined; in floating point, extreme scores can still round to the endpoints, which is why the two steps are usually computed together rather than separately.

The linear score is a raw tally of evidence that can run to any size; the sigmoid is the dial that maps that tally onto a gauge marked 0 to 1, with most of the dial's travel spent near the middle and the ends squeezed almost flat.

saying these in an interview costs you the question

  • Says the sigmoid outputs exactly 0 or 1 for large scores
  • Describes a weight as adding a fixed number of percentage points
  • Claims the sigmoid makes the output a trustworthy probability by itself
  • Thinks squashing changes the ranking of rows by score
  • Cannot state sigmoid(0) = 0.5 without derivation

context

open as a page

Why is a logistic regression's decision boundary a straight line in feature space?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Logistic regression scores each point with one weighted sum. The sigmoid is monotone, so a 0.5 probability cut is exactly the rule score above zero, and the set where the score equals zero is flat: a line, plane or hyperplane.

open as a page

What does the softmax function do to the class scores of a multiclass linear model?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Softmax exponentiates each of the K class scores and divides by their total, turning them into K positive numbers that add up to one. The largest score keeps the largest probability, so the predicted class is unchanged.

open as a page

Why is logistic regression trained with cross-entropy loss rather than squared error?

level: middleimportance: must knowfreq 76%

basics

~20 s

Cross-entropy is the Bernoulli negative log-likelihood, and its gradient on the linear score is simply predicted minus actual. Squared error through a sigmoid is non-convex in the weights and its gradient nearly vanishes exactly where the model is confidently wrong.

open as a page

Why can a linear classifier not separate an XOR pattern over two binary features?

level: middleimportance: must knowfreq 61%

basics

~20 s

A linear classifier cuts feature space with one flat boundary and labels each side uniformly. XOR puts its two positive cells on one diagonal and its two negative cells on the other, and crossing diagonals cannot be split by a straight line.

open as a page

How do you read a multiclass log loss of 1.2 from a 20-language classifier?

level: middleimportance: should knowfreq 44%

basics

~20 s

Multiclass log loss averages the negative log of the probability given to the true class. Uniform guessing over 20 languages scores log 20, about 3.0, so 1.2 beats chance and means an average true-class probability near 0.30.

open as a page

Would you fit one multinomial softmax model or 40 one-vs-rest models for 40 categories?

level: middleimportance: should knowfreq 52%

basics

~20 s

One multinomial fit trains all 40 classes jointly through a shared normaliser, so its probabilities already sum to one. Forty one-vs-rest fits are independent, so their scores need renormalising, while one-vs-one would need 780 models.

open as a page

A clinic no-show model weights 'previous no-shows' negatively - how do you investigate before launch?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A wrong-sign weight tilts the decision boundary against domain sense, so treat it as a suspected defect: check how the label and the feature are coded and computed, whether a correlated predictor absorbs the effect, and whether the sign is stable.

open as a page

Your logistic uptake model's scores will be reused as propensity scores - what changes in how you fit it?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

The label becomes the treatment indicator and the probability level matters more than the ranking. Keep confounders even when they add no accuracy, exclude anything measured after treatment, and do not rebalance the classes or over-penalise the weights.

open as a page

Why does adding 5 to every logit of a softmax classifier leave its probabilities unchanged?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Softmax divides by the sum of the exponentials, so a constant added to every score contributes the same factor to numerator and denominator and cancels. Only differences between class scores are identified, which leaves softmax over-parameterised.

open as a page