Why is a logistic regression's decision boundary a straight line in feature space?
answer
- one weighted sum, then a squash
- the sigmoid never reverses order
- 0.5 out means zero in
- the set where the score equals zero
- flat: line, plane, hyperplane
basics
~20 sLogistic regression scores each point with one weighted sum. The sigmoid is monotone, so a 0.5 probability cut is exactly the rule score above zero, and the set where the score equals zero is flat: a line, plane or hyperplane.
solid answer
~50 sLogistic regression computes a single linear score `s = w.x + b` and squashes it with the sigmoid `p = 1 / (1 + exp(-s))`. The sigmoid is strictly increasing and returns 0.5 exactly when `s = 0`, so the default `p > 0.5` rule is identical to `s > 0`. The boundary is therefore the solution set of `w.x + b = 0`, which is a line in two dimensions, a plane in three, a hyperplane in general. The curve people picture is the S-shaped probability against the score, not the boundary in feature space. The weight vector fixes the boundary's orientation and the intercept slides it: bump `b` and the same line moves parallel to itself, which is why a comment-moderation score built from caps-ratio and link-count keeps its tilt but flags more comments as `b` rises.
go deeper
Be ready to say the boundary is where the linear score equals zero, and that this is the same place the predicted probability equals 0.5. Draw a line on a two-feature scatter if asked.
Explain the mechanics: the sigmoid is strictly increasing, so a probability cut maps to a score cut. Say which parameter rotates the boundary and which one slides it parallel.
Show you use the geometry diagnostically: check that error regions are not concentrated on one side of the line, and recognise when flat-cut-only is the wrong model class for the data you have.
Own the trade you are making by choosing a flat boundary: an auditable, cheap, stable rule whose limits are known in advance, against a flexible model whose shape you cannot describe to a reviewer.
## The two pieces A binary logistic regression does exactly two things to an input. First it forms a **linear score**, one number per row: ``` s = w1*x1 + w2*x2 + ... + wd*xd + b ``` The `w`s are the weights, one per feature, and `b` is the intercept (bias). Second it passes that score through the **sigmoid**, `p = 1 / (1 + exp(-s))`, which squashes any real number into (0, 1) so it can be read as the probability of the positive class. The default decision rule is: predict the positive class when `p > 0.5`. ## Why 0.5 on the probability means 0 on the score The sigmoid is strictly increasing: bigger score, bigger probability, never a reversal. And `sigmoid(0) = 1 / (1 + 1) = 0.5`. Put those together and `p > 0.5` holds **if and only if** `s > 0`. The probability cut is just a relabelled score cut. Nothing about the S-shape survives the translation, because a strictly increasing function preserves the ordering and moves the single cut point 0.5 back to the single cut point 0. So the **decision boundary** — the set of points the model is undecided about — is ``` w1*x1 + w2*x2 + ... + wd*xd + b = 0 ``` That is a linear equation in the features. Its solution set is a line in two dimensions, a plane in three, and a flat *hyperplane* of one dimension less than the feature space in general. Flat, unbent, unbounded. ## The picture people confuse it with The famous S-curve is a plot of `p` against `s` — probability on the vertical axis, score on the horizontal. It is curved, and it has nothing to do with the shape of the boundary. If you instead plot the two features against each other and colour the points by class, the fitted boundary is literally a line you can draw with a ruler. A two-feature university-admissions model — test score on one axis, GPA on the other — is the cleanest demonstration: the fitted rule is one straight cut across the scatter, admits above it, rejects below it. ## What the weights and the intercept each do Split the parameters by their geometric job: - **The weight vector sets the orientation.** It is the normal (perpendicular) direction to the boundary. Raising `w1` relative to `w2` **rotates** the line: the model starts leaning more on the first feature. Flipping the sign of every weight and the intercept keeps the same line but swaps which side is positive. - **The intercept sets the offset.** Changing `b` alone **translates** the line, sliding it parallel to itself without changing its tilt. In a comment-moderation score built from caps-ratio and link-count, raising `b` keeps exactly the same trade between the two features but pushes the whole line down, so more comments land on the flagged side. - **The overall magnitude sets the steepness, not the position.** Multiply every weight *and* the intercept by 10 and the set where `s = 0` is unchanged — the boundary does not move. What changes is how fast `p` swings from near 0 to near 1 as you step away from it. Big weights mean a confident, almost step-like transition; small weights mean a gentle ramp with lots of probabilities near 0.5. A useful related fact: the signed distance from a point to the boundary is `s / ||w||`. Points far on the positive side get large scores and probabilities near 1; points sitting exactly on the boundary get `s = 0` and `p = 0.5` by construction. ## Contours, not just the boundary Every constant-probability contour is also flat. `p = 0.9` is the set where `s` equals a particular constant, another hyperplane, parallel to the boundary. The whole probability surface is a stack of parallel hyperplanes with the sigmoid ramping across them. This is the geometric content of "linear model": all the structure the model can express is one direction plus one offset. ## What this does and does not buy you The consequence to state in an interview is a limitation: whatever the data looks like, the model's answer in the original feature space is one flat cut. If the classes are not arranged so that one flat cut separates them, no amount of data or training time fixes it — the boundary shape is a property of the model class, not of the fit. The usual escape is to hand the model new features (products, powers, buckets) so that a hyperplane in the enlarged space is a curve back in the original one; the model is still linear, in the new coordinates. The flatness is also what makes the model cheap and stable: `d + 1` numbers describe the whole rule, scoring is one dot product, and the direction of every feature's push on the prediction is fixed and inspectable.
- If you multiply every weight and the intercept by ten, what changes about the model's predictions?The boundary does not move at all, because the set where the score is zero is unchanged by a positive rescaling. What changes is confidence: scores are ten times further from zero, so probabilities are pushed towards 0 and 1 and the transition across the boundary becomes almost a step. Same classifications under the 0.5 rule, far more extreme probabilities, and much worse calibration if the scaling was not learned from data.
- Where does a point that sits exactly on the boundary land, and does that matter in practice?Its score is zero, so the sigmoid returns exactly 0.5 and the rule is a tie. In practice it almost never matters with continuous features — the tie set is a measure-zero slice — but with coarse, binary or heavily rounded features many rows can land on it at once, so the implementation's tie-breaking rule becomes a real, testable choice rather than a curiosity.
- How does the boundary change if you add a third feature to a two-feature model?The boundary stops being a line in the plane and becomes a plane in three-dimensional space, still flat and still described by one equation, score equals zero. You can no longer draw it on the scatter, which is why people fall back on projecting onto two features at a time — a projection can make a genuinely separating plane look tangled, so read those plots carefully.
The score is an altitude reading and the sigmoid is just the colour scale printed on the map. Recolouring the map does not move the coastline; sea level is still one flat contour.
saying these in an interview costs you the question
- Says the sigmoid's S-shape makes the boundary curved
- Confuses the probability-versus-score curve with the boundary in feature space
- Thinks each feature gets its own separate cutoff value
- Claims logistic regression can fit any boundary shape given enough data
- Believes scaling all the weights up moves the boundary