Why does a linear probability model fitted with OLS fail on a binary outcome?
answer
- look at the predicted values first
- the outcome takes only two values
- variance depends on the mean here
- predictions can leave the unit interval
- a straight line has no ceiling
basics
~20 sA straight line on a 0/1 outcome has no ceiling or floor, so it predicts impossible values like -0.12 and 1.30. The outcome's variance p(1-p) also changes with the predictors, so the errors are heteroscedastic by construction.
solid answer
~50 sThree problems, in order of severity. First, the fitted line is unbounded, so a churn model can hand back -0.12 for one account and 1.30 for another — values that are not probabilities and cannot be fed into any expected-value calculation. Second, for a binary outcome the conditional variance is `p(1-p)`, which depends on the predictors, so the errors are heteroscedastic by construction and the default standard errors are wrong. Third, the constant slope is implausible near the boundaries: the same one-unit change that moves probability from 0.45 to 0.55 cannot also move it from 0.95 to 1.05. The logit link fixes all three by fitting an S-curve that asymptotes at 0 and 1. That said, the linear model is not worthless — its coefficients read directly in percentage points, which is why it survives as a quick approximation paired with robust standard errors.
go deeper
Recall the headline objection: a straight line on a 0/1 outcome can predict below 0 and above 1, and those numbers are not probabilities. Give a concrete impossible fitted value.
Explain the variance argument as well as the bounds one: a Bernoulli outcome has variance p(1-p), which moves with the predictors, so the errors are heteroscedastic by construction rather than by accident.
Show you would inspect how many fitted values actually fall outside the unit interval before condemning the model, and know that robust standard errors repair inference but not functional form.
Own the call on when a fast, directly interpretable linear approximation is the right tool for a decision and when the bounded model is required, and make sure the caveats travel with the numbers.
## What the linear probability model is Fit ordinary least squares with a 0/1 response and read the fitted value as a probability: ``` Y = b0 + b1*x1 + b2*x2 + error, Y in {0, 1} E[Y | X] = P(Y = 1 | X) ``` Because the expected value of a 0/1 variable *is* the probability that it equals 1, this is not nonsense on its face — the regression really is estimating a conditional probability, and `b1` reads as "a one-unit increase in x1 raises the probability of the event by b1", straight in percentage points. The problems are with the functional form and the error structure. ## Problem 1: unbounded predictions A plane has no ceiling. Fit a churn model on tenure and monthly spend and you will find real rows with fitted values like -0.12 and 1.30. These are not merely inelegant: - They cannot be used in an expected-value calculation. A -0.12 probability of churn multiplied by a customer's lifetime value produces a negative expected loss. - Clipping to [0, 1] afterwards is a patch, not a fix. It leaves the coefficients that produced the impossible values untouched and creates a discontinuity in the fitted function. - The failure concentrates precisely where the decisions matter: at the extreme, highest-risk and lowest-risk accounts. How often this bites depends on the data. If every fitted value sits between 0.2 and 0.8, the linear approximation is nearly harmless; if the base rate is 2% or the predictors are strong, it is everywhere. ## Problem 2: heteroscedasticity by construction For a Bernoulli outcome with success probability `p`, the conditional variance is ``` Var(Y | X) = p * (1 - p) ``` Since `p` varies with the predictors, so does the error variance — it is maximised at `p = 0.5` (variance 0.25) and shrinks toward 0 at both extremes. Homoscedasticity is a core assumption behind the usual OLS standard-error formula, and here it is violated *by the definition of the outcome*, not by accident of the data. The consequence is specific and worth stating precisely: the coefficient estimates remain unbiased for the linear approximation, but the default standard errors, t-statistics and p-values are wrong. Heteroscedasticity-robust standard errors repair the inference without repairing the functional form. A related, more cosmetic point: the residuals can only take two values at any given `X` (namely `1 - fitted` and `-fitted`), so they are nowhere near normal. This matters far less than it is usually claimed to — OLS does not require a normal outcome — so a candidate who names non-normality as the *main* defect has the priorities backwards. ## Problem 3: a constant slope is the wrong shape The linear model asserts that a one-unit change in a predictor moves the probability by the same amount everywhere. Substantively that is rarely credible. Moving an account from a 45% to a 55% churn risk is one thing; the same nominal change applied to an account already at 95% would take it to 105%. Real effects taper as the outcome approaches certainty in either direction. The logistic form encodes exactly that taper: its marginal effect is `b * p * (1 - p)`, largest when `p` is near 0.5 and vanishing at the extremes. ## What the logit link changes Modelling `log(p / (1 - p))` as linear makes the fitted probability `1 / (1 + exp(-(b0 + b1*x1 + ...)))`. That curve is bounded in (0, 1) with asymptotes it never crosses, it saturates smoothly, and the model is fitted by maximum likelihood using the Bernoulli distribution, so the variance structure is part of the model rather than an assumption it violates. The price is interpretability: coefficients now speak in log-odds and need exponentiating into odds ratios or converting into marginal effects before anyone outside the modelling team can act on them. ## Why practitioners still fit the linear model anyway Honesty here impresses interviewers. In applied economics the linear probability model is common and defensible, for concrete reasons: - **Direct interpretation.** The coefficient is already an average partial effect in percentage points, with no conversion step. - **Close agreement in the middle.** When fitted probabilities stay roughly within 0.2 to 0.8, its coefficients approximate the logistic model's average marginal effects well. - **Fixed effects and instruments.** Linear estimators handle high-dimensional fixed effects and instrumental-variable designs cleanly, where the nonlinear analogues are harder. The defensible version always comes with robust standard errors, a check on how many fitted values fall outside [0, 1], and an acknowledgment that the model is an approximation to a conditional probability rather than a description of one. The indefensible version is fitting OLS to a 0/1 column, reading the default p-values, and reporting a negative probability with a straight face.
- If the linear probability model is so flawed, why do practitioners still use it?Its coefficients are already average partial effects in percentage points, needing no conversion, and when fitted probabilities stay in the middle range they track the logistic model's average marginal effects closely. It also handles high-dimensional fixed effects and instrumental-variable designs cleanly. The defensible version always pairs it with heteroscedasticity-robust standard errors.
- What does the logit link change about the shape of the fitted curve?It replaces the straight line with an S-curve that asymptotes at 0 and 1 without crossing them. The effect of a one-unit change is largest near a fitted probability of 0.5 and compressed near either extreme, since the marginal effect is b times p times (1-p). That taper is what a constant slope cannot express.
- Does heteroscedasticity bias the coefficients in a linear probability model?No. It leaves the coefficient estimates unbiased for the best linear approximation to the conditional probability; what it breaks is the standard errors, and therefore the t-statistics, p-values and confidence intervals computed from them. Robust standard errors fix the inference while leaving the functional-form problem untouched.
saying these in an interview costs you the question
- Says clipping predictions to zero and one fixes the model
- Names non-normal residuals as the main defect
- Claims heteroscedasticity here biases the coefficients
- Cannot say why a fitted value of 1.30 matters
- Treats it as a sample-size problem that more data solves