How does iteratively reweighted least squares fit a logistic regression model?
answer
- no closed form, so iterate
- Newton's method wearing a different name
- weights from the current fitted probabilities
- p times one minus p peaks at a half
- separation makes coefficients run away
basics
~10 sIRLS runs Newton's method on the log-likelihood. Each iteration computes fitted probabilities, forms weights p(1-p), and solves a weighted least-squares problem on a working response, repeating until the coefficients stop moving.
solid answer
~50 sThe logistic log-likelihood has no closed-form maximiser, so it is fitted iteratively. Its gradient is `X^T (y - p)` and its Hessian is `-X^T W X` with `W` diagonal holding `p_i(1 - p_i)` from the current fitted probabilities. Plugging those into the Newton update gives `beta_new = beta + (X^T W X)^-1 X^T (y - p)`, which is algebraically identical to a weighted least-squares fit of a working response `z = X beta + W^-1 (y - p)` with weights `W` — hence "iteratively reweighted least squares". Observations with fitted probabilities near 0.5 carry the most weight, since that is where the model is most sensitive; observations the model is already confident about are nearly ignored. The log-likelihood is concave, so with non-separable data there is a single maximum and convergence is usually a handful of iterations. Perfect separation is the classic failure: weights collapse toward zero, coefficients run off toward infinity, and the fit never settles without a penalty term.
go deeper
Recall that a logistic fit has no closed-form solution and is found by iterating to convergence, unlike a linear least-squares fit which is solved directly.
Explain the loop: fitted probabilities give weights p(1-p), those weights define a weighted least-squares problem, and repeating it is the same as taking Newton steps.
Demonstrate diagnosis: recognise separation from exploding coefficients and vanishing weights, spot ill-conditioning from collinear predictors, and know why loosening the tolerance is the wrong fix.
Own the decision of how to handle data that does not identify a finite fit — penalisation, feature consolidation, or declaring the effect unestimable — and what each choice costs downstream.
## Why an iterative fit at all A logistic model gives each observation a fitted probability `p_i` that depends on the coefficient vector `beta` through the linear predictor `eta_i = x_i^T beta`. The log-likelihood is ``` l(beta) = sum_i [ y_i * eta_i - log(1 + exp(eta_i)) ] ``` Setting its derivative to zero produces equations that are nonlinear in `beta` — unlike ordinary least squares, there is no formula that hands you the answer. So the fit is an optimization: maximise a smooth concave function of p parameters, iteratively. ## The derivatives Differentiating gives the score vector ``` grad l = X^T (y - p) ``` where `y` is the vector of 0/1 outcomes, `p` the vector of current fitted probabilities, and `X` the n-by-p design matrix. The interpretation is clean: each column of X is weighted by how far the outcomes sit from what the model currently predicts. Differentiating again gives ``` Hessian l = -X^T W X, W = diag(p_i (1 - p_i)) ``` Two facts about this Hessian matter. First, it depends only on the fitted probabilities and not on the observed outcomes `y` at all — a special property of this link function. Second, `X^T W X` is positive semidefinite for any non-negative weights, so the log-likelihood is concave: no local maxima to get trapped in, and a unique maximiser whenever `X` has full column rank and the data are not separable. ## The Newton step, rewritten Newton's method for maximising subtracts the inverse Hessian times the gradient with the appropriate sign: ``` beta_new = beta + (X^T W X)^-1 X^T (y - p) ``` Now do a small algebraic rearrangement. Define the **working response** ``` z = X beta + W^-1 (y - p) ``` that is, the current linear predictor plus a residual scaled by the inverse weight. Substituting shows ``` beta_new = (X^T W X)^-1 X^T W z ``` which is exactly the formula for a weighted least-squares regression of `z` on `X` with weights `W`. So every Newton step is a weighted linear regression. Recompute `p`, recompute `W` and `z`, refit. That is the entire algorithm, and it is where the name comes from: iteratively **reweighted** least squares. ## Reading the weights The weight for observation i is `p_i (1 - p_i)`. This is maximised at `p_i = 0.5`, where it equals 0.25, and falls toward zero as the fitted probability approaches 0 or 1. Observations the model is uncertain about therefore dominate the update, while ones it is already confident about barely influence it. This is a genuine property of the curvature: the log-likelihood surface is most sensitive to coefficient changes for observations sitting near the decision boundary. ## Convergence and its failures Because the objective is concave and the Newton step is exact for the local quadratic model, IRLS typically converges in fewer than ten iterations from a zero starting vector, and the convergence is quadratic near the optimum. Standard stopping rules watch the change in the log-likelihood or the relative change in coefficients. The well-known failure is **complete or quasi-complete separation**: some linear combination of the predictors perfectly (or almost perfectly) splits the two classes. The likelihood then has no finite maximiser — it keeps increasing as the coefficients grow — so the fitted probabilities are driven to 0 and 1, the weights `p(1-p)` collapse toward zero, `X^T W X` becomes numerically singular, and the coefficients diverge while the algorithm reports huge, meaningless estimates. Diagnosing it means looking at coefficient magnitudes and fitted probabilities rather than trusting a convergence flag. The standard remedies are to add a penalty that keeps coefficients finite, to drop or combine the offending predictor, or to accept that the data do not identify the effect. A second, milder failure is **collinearity**: if `X` has near-dependent columns, `X^T W X` is ill-conditioned and coefficient estimates are unstable across refits even when the algorithm converges. ## What the question is testing An interviewer asking this wants to see that you know a logistic fit is an optimization rather than a formula, that you can identify the algorithm as Newton's method in disguise, and that you can explain the reweighting in terms of curvature. Naming the separation failure and its symptom — exploding coefficients with vanishing weights — is what separates a candidate who has fitted these models in anger from one who has only read about them.
- Why does an observation with a fitted probability near 0.5 get the largest weight?The weight is `p(1 - p)`, which peaks at 0.25 when p is 0.5 and shrinks toward zero as p approaches either extreme. It measures how much the fitted probability responds to a change in the linear predictor, so uncertain observations near the boundary carry the most information about where the coefficients should move, while confidently classified ones barely shift the update.
- What happens to the fit when a predictor perfectly separates the two classes?The likelihood has no finite maximiser, so coefficients grow without bound. Fitted probabilities are pushed to 0 and 1, the weights `p(1-p)` collapse, the weighted design matrix becomes numerically singular, and the algorithm either stops on an iteration cap or reports enormous estimates. Adding a penalty that keeps coefficients finite, or removing the offending predictor, is the usual response.
- How many iterations would you expect a well-behaved logistic fit to take, and what does an unusually high count suggest?Typically under ten, because the objective is concave and each step is a full Newton step with quadratic local convergence. A run that grinds on for dozens of iterations usually signals near-separation, severe collinearity making the weighted design matrix ill-conditioned, or extreme predictor scaling. Inspect coefficient magnitudes and fitted probabilities before touching the tolerance.
saying these in an interview costs you the question
- Claims logistic regression has a closed-form solution
- Says the weights are constant across iterations
- Cannot connect the reweighting to Newton's method
- Treats non-convergence as a tolerance setting to loosen
- Believes the log-likelihood has multiple local maxima