What is the leverage (hat value) of an observation in linear regression, and what determines it?
answer
- no y appears anywhere in it
- distance from the predictor centroid
- diagonal of the hat matrix
- the values sum to the coefficient count
- average is p/n, flag twice that
basics
~20 sLeverage (the hat value h_ii) measures how unusual a row's predictor values are compared with the rest of the data. It depends only on X, never on the response y, and bounds how far that row can pull its own fitted value.
solid answer
~50 sIn OLS the fitted values are `yhat = H y`, where `H = X(X'X)^-1 X'` is the hat matrix. The leverage of observation `i` is the diagonal element `h_ii`, which is exactly how much `yhat_i` moves when `y_i` moves by one unit. Because `H` is built from the design matrix alone, leverage is a property of the predictors only — the response never enters it. With one predictor, `h_ii = 1/n + (x_i - xbar)^2 / sum_j (x_j - xbar)^2`, so leverage grows with distance from the centre of the x-values. The hat values sum to the number of estimated coefficients `p`, so the average is `p/n`, and a common flag is `h_ii > 2p/n`. High leverage is potential influence, not realised influence: it says the row *could* move the line, not that it has.
go deeper
Be ready to say in one sentence that leverage measures how unusual a row's predictor values are, and that it has nothing to do with how far the point is from the line.
Explain the mechanics: yhat = Hy, leverage is the diagonal of the hat matrix, hat values sum to the number of coefficients, so the average is p/n and 2p/n is the usual flag.
Show you use leverage as a screen, not a verdict. Talk about extrapolation risk, designed extremes that are legitimately high-leverage, and why you always pair hat values with residuals before touching any row.
Own the framing that leverage is a property of the design. If your sampling or feature set repeatedly puts single rows in near-unit leverage, argue about fixing the design or the predictor parameterisation rather than policing individual observations.
## The hat matrix Fit an ordinary least squares regression with `n` observations and a design matrix `X` that has `p` columns — `p = k + 1` when you fit an intercept plus `k` predictors. The coefficient estimates are `b = (X'X)^-1 X' y`, so the fitted values are ``` yhat = X b = X (X'X)^-1 X' y = H y ``` The matrix `H = X (X'X)^-1 X'` is called the **hat matrix**, because it is the operator that puts the hat on `y`. It is the projection onto the column space of `X`: it takes the observed responses and squashes them onto the plane the model is allowed to reach. **Leverage** is the `i`-th diagonal element of that matrix, written `h_ii`. Reading the definition row by row, `yhat_i = sum_j h_ij y_j`, so `h_ii` is the weight observation `i` places on *itself* when its own fitted value is computed. Equivalently `h_ii = d(yhat_i)/d(y_i)`: raise this row's response by one unit, holding everything else fixed, and its fitted value rises by `h_ii`. ## Leverage depends on X only The single most important structural fact is that `H` is assembled entirely from `X`. No response value appears anywhere in it. That means leverage can be computed before you have seen a single `y`, and two datasets with identical predictors but wildly different responses have identical leverages. A candidate who says "that point has high leverage because it is far above the line" has confused leverage with the residual. With a single predictor the formula becomes concrete: ``` h_ii = 1/n + (x_i - xbar)^2 / sum_j (x_j - xbar)^2 ``` So leverage is a rescaled squared distance from the centre of the predictor values, floored at `1/n`. With several predictors the same reading holds in higher dimensions: `h_ii` is a monotone function of the Mahalanobis distance of that row's predictor vector from the predictor mean, measured in the metric of the predictors' own covariance. This is why a row can have high leverage while every individual predictor value looks ordinary — it is the *combination* that is unusual, for example a very short person with a very large shoe size in a fit that uses both. ## Arithmetic worth memorising - The hat values sum to the trace of `H`, which equals `p`, the number of estimated coefficients. So the **average leverage is `p/n`**. - With an intercept in the model, `1/n <= h_ii <= 1`. - A common rule of thumb flags `h_ii > 2p/n` as high leverage; some texts use `3p/n` on small samples. These are screening heuristics, not tests — nothing is significant at `2p/n`. - `Var(yhat_i) = sigma^2 h_ii` and `Var(e_i) = sigma^2 (1 - h_ii)`, where `e_i` is the raw residual. High-leverage rows therefore have systematically *small* residual variance: the line is dragged toward them, so their raw residual understates how odd they are. That is the whole motivation for rescaling residuals before you judge them. - The extreme case `h_ii = 1` forces `e_i = 0`. It happens, for instance, when a dummy variable is one for exactly that row and zero everywhere else: the model spends a whole coefficient fitting that observation perfectly. ## Leverage is potential, not verdict Because leverage ignores `y`, it cannot tell you that a point *has* distorted the fit. Consider a wealth-on-education regression where one row sits far out at the top of the education range. If its wealth lands exactly on the line the other rows imply, deleting it barely moves the slope — high leverage, near-zero influence. If instead its wealth is wildly off that line, the same leverage now converts a large discrepancy into a large swing in the coefficients. So the practical workflow is: leverage tells you which rows are in a position to matter, the residual tells you which rows disagree with the model, and an influence measure combines the two into how much the fit actually changed. High leverage on its own is a reason to look, never a reason to delete. ## Common traps Do not treat leverage as an error flag — a designed experiment deliberately puts observations at the extremes of the predictor range precisely because those points estimate the slope most efficiently. Do not compare hat values across models with different predictor sets; adding a column changes `p`, changes `H`, and changes every hat value. And remember that leverage is defined relative to the rows you fitted: a new observation far outside that predictor cloud is an extrapolation, which is a related but separate worry.
- Can an observation have high leverage yet almost no effect on the fitted coefficients?Yes, and it is the standard counterexample. A point far out in the predictor range whose response lands exactly on the line the other data imply has large `h_ii` but a near-zero residual, so removing it barely moves the slope. Leverage describes the position from which a row could pull; influence describes whether it actually pulled.
- What is the range of hat values, and what does h_ii = 1 mean?With an intercept in the model, `1/n <= h_ii <= 1`. A value of exactly 1 means the fit passes through that observation perfectly, forcing its residual to zero — typically because some predictor combination, such as an indicator that is on for that row alone, is dedicated to it. That row contributes nothing to estimating error variance.
- Why do high-leverage rows tend to have unusually small raw residuals?Because `Var(e_i) = sigma^2 (1 - h_ii)`. As `h_ii` approaches one, the residual's own variance shrinks toward zero — the fitted surface is pulled toward the point. Judging such a row by its raw residual therefore understates how discrepant it is, which is why residuals are divided by `s * sqrt(1 - h_ii)` before being compared.
Leverage is the length of the lever arm on a see-saw. A point far out along the x-axis sits at the very end, so a small change in its height swings the whole plank; whether it actually swings anything depends on where its height happens to be.
saying these in an interview costs you the question
- Says leverage depends on how far y sits from the line
- Uses leverage and residual size interchangeably
- Treats every high-leverage row as bad data to delete
- Thinks leverage exists only in simple one-predictor regression
- Cannot say that hat values average p/n