What happens to a closed-form least-squares fit when two features are nearly identical?
answer
- the matrix is nearly non-invertible
- the data constrains the sum, not the split
- huge offsetting coefficients, unstable across refits
- variance, not bias
- predictions survive; interpretation does not
basics
~20 sNear-duplicate columns leave the cross-product matrix almost singular, so the two coefficients come out enormous, opposite in sign and wildly unstable across refits. Predictions inside the training range stay fine; the individual coefficients are not interpretable.
solid answer
~50 sOn a turbine efficiency regression with 40 sensor channels, two thermocouples on the same shaft read within 0.2 degrees of each other, so `X'X` is near-singular. The data pins down the pair's combined effect but says almost nothing about how to split it, and the exact solve answers anyway: you get something like `+840` and `-837` where every other channel is single-digit, and the pair swings by orders of magnitude if you bootstrap or drop 1% of rows. Standard errors blow up because the coefficient covariance is proportional to `(X'X)^-1`. Crucially the estimates are not biased and in-range predictions are stable — the offsetting weights cancel. What breaks is interpretation and extrapolation. I would confirm it with the condition number of the standardised design or variance inflation factors, then average the two thermocouples into one shaft-temperature channel.
code
python · 20 linesx1 = [1.0, 2.0, 3.0, 4.0, 5.0]
x2 = [1.0001, 2.0002, 2.9998, 4.0001, 5.0002] # a near-duplicate sensor
def solve(y): # 2x2 normal equations, no intercept
a = sum(u * u for u in x1)
b = sum(u * v for u, v in zip(x1, x2))
c = sum(v * v for v in x2)
d = sum(u * t for u, t in zip(x1, y))
e = sum(v * t for v, t in zip(x2, y))
det = a * c - b * b # nearly zero: the design is ill-conditioned
return (c * d - b * e) / det, (a * e - b * d) / det
y1 = [3.0, 6.0, 9.0, 12.0, 15.0]
y2 = [3.01, 5.99, 9.0, 12.0, 15.0] # two targets nudged by 0.01
w = solve(y1)
v = solve(y2)
print("fit A:", round(w[0], 1), round(w[1], 1), "sum", round(w[0] + w[1], 3))
print("fit B:", round(v[0], 1), round(v[1], 1), "sum", round(v[0] + v[1], 3))
# fit A: 3.0 0.0 sum 3.0
# fit B: 10.0 -7.0 sum 3.0go deeper
Recall that two features carrying nearly the same information make the fitted coefficients unreliable, and that the standard first check is the correlation matrix for pairs above about 0.99.
Explain the mechanism: near-duplicate columns make the cross-product matrix nearly singular, its inverse huge, and the coefficient variance enormous. Be able to say why the sum of the pair is well determined while the split is not.
Demonstrate the diagnosis and the judgment: condition number on standardised columns or variance inflation factors, a bootstrap stability check, and a clear statement that in-range prediction survives while interpretation and extrapolation do not. Then name the remedy you chose and why.
Own the downstream consequence: coefficient tables from collinear designs get read as causal drivers in decision meetings. Be ready to argue for a feature-intake discipline — deduplicating redundant sensor channels at the source — rather than repairing it per model.
## What "nearly identical" does to the arithmetic The exact least-squares weights come from solving `(X'X) w = X'y`, where `X'X` is the matrix of cross-products between feature columns. If two columns are almost the same vector, `X'X` is almost **singular**: its smallest eigenvalue is close to zero, its determinant is close to zero, and the system it defines is close to having no unique solution. Concretely, on a wind-turbine efficiency regression with 40 sensor channels, two thermocouples mounted on the same shaft read within 0.2 degrees of each other all day. Call the columns `t1` and `t2`. Because `t2` is essentially `t1` plus a whisker of noise, the data can only tell you about the *combined* effect `w1 + w2`. It contains almost no information about how that total should be split between the two. The least-squares machinery still returns an answer — but the answer to a question the data barely constrains. ## The symptoms **Enormous, offsetting coefficients.** You get pairs like `+840` on one thermocouple and `-837` on the other, when every other channel carries a coefficient in single digits. Their sum is sensible; the individuals are not. **Wild instability.** The classic diagnostic: refit on a bootstrap resample, or drop 1% of rows, and the pair swings by orders of magnitude and can swap signs. The attached snippet shows the minimal version — two columns matching to four decimal places, and a 0.01 nudge in two target values moves the coefficients from `(3.0, 0.0)` to `(10.0, -7.0)` while their sum stays pinned at exactly `3.0`. **Inflated standard errors.** The coefficient covariance is proportional to `(X'X)^-1`. A near-zero eigenvalue makes entries of that inverse huge, so the standard errors of the affected coefficients blow up and their significance tests turn insignificant — even though the pair jointly matters a lot. ## The symptom you will *not* see Collinearity does **not** bias the coefficients: the estimator is still centred on the truth, it is just extremely high-variance in the collinear directions. And it does **not** wreck predictions inside the region of the data. The direction the model is confused about — how much of the effect to attribute to `t1` versus `t2` — is a direction in which new data barely varies either, so the offsetting weights cancel and the fitted values are stable. Predictions degrade only when you **extrapolate**: feed the model a row where the two thermocouples disagree by 5 degrees (one sensor drifts, one is replaced) and the `+840 / -837` pair produces a wildly wrong number. That split is the heart of the senior answer: **collinearity is an interpretation and robustness problem, not usually an accuracy problem.** ## How to detect it before it bites - **Condition number of the design.** The ratio of the largest to the smallest singular value of `X`. Standardise the columns first, otherwise a column measured in watts and one in megawatts produce a huge condition number that has nothing to do with genuine redundancy. Large values (rules of thumb start around 30) flag ill-conditioning. - **Variance inflation factors.** Regress each feature on all the others; `VIF_j = 1 / (1 - R_j^2)`. A near-perfect fit for feature `j` from the rest means that column adds almost no independent information. - **A stability check.** Bootstrap the fit and look at the spread of each coefficient. This costs nothing to run and is the most convincing evidence to show a stakeholder. - **Just look at the correlation matrix** for pairs above 0.99. It misses collinearity spread across three or more columns, which the condition number catches, but it finds the duplicated-sensor case immediately. ## The exactly-singular cases Sometimes `X'X` is not merely near-singular but genuinely non-invertible, and the exact solve has no unique answer at all. **The dummy trap.** One-hot encode a categorical with `K` levels into `K` columns *and* keep an intercept: the `K` columns sum to the all-ones intercept column, an exact linear dependence. Drop one level, or drop the intercept. **More features than rows.** A gene-expression design with 20,000 probes and 300 samples has rank at most 300, so the 20,000 x 20,000 cross-product matrix is singular by construction. There are infinitely many weight vectors that drive the training residuals to exactly zero. A pseudoinverse will hand you one of them — the minimum-norm solution — but it is a convention, not information the data provided. In this regime unpenalised least squares is simply not the right tool; you need shrinkage or feature selection to make the problem well-posed, and that decision is the modelling step, not a numerical detail. ## What to do about the turbine pair Within the plain least-squares setting the options are: **drop one** of the two thermocouples; **combine them** into a single averaged shaft-temperature channel, which is what the physics says they measure and which removes the redundancy honestly; or **accept it** and do not interpret those two coefficients, if you only need predictions in the operating range you trained on. Choose deliberately and write down which you chose — a coefficient table shipped without that note is how a `+840` ends up in a slide as "temperature is the dominant driver".
- Does collinearity bias the coefficients or just inflate their variance?Variance only. Ordinary least squares stays unbiased under collinearity — the estimator is still centred on the true coefficients. What collapses is precision: the near-zero eigenvalue of the cross-product matrix makes entries of its inverse enormous, so the coefficient covariance and hence the standard errors explode in exactly the directions the columns are redundant in. That is why the pair can be jointly significant while neither one is individually significant.
- The design has 20,000 gene-expression probes and only 300 samples. What does the exact solve give you?Nothing usable. The design has rank at most 300, so the 20,000 by 20,000 cross-product matrix is singular by construction and there is no unique solution — infinitely many weight vectors drive the training residuals to exactly zero. A pseudoinverse returns the minimum-norm one, but that is a convention, not information the data supplied. In this regime plain least squares is the wrong tool; the problem needs shrinkage or feature selection to be well-posed at all.
- Your only goal is prediction accuracy in the normal operating range. Do you still have to fix the collinearity?Not necessarily. The direction the model is confused about is a direction new data barely varies in either, so the offsetting weights cancel and in-range predictions are stable. I would still write down that the coefficients are uninterpretable, so nobody reads +840 off a table as a dominant driver, and I would add a guard for the extrapolation case: if one thermocouple drifts or is replaced and the two channels start disagreeing by degrees, that huge coefficient pair produces a badly wrong prediction.
Two people push a stalled car together. You can measure the total force on the car precisely, but the car's motion cannot tell you who pushed how hard, and the tiniest measurement wobble reassigns the whole effort from one to the other.
saying these in an interview costs you the question
- Claims collinearity biases the coefficients toward zero
- Says predictions become unreliable at every input
- Reads a huge coefficient as evidence of a strong effect
- Thinks more features than rows just makes the fit slow
- Keeps all one-hot levels alongside an intercept
- Checks the condition number without standardising the columns first