Two feature columns correlate at 0.999 - is that matrix rank deficient?
answer
- 0.999 is not 1
- rank is exact, correlation is graded
- independent on paper, fragile in arithmetic
- distance to the nearest dependent matrix
- derived columns versus measured ones
basics
~20 sNo. A correlation of 0.999 is not an exact linear identity, so the two columns stay independent and the matrix is full rank. It sits a hair from rank deficiency, so arithmetic on it behaves almost as badly.
solid answer
~50 sRank is an exact algebraic property: columns are dependent only if some non-trivial weighted combination is exactly the zero vector, which for two columns means one is an exact affine function of the other - correlation exactly +1 or -1. At 0.999 no such identity holds, the columns are independent, and the matrix is full rank on paper. The catch is that rank is all-or-nothing while nearness to dependence is a matter of degree. Near-collinear columns span a genuinely two-dimensional space that is extremely thin in one direction: small changes in the data, or the last bits of floating-point arithmetic, swing the weights that reconstruct one column from the other. So exact rank is the wrong instrument for measured data; the useful question is how far the matrix sits from a lower-rank one. Two exactly proportional columns are the limit case, where independence disappears outright.
go deeper
Know that high correlation and exact linear dependence are different claims, and that only an exact identity between columns lowers the rank.
Explain why rank is all-or-nothing while near-collinearity is a matter of degree, and describe geometrically what an almost-flat two-dimensional span looks like.
Show the judgment to trace a near-duplicate column back to its origin - a derived definition versus two genuinely similar measurements - and to say which problem you actually have.
Set the team standard for how redundancy among features is measured and reported, instead of letting each project invent its own correlation threshold.
## Two different claims **Exact linear dependence** is an algebraic statement: there exist weights, not all zero, with c1*a1 + c2*a2 = 0. For two columns this means one is a scalar multiple of the other. It is a yes/no property with no middle ground, and it is exactly what makes the rank drop. **Correlation** is a statistical summary of two centred columns, taking any value in [-1, 1]. It equals +1 or -1 precisely when one column is an exact affine function of the other, that is, when the centred columns are exactly proportional. Anything strictly inside that range - 0.999 included - means no exact identity holds. So the direct answer is: at 0.999 the columns are linearly independent, and a matrix with those two columns has rank 2, full column rank. Nothing about the algebra is degenerate. ## Why that answer is not the end of it Rank is a **discontinuous** function of a matrix's entries. Perturb one entry of an exactly rank-deficient matrix by 10^-12 and the rank jumps from 2 to 3, even though the matrix is, in every meaningful sense, the same matrix. Conversely a matrix at correlation 0.999 is full rank, yet it is a whisker away from one that is not. A property that flips on an invisible change is a poor instrument for describing measured data, where every entry carries noise anyway. The honest framing is **distance to rank deficiency**: how large a perturbation would it take to make the columns exactly dependent? For exactly proportional columns the distance is zero. At 0.999 the distance is tiny. At correlation 0.2 it is large. That is a continuous quantity with a continuous meaning, which is what real data needs. ## The geometry Two independent columns span a plane. When the columns are nearly proportional, that plane is real but the data barely explores one of its directions: almost all of the variation runs along a single line, with a sliver of spread perpendicular to it. The span is two-dimensional in the algebraic sense while being effectively one-dimensional in the practical sense. Because the second direction is so thinly supported, the weights needed to express any given vector in terms of the two columns become enormous and delicately balanced - a large positive weight on one column nearly cancelled by a large negative weight on the other. Nudge the data and those weights move dramatically, even though the vector they reconstruct barely changes. That sensitivity, not the rank, is what actually causes trouble. ## Exact versus near, in real datasets A useful discriminator is *where the relationship comes from*. - **Exact dependence** almost always comes from a **definition**: a column computed as the sum, difference, ratio or unit conversion of others; a duplicated column under two names; a quantity recorded once in two systems and copied. Exact identities are manufactured, not measured. - **Near-collinearity** almost always comes from **measurement**: two sensors observing the same physical thing, two questionnaire items asking nearly the same question, two metrics driven by the same underlying quantity. Noise makes an exact identity essentially impossible. So when you see a correlation of 0.999, the first move is not to pick a threshold - it is to read the definition of the two columns in the pipeline. If one is derived from the other, you have exact dependence hiding behind rounding, and the fix is structural. If they are two genuine measurements, you have near-collinearity, and the question becomes how much independent information the second one really contributes. ## Floating point blurs the line further In finite precision, even an exactly dependent pair of columns rarely produces an exact zero when you test the identity: rounding leaves residue at the level of machine precision. Conversely a near-dependent pair produces residue barely larger than that. Any procedure that decides rank by computing whether something equals zero is therefore making a threshold judgement, whether or not it admits to it. This is why 'what is the rank of this data matrix?' is not really a well-posed question for measured data; 'how close is it to rank deficient, and does that closeness matter for what I am computing?' is. ## What a strong answer sounds like State the algebra first - 0.999 is not 1, so the columns are independent and the matrix is full rank. Then immediately explain why the algebra is not the point: rank is a knife-edge property, near-dependence is continuous, and the practical consequence is extreme sensitivity of any quantity derived from reconstructing one column from the others. Finish by distinguishing derived columns (exact dependence, a definitional problem) from co-varying measurements (near-collinearity, a data problem). Candidates who only say 'yes, it is basically singular' have skipped the distinction the question is testing.
- If the matrix is technically full rank, why do people still treat near-collinear columns as a problem?Because the useful question is not whether an exact identity exists but how much independent information the second column adds. When the answer is almost none, any quantity obtained by reconstructing one column from the others becomes extremely sensitive: a small perturbation in the data produces a large swing in the numbers reported. The algebra says full rank; the arithmetic says barely.
- What would convince you two columns are exactly dependent rather than merely near-collinear?An algebraic reason drawn from how the data was built - a column defined as a sum, difference or rescaling of others - rather than a number close to one. Measured quantities essentially never produce exact identities; derived quantities always do. Reading the column's definition in the pipeline settles it faster and more reliably than any numeric threshold.
- Does centring or rescaling the columns change the rank of a matrix?Rescaling a column by a non-zero constant never changes the rank, since it does not change which combinations vanish. Centring can change it: subtracting column means is itself a linear operation, and it can create a dependence that was not there, for instance when a constant column is present and becomes all zeros after centring.
saying these in an interview costs you the question
- Says correlation 0.999 makes the matrix singular
- Treats rank as a continuous measure of redundancy
- Thinks full rank means the matrix is well behaved
- Uses a fixed correlation cutoff as proof of exact dependence