Why can a correlation or mutual-information filter keep two 0.97-correlated columns yet drop a feature that only matters in combination?
answer
- one column at a time against the target
- marginal cannot answer conditional
- twins each score well separately
- agreement of two flags, zero alone
- mutual information fixes curves, not context
basics
~20 sA correlation or mutual-information filter scores each column against the target on its own. Two near-duplicates both score well and survive; a feature that matters only alongside another scores near zero alone and is cut. The blind spot is univariate scoring.
solid answer
~50 sA univariate filter asks one question per column: how strongly does this column alone relate to the target. That question has no way to notice that two columns are saying the same thing. List price and net price at `0.97` correlation both score highly, both get kept, and you have paid for two columns' worth of pipeline to buy one column's worth of information. The same blind spot runs the other way: if the target depends on the agreement of two flags — high when both are on or both off, low when they disagree — then each flag alone is uninformative, its mutual information with the target is essentially zero, and the filter drops the pair before any model could have used it. Swapping correlation for mutual information fixes nonlinearity, not univariateness. If interactions or redundancy matter, you need a redundancy-aware filter or a family that evaluates columns jointly.
go deeper
Know that these filters score each column against the target separately, and that separately means they cannot tell you a column is redundant or that two columns only matter together.
Explain both failure directions from the single cause, and be precise that mutual information generalises correlation to nonlinear dependence but is still a marginal, one-column-at-a-time statistic.
Demonstrate the working discipline: use filters to shrink a wide table, engineer suspected interactions explicitly, and check pairwise redundancy before accepting a top-k list as the feature set.
Own the framing that a marginal statistic cannot answer a conditional question, and set the team norm that a filter ranking is a screening artefact, never the evidence quoted for which features drive the model.
## Two statistics, one shared limitation Pearson correlation between a column and the target measures the strength of a **linear** association, on a scale from -1 to 1. Mutual information measures how much knowing the column reduces uncertainty about the target; it is zero exactly when the two are independent, and it is nonzero for curved, threshold-shaped or otherwise nonlinear relationships that correlation would report as roughly zero. So mutual information is strictly the more general dependence measure of the two. But both are computed **one column at a time against the target**. That property — univariateness — is the source of the two failures in the question, and upgrading correlation to mutual information does nothing about it. ## Failure one: redundancy survives Suppose a pricing table carries list price and net price correlated at 0.97. If list price relates to the target, net price relates to the target almost identically. A univariate filter scores them separately, both score near the top, and both are kept. Nothing in the procedure can express the thought `I already have this information`. The damage is real but usually not catastrophic accuracy loss. You keep a column you are paying to collect, store, monitor and serve for no additional signal. Per-feature credit gets split between the twins, so any downstream story about which feature matters becomes unstable — small data changes flip which twin looks more important. And for a linear model, near-duplicate inputs make the individual coefficient estimates unstable even when predictions are fine. The fix inside the filter family is to stop scoring columns in isolation: redundancy-aware selection such as minimum-redundancy–maximum-relevance picks the next column by trading off its relevance to the target against its similarity to the columns already chosen. That is still cheap — it needs pairwise statistics, not model fits — and it directly addresses the twin problem. A crude, common alternative is a pre-pass that clusters columns by pairwise correlation and keeps one representative per cluster. ## Failure two: joint signal is invisible Now the other direction. Take two binary channels and a target that is 1 when they agree and 0 when they disagree. Marginally, each channel is uninformative: whichever value it takes, the target is still 50/50 until you know the other channel. Correlation with the target is about zero. Mutual information with the target is about zero, because the two really are independent taken one at a time. A univariate filter deletes both, and every downstream model — however powerful — is now solving a problem from which the answer has been removed. This is not a textbook curiosity. In a 20,000-column gene-expression table with 80 samples, a gene that only matters conditional on the state of a co-expressed gene shows almost no marginal association, and a top-k filter on marginal scores never lets it through. The same shape appears in sensor data whenever the meaningful quantity is a difference or a ratio between two channels rather than either channel's level. ## What actually helps Three honest options, in rising cost. 1. **Engineer the interaction explicitly.** If you suspect the ratio or the difference or the agreement matters, build that column and let the filter score it. Filters are excellent at scoring features you thought to create. 2. **Use a redundancy-aware or multivariate filter.** Still cheap, fixes the twin problem, and partially fixes the joint-signal problem. 3. **Evaluate columns jointly.** Families that judge subsets by what the fitted model actually does with them see interactions and redundancy by construction, because the model does. You pay for that in model fits and in a result that is tied to the learner you used. ## The interview point When you are asked why a filter missed something, the wrong answer is `because correlation is linear`. That explains only half of one failure. The right frame is: a univariate filter answers the question `is this column useful by itself`, and neither of the two failures above is about a column by itself. One is about a column being redundant **given another column**; the other is about a column being useful **only given another column**. Both are conditional questions, and a marginal statistic cannot answer a conditional question no matter how sophisticated the marginal statistic is. A reasonable working discipline is to use univariate filters exactly where they are strong — cheaply cutting a very wide table down to a size that a joint method can handle — and never to treat the filter's ranking as a final verdict on which features matter.
- Does swapping correlation for mutual information fix the interaction blindness?No. Mutual information captures nonlinear dependence between one column and the target, so it rescues threshold and curved relationships that correlation misses. It is still computed one column at a time, so a feature that is informative only jointly with another still measures as independent of the target and still gets dropped.
- When is keeping both of two 0.97-correlated columns actually acceptable?When nothing downstream is hurt by it: the model tolerates near-duplicates, serving cost is not a constraint, and no one is reading per-feature attributions as a decision. The harm from twins is wasted cost and unstable per-feature credit more than lost accuracy, so if none of those bind, the twin is a nuisance rather than a defect.
- How would you catch the interaction case before selection ever runs?Bring domain knowledge to the table and construct the candidate interaction — the difference, ratio, or agreement of the two channels — as an explicit column. A filter scores whatever you hand it, so the cheap fix is usually to hand it the right feature rather than to buy a more expensive search.
It is like hiring by reading each résumé alone: you happily hire two identical candidates, and you pass over the one whose value only shows up next to a particular colleague.
saying these in an interview costs you the question
- Says the problem is that correlation is only linear
- Thinks mutual information detects feature interactions
- Believes a univariate filter removes duplicate columns
- Treats a filter ranking as a final verdict on importance
- Assumes a zero univariate score means a useless feature