skip to content

What does a variance-threshold filter remove from a feature matrix, and when does it drop something useful?

level: juniorimportance: nice to knowfreq 30%

answer

  1. it never looks at the target
  2. constant columns carry nothing
  3. change the units, change the number
  4. rare binary flag, tiny variance
  5. p times one minus p

basics

~20 s

A variance-threshold filter drops every column whose spread across rows falls below a cutoff, removing constant and near-constant features. It ignores the target and depends on units, so it can discard a rare binary flag that predicts strongly.

solid answer

~50 s

It computes the variance of each column on its own and deletes the ones below a chosen cutoff. At a cutoff of zero that is pure hygiene: a constant column carries no information for any model, so dropping it is always safe. Above zero it becomes a judgment call with two traps. First, variance depends on units — the same measurement in millimetres has a million times the variance it has in metres — so a single cutoff across mixed-unit columns is meaningless, and standardising first sets every variance to 1 and makes the filter a no-op. Second, it is unsupervised: a binary flag that is 1 in 1% of rows has variance about `0.0099`, so a cutoff of `0.01` deletes it even if it is the single best predictor of a rare event. I use it only as the first cheap pass that makes an expensive selection step tractable.

go deeper

for a junior

Be ready to state the rule in one line: it drops columns that barely vary, using only the column itself. Know that a constant column is always safe to remove and that the target is never consulted.

for a middle

Explain why the cutoff is unit-dependent and why standardising first defeats the filter entirely. Be able to compute a binary column's variance as the positive rate times one minus the positive rate.

for a senior

Show the operational habit: log which columns the threshold removed and eyeball the list before trusting it, because rare indicator flags in imbalanced problems are exactly what a naive cutoff deletes.

for a principal

Own the framing that this is data hygiene, not selection on signal, and keep it out of any pipeline stage that is supposed to be making evidence-based keep-or-drop decisions about predictive value.

## What the filter actually computes For each column `x`, variance is the average squared distance of its values from their own mean: `var(x) = mean((x - mean(x))^2)`. A variance-threshold filter computes that one number per column and keeps only the columns whose variance exceeds a cutoff. Nothing about the target is involved, and nothing about any other column is involved — it is the cheapest possible screening pass, one sweep over the data, and it scales to tens of thousands of columns without difficulty. ## The one case that is always safe A column with variance exactly zero is constant: every row has the same value. No model can use it, because there is no variation to associate with variation in the target. A linear model gives it an unidentifiable coefficient, a tree can never split on it, a distance metric gets a constant offset from it. Dropping zero-variance columns is hygiene rather than selection, and it is the reason this filter is usually the first step in a long pipeline. In a wide problem — say a 20,000-column microarray with 80 samples — a large share of columns are flat or nearly flat across those 80 samples, and removing them before anything expensive runs is pure profit. ## Trap one: variance carries units Variance is measured in the square of the column's units. A distance column recorded in metres and the same column recorded in millimetres differ in variance by a factor of a million. A price column in whole currency units and a probability column bounded in `[0, 1]` are not on any common footing at all. So a single numeric cutoff applied across a mixed-unit table does not express one consistent idea of `too flat`; it mostly ranks columns by how large their units happen to be. The obvious fix breaks the filter. If you standardise every column to mean 0 and standard deviation 1 first, then every variance equals 1 by construction, and the threshold either keeps everything or deletes everything. The practical resolutions are: apply the filter before scaling and only at or near zero variance; apply it within groups of columns that genuinely share a scale; or, for binary indicator columns, think in terms of the positive rate instead of the variance. ## Trap two: low variance is not low signal Because the filter is unsupervised, it has no way to distinguish a boring column from a rare but decisive one. For a binary column that is 1 in a fraction `p` of rows, the variance is exactly `p * (1 - p)`. At `p = 0.01` that is `0.0099`; at `p = 0.001` it is about `0.000999`. Any cutoff meant to remove `almost constant` columns therefore removes rare flags. In fraud detection, defect detection or adverse-event modelling, the rare flag is frequently the strongest single predictor in the table — precisely because it fires only when the rare outcome is in play. Deleting it because it is rare inverts the thing you were trying to do. The mirror-image misconception is the belief that high variance means high value. A column of random noise has as much variance as you like and zero relationship to the target. Variance measures how much a column moves, not whether the movement means anything. ## Where it belongs Treat a variance threshold as a hygiene pass, not as feature selection on signal. It answers one narrow question — does this column vary at all — and it says nothing about whether a column relates to the target, and nothing about whether two columns say the same thing. Those questions belong to later, more expensive stages. Set the cutoff at or very near zero unless you have a specific, scale-aware reason to raise it, and if you do raise it, check by hand which columns fell out before you accept the result. A one-line log of dropped column names has saved more models than a clever threshold ever has.

  • If you standardise every column before applying the filter, what happens?
    Standardising forces every column to variance 1, so a single cutoff either keeps all columns or deletes all of them. Apply the filter before scaling, on raw columns, and restrict it to constant or near-constant removal — or compare only columns that already share a scale.
  • How would you set a cutoff for binary indicator columns?
    Reason about the positive rate rather than a variance number, since variance is `p * (1 - p)`. Decide the smallest number of positive rows you could estimate an effect from — a few dozen, typically — and translate that into a rate. That keeps the decision about statistical support instead of about units.

It is like clearing a bookshelf by throwing out every book thinner than a centimetre: fast and mostly harmless, but the thinnest volume may be the one that matters.

saying these in an interview costs you the question

  • Thinks a variance threshold uses the target label
  • Standardises first, then wonders why nothing is dropped
  • Assumes low variance implies low predictive value
  • Applies one cutoff across columns with different units
  • Calls high variance evidence that a column is informative

context