skip to content

Filter, Wrapper, Embedded

Variance and correlation filters, recursive feature elimination and forward search, and how the embedded family differs from both. Interviewers compare their compute cost and bias.

on this pageshow

questions

5

How do filter, wrapper and embedded feature selection differ in cost and in what they can see?

level: middleimportance: must knowfreq 68%

answer

  1. count model fits per family
  2. one pass, many fits, one fit
  3. does the method see columns together
  4. wrapper answers are learner-specific
  5. filters cascade into the expensive step

basics

~20 s

Filters rank columns with a cheap statistic and never train the model. Wrappers train it many times on candidate subsets and keep the best. Embedded selection falls out of one model fit. Cost and model-specificity rise across the three.

solid answer

~50 s

A filter scores each column against the target with a statistic — a correlation, a mutual-information estimate, a group-difference test — and keeps the top ones. It costs one pass over the data, is independent of whichever model comes later, and is blind to how columns behave together. A wrapper searches over subsets and scores each candidate by actually training the model and measuring held-out performance; recursive feature elimination and forward or backward stepwise search are the standard forms. It sees interactions and redundancy because the model does, but it costs tens to thousands of fits and the search itself can overfit the scores it is optimising. Embedded selection comes free with a single fit, because the learner's own training procedure decides which columns end up with influence — but the resulting set is tied to that learner. In practice I cascade: a cheap filter to cut a very wide table down, then a wrapper or embedded step on what survives.

go deeper

for a junior

Be ready to name the three families and give one example of each, and to say that a filter runs before any model while a wrapper needs the model trained repeatedly.

for a middle

Explain the cost in model fits and the information each family has access to, especially why a one-column-at-a-time score cannot see how columns behave together.

for a senior

Show that you choose the family from the problem shape — table width, cost of a fit, row count — and that you cascade a cheap stage into an expensive one rather than picking one family dogmatically.

for a principal

Own the tradeoff that a wrapper's accuracy is bought with compute and with a subset that is tied to one learner, and decide when a model-agnostic, cheaper, more stable selection is the better organisational bet.

## The three families Every feature-selection method answers the same question — which columns do we keep — but they differ in what evidence they use and how much computation they spend getting it. The standard taxonomy has three families. **Filters** score columns using a statistic computed from the data alone, before and independently of any model. Typical scores are the absolute correlation between a column and the target, an estimate of mutual information between a column and the target, a difference-of-means style test statistic across classes, or an unsupervised hygiene check such as a variance threshold. You compute one number per column, sort, and keep the top k or everything above a cutoff. **Wrappers** treat the learner as a black box and search the space of subsets, scoring each candidate subset by training the model on it and measuring held-out performance. Forward stepwise selection starts from nothing and repeatedly adds the column that most improves the score. Backward elimination starts from everything and repeatedly removes the least useful. Recursive feature elimination fits the model, ranks columns by the model's own weights, drops the weakest, and repeats. **Embedded** methods fuse selection into a single model fit: the training procedure itself decides which columns end up carrying influence, so you get a selected set as a side effect of fitting once. The specific mechanisms — a sparsity-inducing penalty on a linear model, or a tree learner's own choice of which columns to split on — are each covered under their own topics; what matters for the taxonomy is the interface, which is that you pay for one fit and receive a subset. ## Cost, in the units that matter Count model fits, because that is what dominates. - A filter costs **zero fits**. It is `p` cheap statistics for `p` columns, one sweep of the data. A 20,000-column table is not a problem. - A wrapper costs **many fits**. Forward selection over `p` candidates that stops after `k` steps evaluates roughly `p + (p-1) + ... ` candidates, on the order of `p * k` fits — and multiply that by the number of cross-validation folds used to score each candidate. Sixty candidate columns and four steps is already several hundred fits per fold. - An embedded method costs **one fit**, plus whatever hyperparameter tuning that model needed anyway. The practical consequence is that wrappers are affordable when `p` is in the dozens and the model trains in seconds, and infeasible when `p` is in the thousands and the model trains in minutes. ## What each family can and cannot see The cost ordering mirrors an information ordering. A univariate filter looks at one column at a time against the target. It therefore cannot notice that two columns say the same thing, and it cannot notice that a column is worthless alone but decisive in combination with another. It is also agnostic about the downstream model, which is a genuine virtue when you do not yet know what you will train, and a genuine limitation when the model has strong preferences. A wrapper sees exactly what the model sees, including interactions and redundancy, because the evidence is the model's own held-out performance on that subset. That fidelity is why it is the most accurate family in principle. It comes with two costs beyond compute: the chosen subset is tuned to one learner and does not necessarily transfer, and searching many subsets and reporting the winner produces an optimistically biased score unless the whole procedure is judged on data the search never touched. An embedded method also sees columns jointly, since the fit considers them together, but only through the lens of that one model's inductive bias. A subset chosen by one model family is a statement about that family, not a universal ranking of feature value. ## How they are actually combined The families are not competitors so much as stages. With 20,000 gene-expression columns and 80 samples, no wrapper is going to run and no wrapper's result would be stable at that sample size; a variance pass and then a univariate filter down to a few hundred columns makes the problem tractable, and something joint runs afterwards. With 60 sensor channels and plenty of rows, the filter stage buys almost nothing and the wrapper is affordable, so you skip straight to it. The question an interviewer is really probing is whether you pick a family from the shape of the problem — how wide the table is, how expensive one fit is, how many rows you have, and whether the selected set must be defensible to someone who did not choose the model.

  • With 20,000 columns and 80 rows, which family do you reach for first and why?
    A filter. A wrapper needs hundreds or thousands of fits over a space that huge, and with 80 rows every held-out score it computes is so noisy that the search would chase noise. A cheap univariate pass cuts the table to a few hundred columns; anything joint runs after that on a survivable problem size.
  • Does a subset chosen by a wrapper transfer to a different learner?
    Not reliably. The wrapper optimised held-out score for one model, so the subset encodes that model's inductive bias — what a linear learner needs is not what a tree-based learner needs. If you switch learners, re-run the selection, or use a filter stage whose output is model-agnostic by construction.
  • When would you skip feature selection entirely?
    When the table is narrow relative to the rows, the model already handles irrelevant columns gracefully, and nothing downstream needs a shorter list. Selection buys inference cost, data-collection cost and reviewability — if none of those is binding, the search is cost and risk with no return.

saying these in an interview costs you the question

  • Calls a wrapper strictly better without mentioning its cost
  • Thinks filters are model-specific
  • Believes a wrapper's chosen subset transfers to any learner
  • Cannot say how many model fits each family needs
  • Treats the three families as mutually exclusive choices

context

open as a page

Why can a correlation or mutual-information filter keep two 0.97-correlated columns yet drop a feature that only matters in combination?

level: middleimportance: should knowfreq 52%

basics

~20 s

A correlation or mutual-information filter scores each column against the target on its own. Two near-duplicates both score well and survive; a feature that matters only alongside another scores near zero alone and is cut. The blind spot is univariate scoring.

open as a page

Recursive feature elimination over 400 marketing-attribution columns runs for hours — what is it doing and how do you cut the cost?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Recursive feature elimination refits the model, ranks columns by its weights, drops the weakest, and repeats — hundreds of fits, times cross-validation folds. Cut cost by dropping a block per round, pre-screening with a cheap filter, or ranking with a faster model.

open as a page

Forward stepwise selection picked 8 of 60 sensor channels; why is the winning subset's cross-validated score optimistic?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Forward stepwise selection evaluated hundreds of candidate subsets and reported the winner. The maximum of many noisy estimates is biased upward, so part of the winner's margin is luck — even when the selector saw only training rows.

open as a page

What does a variance-threshold filter remove from a feature matrix, and when does it drop something useful?

level: juniorimportance: nice to knowfreq 30%

basics

~20 s

A variance-threshold filter drops every column whose spread across rows falls below a cutoff, removing constant and near-constant features. It ignores the target and depends on units, so it can discard a rare binary flag that predicts strongly.

open as a page