How do you choose between scikit-learn's SelectKBest, SelectFromModel and RFE?
answer
- three families, three price tags
- score once, fit once, or fit many times
- one of them ignores interactions entirely
- one needs coef_ or feature_importances_
- one needs non-negative inputs
basics
~20 sSelectKBest scores each feature independently against the target and is cheapest. SelectFromModel fits one estimator and thresholds its coefficients or importances. RFE refits repeatedly, dropping the weakest features each round — the most expensive and the only one that reacts to what removal does.
solid answer
~40 sThe three sit on a cost-versus-fidelity ladder. `SelectKBest(score_func=..., k=...)` and `SelectPercentile` are univariate filters: each feature is scored on its own with `f_classif`, `chi2`, `mutual_info_classif` or `f_regression`, so they are fast but blind to interactions and to redundancy between correlated features. `SelectFromModel(estimator, threshold=...)` fits one estimator and keeps features whose `coef_` or `feature_importances_` clear the threshold — so the estimator must expose one of those attributes; with an L1-penalised linear model it is effectively an embedded selector for one fit's cost. `RFE(n_features_to_select, step)` and `RFECV` refit the estimator repeatedly, discarding the weakest features each round, which captures how removing one feature changes the rest but costs many fits. All three are transformers exposing `get_support()`, and all must sit inside the pipeline so each fold selects on its own training data.
code
python · 7 linesfrom sklearn.datasets import make_classification
from sklearn.feature_selection import SelectKBest, f_classif
X, y = make_classification(n_samples=200, n_features=20, random_state=0)
sel = SelectKBest(score_func=f_classif, k=5).fit(X, y)
print(sel.get_support(indices=True))
print(sel.transform(X).shape) # (200, 5)go deeper
Know that these are transformers: fit learns which columns to keep, transform returns the narrowed matrix, and get_support() tells you which ones survived.
Be able to place the three on a cost ladder and say what each inspects — a per-feature statistic, one model's importances, or repeated refits — plus the chi2 non-negativity rule.
Demonstrate that you keep the selector inside the pipeline so each fold selects on its own training data, and that you check selection stability across folds before trusting a chosen feature set.
Own the prior question of whether to select at all. Weigh regularisation, native categorical or sparse handling, and the operational cost of a narrower feature set against the compute and instability that wrapper methods introduce.
## The three families in sklearn.feature_selection **Unsupervised.** `VarianceThreshold(threshold=0.0)` drops columns whose variance falls below the threshold. It never looks at the target, so it cannot overfit and is safe to run anywhere, but it only removes constants and near-constants. **Filters (univariate).** `SelectKBest`, `SelectPercentile` and `GenericUnivariateSelect` all take a `score_func` and keep the top features by that score. The scoring functions matter: - `f_classif` — ANOVA F-value, the classification default. - `f_regression` — the regression analogue. - `chi2` — requires non-negative features, because it is a chi-squared test on counts. Feed it standardised data and it raises. This is the single most-asked detail of the module. - `mutual_info_classif` / `mutual_info_regression` — capture non-linear dependence, at a much higher compute cost and with a `random_state` because the estimator is stochastic. Filters score each feature in isolation, so two perfectly correlated strong features both survive, and a feature that only matters in combination with another is invisible. **Embedded.** `SelectFromModel` wraps an estimator, fits it once, and keeps features whose importance clears `threshold` — a number, or a string expression like `'median'` or `'1.5*mean'`. `max_features` caps the count. The estimator must expose `coef_` or `feature_importances_` after fitting, which is why `LogisticRegression`, `Lasso`, `RandomForestClassifier` and the boosting estimators work while `SVC(kernel='rbf')` and `KNeighborsClassifier` do not. `prefit=True` lets you pass an already-fitted estimator, in which case the wrapper only transforms. This family is a good default: one fit, and the importances reflect the model's own view of the features rather than a generic statistic. **Wrappers.** `RFE` fits the estimator, ranks features by importance, drops the weakest `step`, and repeats until `n_features_to_select` remain, exposing `ranking_` and `support_`. `RFECV` wraps that in cross-validation to choose the number of features rather than making you guess, with `min_features_to_select` and `scoring`. `SequentialFeatureSelector` goes further still, evaluating candidate additions or removals by cross-validated score (`direction='forward'` or `'backward'`), which means an inner CV loop for every candidate at every step — the most expensive option by a wide margin, and the only one that optimises the metric you actually care about rather than a proxy. ## Cost, concretely For `p` features: a filter is one pass. `SelectFromModel` is one estimator fit. `RFE` with `step=1` down to `k` features is roughly `p - k` fits. `RFECV` multiplies that by the number of folds. `SequentialFeatureSelector` is on the order of `p^2` cross-validated fits. On a wide matrix that difference is the difference between seconds and hours, and it is the first thing to say when asked to choose. ## The wide-data case With thousands of features and a few hundred rows, everything overfits — including the selector. Two disciplines matter. First, put the selector inside the pipeline so it is refitted on each fold's training part; scoring features against the full target once and then cross-validating the survivors produces an optimistic number that will not survive contact with new data. Second, prefer a cheap filter to prune the obviously useless bulk, then an embedded selector on what remains, rather than running a wrapper over the raw width. It is also worth saying that selection is not always the right tool. A regularised model often matches selection while keeping one fit, and tree ensembles tolerate irrelevant features better than linear models do. Selection earns its keep when you need a smaller feature set for cost, latency, or explanation reasons rather than purely for accuracy. ## Inspecting the result Every selector implements `get_support()` for a boolean mask, or `get_support(indices=True)` for positions, and `get_feature_names_out()` for names when the input carried them. Log which features survived across folds: a selection that changes wildly between folds is telling you the choice is noise, not signal. ## How to answer Rank the three by cost, name what each actually inspects, give the `chi2` non-negativity constraint and the `coef_`/`feature_importances_` requirement as the concrete API facts, and close on the discipline of refitting the selector per fold.
- Why does chi2 raise on your standardised features?`chi2` is a chi-squared test over non-negative quantities such as counts or frequencies, so it rejects negative input. Standardising centres columns on zero and produces negatives. Either score the raw non-negative features, switch to `f_classif` or `mutual_info_classif`, or scale with `MinMaxScaler` instead of `StandardScaler`.
- Which estimators can you pass to SelectFromModel?Any that exposes `coef_` or `feature_importances_` after fitting — linear models including `Lasso` and `LogisticRegression`, decision trees and forests, and gradient boosting. `SVC` with a non-linear kernel and `KNeighborsClassifier` expose neither, so they cannot be used. `prefit=True` accepts an already-fitted estimator.
- With 5,000 features and 800 rows, would you reach for SequentialFeatureSelector?Not as a first move. It runs an inner cross-validation for every candidate at every step, which is on the order of thousands of CV loops here. Prune with a cheap filter or an L1-penalised `SelectFromModel` first, and only consider a wrapper on the reduced set — if a regularised model has not already made selection unnecessary.
saying these in an interview costs you the question
- Selecting features on the full dataset before cross-validating
- Assuming univariate scores account for feature interactions
- Passing standardised (negative) values to chi2
- Expecting SelectFromModel to work with an RBF-kernel SVC
- Treating more selection as always better than regularisation