skip to content

Scikit-learn

Scikit-learn is the workhorse for everything that is not deep learning: a uniform estimator API, preprocessing, pipelines, cross-validation, and metrics. Data-science interviews lean on it heavily because it forces you to talk about methodology, not just models.

on this pageshow

explore

questions

page 2 of 2

When does Pipeline(memory=...) in scikit-learn actually save time, and what does it cost?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Setting memory on a Pipeline caches fitted transformers on disk, keyed by the transformer, its parameters and its input. It pays off during a search where many candidates share an identical, expensive prefix — and buys nothing when every candidate changes an early step.

open as a page

How do you choose between scikit-learn's SelectKBest, SelectFromModel and RFE?

level: seniorimportance: should knowfreq 44%

basics

~20 s

SelectKBest scores each feature independently against the target and is cheapest. SelectFromModel fits one estimator and thresholds its coefficients or importances. RFE refits repeatedly, dropping the weakest features each round — the most expensive and the only one that reacts to what removal does.

open as a page

What does class_weight='balanced' compute, and which scikit-learn estimators accept it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

class_weight='balanced' sets each class's weight to n_samples / (n_classes * count of that class), so rare classes contribute proportionally more loss. LogisticRegression, SVC, LinearSVC, the tree and forest classifiers and HistGradientBoostingClassifier accept it; GradientBoostingClassifier does not.

open as a page

When should you use LinearSVC or SGDClassifier instead of SVC in scikit-learn?

level: seniorimportance: should knowfreq 46%

basics

~20 s

SVC's kernel solver has fit time scaling at least quadratically in the number of samples, so it becomes impractical past roughly tens of thousands of rows. For large linear problems reach for LinearSVC, backed by liblinear, or SGDClassifier, whose cost is linear in the sample count.

open as a page

Which preprocessing belongs inside a persisted scikit-learn Pipeline versus upstream ETL?

level: principalimportance: should knowfreq 30%

basics

~20 s

Anything with state learned from training data — imputer statistics, encoder categories, scaler means, vectorizer vocabularies, selection masks — must live inside the fitted Pipeline so training and serving share one artifact. Stateless business joins and label construction stay upstream.

open as a page

What are scikit-learn estimator tags, and when do you implement __sklearn_tags__?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Estimator tags are declarative metadata describing what an estimator can accept and do — sparse input, NaN tolerance, whether y is required, the estimator type. Since scikit-learn 1.6 they are a dataclass returned by sklearn_tags, which you override only when your estimator departs from the defaults.

open as a page

showing 31–36 of 36