Scikit-learn
Scikit-learn is the workhorse for everything that is not deep learning: a uniform estimator API, preprocessing, pipelines, cross-validation, and metrics. Data-science interviews lean on it heavily because it forces you to talk about methodology, not just models.
on this pageshowhide
explore
- Estimator API6 questions
- Preprocessing & Features6 questions
- Pipelines6 questions
- Model Selection & CV6 questions
- Supervised Estimators6 questions
- Metrics & Evaluation6 questions
questions
page 2 of 2When does Pipeline(memory=...) in scikit-learn actually save time, and what does it cost?
basics
~20 sSetting memory on a Pipeline caches fitted transformers on disk, keyed by the transformer, its parameters and its input. It pays off during a search where many candidates share an identical, expensive prefix — and buys nothing when every candidate changes an early step.
How do you choose between scikit-learn's SelectKBest, SelectFromModel and RFE?
basics
~20 sSelectKBest scores each feature independently against the target and is cheapest. SelectFromModel fits one estimator and thresholds its coefficients or importances. RFE refits repeatedly, dropping the weakest features each round — the most expensive and the only one that reacts to what removal does.
What does class_weight='balanced' compute, and which scikit-learn estimators accept it?
basics
~20 sclass_weight='balanced' sets each class's weight to n_samples / (n_classes * count of that class), so rare classes contribute proportionally more loss. LogisticRegression, SVC, LinearSVC, the tree and forest classifiers and HistGradientBoostingClassifier accept it; GradientBoostingClassifier does not.
When should you use LinearSVC or SGDClassifier instead of SVC in scikit-learn?
basics
~20 sSVC's kernel solver has fit time scaling at least quadratically in the number of samples, so it becomes impractical past roughly tens of thousands of rows. For large linear problems reach for LinearSVC, backed by liblinear, or SGDClassifier, whose cost is linear in the sample count.
Which preprocessing belongs inside a persisted scikit-learn Pipeline versus upstream ETL?
basics
~20 sAnything with state learned from training data — imputer statistics, encoder categories, scaler means, vectorizer vocabularies, selection masks — must live inside the fitted Pipeline so training and serving share one artifact. Stateless business joins and label construction stay upstream.
What are scikit-learn estimator tags, and when do you implement __sklearn_tags__?
basics
~20 sEstimator tags are declarative metadata describing what an estimator can accept and do — sparse input, NaN tolerance, whether y is required, the estimator type. Since scikit-learn 1.6 they are a dataclass returned by sklearn_tags, which you override only when your estimator departs from the defaults.
showing 31–36 of 36