skip to content

Scikit-learn

Scikit-learn is the workhorse for everything that is not deep learning: a uniform estimator API, preprocessing, pipelines, cross-validation, and metrics. Data-science interviews lean on it heavily because it forces you to talk about methodology, not just models.

on this pageshow

explore

questions

page 1 of 2

In scikit-learn, what does a trailing underscore in coef_ or mean_ mean?

level: juniorimportance: must knowfreq 56%

answer

  1. two kinds of attribute, one naming rule
  2. the suffix is not privacy
  3. set only inside fit
  4. how the library detects fitted state
  5. check_is_fitted scans for it

basics

~20 s

A trailing underscore marks state learned during fit, such as coef_, mean_ or n_features_in_. It does not exist on a freshly constructed estimator, and check_is_fitted uses the presence of such attributes to decide whether the estimator has been fitted.

solid answer

~40 s

scikit-learn splits an estimator's attributes into two disjoint groups. Names without a trailing underscore are the hyperparameters you passed to `__init__` — they exist from construction and are what `get_params` reports. Names ending in an underscore are *fitted* state, created only inside `fit`: `coef_`, `mean_`, `classes_`, `n_features_in_`. The convention is load-bearing, not cosmetic: `check_is_fitted(self)` decides an estimator is fitted by looking for any attribute that ends in an underscore and is not a dunder, and raises `NotFittedError` otherwise, which is how calling `predict` before `fit` produces a clear error. `fit` must also return `self` so chaining works, and by default it *resets* — a second `fit` discards the previous fitted attributes rather than continuing training.

go deeper

for a junior

Recall the split: no underscore means a constructor hyperparameter, a trailing underscore means something learned during fit. Say that reading coef_ before fitting fails, and that fit returns self.

for a middle

Explain the mechanism behind it — check_is_fitted scans for trailing-underscore attributes and raises NotFittedError, so the convention is what makes the guard work in a custom estimator.

for a senior

Talk about the operational edges: fit resetting rather than accumulating, warm_start and partial_fit as the deliberate exceptions, and n_features_in_/feature_names_in_ catching schema drift between training and serving.

for a principal

Own the consistency argument — a naming convention that generic code can rely on removes the need for per-estimator metadata, and any in-house estimator that ignores it becomes invisible to the whole ecosystem of wrappers built on it.

## Two disjoint groups of attributes Every scikit-learn estimator carries two kinds of state, and the naming tells you which is which at a glance. **Hyperparameters** are the arguments to `__init__`, stored under their own names with no underscore: `alpha`, `n_estimators`, `C`, `with_mean`. They exist the instant the object is constructed, they are what `get_params` reports and `set_params` writes, and they are the only thing that survives a `clone`. **Fitted attributes** are learned from data. They are created inside `fit` and their names end in a single trailing underscore: `coef_`, `intercept_`, `mean_`, `scale_`, `classes_`, `n_iter_`, `feature_importances_`. On a freshly constructed estimator they do not exist at all — touching one raises `AttributeError`. That separation is what lets generic library code reason about an object it has never seen. Configuration is introspectable and reproducible; learned state is disposable and rebuilt from data. ## The underscore is not "private" The common confusion is with Python's own convention, where a *leading* underscore signals "internal, do not touch". scikit-learn's marker is on the other end of the name, and it means the opposite: `coef_` is public API, the thing you are meant to read after fitting. A leading underscore in scikit-learn still means internal, so `self._cache` is private scratch, while `self.coef_` is a documented result. ## check_is_fitted and NotFittedError `sklearn.utils.validation.check_is_fitted(estimator)` is how estimators guard `predict`, `transform` and `score`. With no explicit attribute list, it scans the instance's `__dict__` for any name that ends in `_` and is not a dunder. If it finds none, it raises `sklearn.exceptions.NotFittedError` with a message telling you to call `fit` first. Two practical implications. First, this is why the convention must be followed in a custom estimator: if you store learned state as `self.mean` instead of `self.mean_`, `check_is_fitted` cannot see it and will insist the estimator is unfitted forever. Second, it explains a subtler bug — assigning *any* underscore-suffixed attribute in `__init__` makes the estimator look fitted from birth, so learned state must be created in `fit` and nowhere else. You can also pass explicit names, `check_is_fitted(self, "mean_")`, when you want a specific attribute checked. ## The fitted attributes almost every estimator sets Input validation during `fit` records two standard pieces of state for you: - `n_features_in_` — the number of columns of `X` seen during `fit`. At `predict` time the estimator compares the incoming width against it and raises a clear `ValueError` on a mismatch instead of producing nonsense. - `feature_names_in_` — an array of column names, set **only** when `X` had string feature names, i.e. when you fitted on a pandas DataFrame. Fit on a DataFrame and predict on a raw NumPy array and you get a warning about missing feature names; fit on an array and predict on a DataFrame gets one too. In a custom estimator, calling `validate_data(self, X)` from `sklearn.utils.validation` inside `fit` sets both, and `validate_data(self, X, reset=False)` inside `transform`/`predict` checks the new data against them rather than overwriting them. (Before scikit-learn 1.6 the same logic lived on the private `BaseEstimator._validate_data` method.) ## fit resets; it does not accumulate Calling `fit` a second time on the same object is a fresh fit. The previous fitted attributes are overwritten, and any state from the earlier run is discarded — estimators are expected to behave as if they had just been cloned. This is deliberate: cross-validation and hyperparameter search reuse configurations constantly, and silent accumulation would make results depend on call history. There are two explicit opt-outs. Some iterative estimators accept `warm_start=True`, which tells `fit` to continue from the existing solution instead of resetting — useful for adding trees to a forest or extra iterations to a linear model. And a separate method, `partial_fit`, exists on estimators that support genuine incremental learning over mini-batches; it deliberately does not reset, and for classifiers usually needs the full `classes` list on the first call because it cannot infer the label set from one batch. ## fit returns self `fit` must end with `return self`. It is what makes `model.fit(X, y).predict(X_test)` and `scaler.fit(X).transform(X)` work, and generic library code relies on it. Returning `None`, or returning the transformed data, breaks pipelines and searches in confusing ways. ## Writing it correctly In a custom estimator the shape is fixed: `__init__` assigns hyperparameters only; `fit` validates input, computes results into trailing-underscore attributes, and returns `self`; `predict`/`transform` start with `check_is_fitted(self)` and then use only the fitted attributes and the hyperparameters. Follow that and the object becomes usable by every generic tool in the library without further work.

  • What exception does calling predict() before fit() raise, and what produces it?
    `sklearn.exceptions.NotFittedError`, raised by `check_is_fitted` at the top of `predict`. With no attribute list given, it scans the instance for any non-dunder attribute whose name ends in an underscore; finding none, it reports that the estimator is not fitted and tells you to call `fit` first. That is precisely why a custom estimator must store learned state with the trailing-underscore suffix.
  • Does calling fit() twice on the same estimator continue training from where it stopped?
    No. By default `fit` resets: the fitted attributes are recomputed from scratch and the earlier solution is discarded, so a refit behaves like fitting a clone. Two opt-ins exist — `warm_start=True` on estimators that support it continues from the current solution, and `partial_fit` performs genuine incremental learning over mini-batches without resetting.
  • What are n_features_in_ and feature_names_in_, and when is the second one absent?
    Both are set during `fit` by input validation. `n_features_in_` is the column count seen at fit time, checked again at predict time so a width mismatch raises a clear error. `feature_names_in_` holds the column names, and is set only when `X` carried string feature names — that is, when you fitted on a DataFrame. Fitting on a NumPy array leaves it unset, and mixing the two between fit and predict triggers a warning.

saying these in an interview costs you the question

  • Thinks the trailing underscore means the attribute is private
  • Sets learned attributes in __init__ rather than in fit
  • Expects fit to continue training by default on a second call
  • Thinks check_is_fitted reads a boolean flag you must set yourself
  • Forgets to return self from fit

context

open as a page

In scikit-learn, how do you read the output of confusion_matrix?

level: juniorimportance: must knowfreq 58%

basics

~20 s

confusion_matrix returns a square array where rows are true classes and columns are predicted classes, so C[i, j] counts samples of true class i predicted as class j. Classes appear in sorted order unless you pass labels= to fix the order.

open as a page

In scikit-learn, what does the stratify argument of train_test_split do?

level: juniorimportance: must knowfreq 80%

basics

~20 s

stratify=y tells train_test_split to preserve each class's proportion in both halves instead of splitting purely at random. Without it a rare class can land unevenly in the test set, or be missing from it entirely.

open as a page

In scikit-learn, what happens at each step when you call Pipeline.fit() then predict()?

level: juniorimportance: must knowfreq 72%

basics

~20 s

fit runs fit_transform on every step except the last, feeding each output into the next, then fit on the final estimator. predict runs transform only on those same intermediate steps — never fit again — and calls predict on the final estimator.

open as a page

When is scikit-learn's OrdinalEncoder a safe choice instead of OneHotEncoder?

level: juniorimportance: must knowfreq 68%

basics

~20 s

OrdinalEncoder maps each category to an integer in one column, which implies an ordering. That is safe for genuinely ordered features and for tree-based models that split on thresholds, but misleading for linear models, SVMs and distance-based estimators.

open as a page

Why must a scikit-learn estimator's __init__ store every argument unchanged?

level: middleimportance: must knowfreq 62%

basics

~20 s

In scikit-learn, get_params reads the init signature and returns the same-named attributes, and clone() feeds those values straight back into the constructor. Any renaming, conversion or validation inside init breaks that round-trip and makes clone raise RuntimeError.

open as a page

Which scikit-learn base classes do you inherit to write a custom transformer?

level: middleimportance: must knowfreq 58%

basics

~10 s

Inherit both, mixin first: class MyTransformer(TransformerMixin, BaseEstimator). BaseEstimator supplies get_params, set_params, the repr and the default tags; TransformerMixin supplies fit_transform and set_output. You then implement only init, fit returning self, and transform.

open as a page

In scikit-learn, what does precision_score's average parameter control?

level: middleimportance: must knowfreq 68%

basics

~20 s

The average parameter decides how per-class precision values are reduced to one number. It defaults to 'binary', which scores only one class and rejects multiclass targets; multiclass data needs 'macro', 'micro', 'weighted', or None for the per-class array.

open as a page

Why does scikit-learn's roc_auc_score need scores rather than predict() labels?

level: middleimportance: must knowfreq 64%

basics

~20 s

ROC-AUC measures how well a continuous score ranks positives above negatives, so it needs predict_proba(X)[:, 1] or decision_function(X). Hard 0/1 labels from predict() collapse the curve to one threshold and silently return a lower, meaningless number instead of raising.

open as a page

What splitter does scikit-learn use when you pass cv=5 to cross_val_score?

level: middleimportance: must knowfreq 68%

basics

~20 s

An integer cv is expanded for you: StratifiedKFold when the estimator is a classifier and the target is binary or multiclass, plain KFold otherwise. Neither shuffles — folds are contiguous blocks of the data in its current row order.

open as a page

When should you use RandomizedSearchCV instead of GridSearchCV in scikit-learn?

level: middleimportance: must knowfreq 74%

basics

~20 s

GridSearchCV fits every combination in param_grid, so its cost multiplies with each added parameter. RandomizedSearchCV draws a fixed n_iter samples from param_distributions, letting you cap the budget and sample continuous ranges instead of a hand-picked ladder.

open as a page

How do you address a Pipeline step's parameters in a scikit-learn GridSearchCV param_grid?

level: middleimportance: must knowfreq 68%

basics

~20 s

Use the step name, a double underscore, then the parameter name: "clf__C": [0.1, 1.0]. Nesting repeats the pattern for nested estimators, and using a step's bare name as the key replaces the whole step with another estimator or with "passthrough".

open as a page

In scikit-learn, why must a scaler live inside the Pipeline passed to cross_val_score?

level: middleimportance: must knowfreq 78%

basics

~20 s

cross_val_score refits whatever estimator you hand it on each training fold. A scaler fitted outside it has already seen every fold's held-out rows, so its means and variances encode validation data and the reported score comes out optimistically biased.

open as a page

In scikit-learn, how does ColumnTransformer apply transforms per column, and what happens to unlisted columns?

level: middleimportance: must knowfreq 66%

basics

~20 s

ColumnTransformer fits each listed transformer on its own column subset and concatenates the results side by side in the order the transformers are declared. Columns you did not list are dropped, because remainder defaults to 'drop'.

open as a page

How does scikit-learn's OneHotEncoder handle categories unseen during fit()?

level: middleimportance: must knowfreq 74%

basics

~10 s

By default OneHotEncoder raises a ValueError at transform time for any category it did not see in fit(). Setting handle_unknown='ignore' encodes the unknown value as an all-zero row instead, so scoring continues.

open as a page

What does scikit-learn's StandardScaler learn in fit(), and why never fit it on test data?

level: middleimportance: must knowfreq 78%

basics

~20 s

StandardScaler.fit() computes and stores each column's mean (mean_) and standard deviation (scale_). Test data must only go through transform(), because refitting replaces those statistics with test-set numbers the model would never have at serving time.

open as a page

When should you prefer HistGradientBoostingClassifier over GradientBoostingClassifier?

level: middleimportance: must knowfreq 56%

basics

~20 s

Prefer HistGradientBoostingClassifier on anything beyond a few thousand rows. It bins each feature into at most 255 integer bins so split-finding stops scaling with the number of distinct values, handles missing values natively, and accepts categorical columns directly.

open as a page

In scikit-learn, is LogisticRegression regularized by default, and what does C control?

level: middleimportance: must knowfreq 74%

basics

~20 s

Scikit-learn's LogisticRegression is regularized by default: penalty='l2' with C=1.0. C is the inverse of regularization strength, so a smaller C shrinks coefficients harder and a larger C fits closer to unpenalized. Pass penalty=None to switch it off.

open as a page

Why can two RandomForestClassifier fits on the same data give different predictions?

level: juniorimportance: should knowfreq 47%

basics

~20 s

A random forest is randomized twice: each tree trains on a bootstrap resample of the rows, and each split considers a random subset of features. With random_state left at None those draws differ per fit. Pass an integer random_state for reproducible results.

open as a page

What does sklearn.base.clone() copy from an estimator, and what does it drop?

level: middleimportance: should knowfreq 45%

basics

~20 s

clone() returns a new, unfitted estimator of the same class built from get_params(deep=False): hyperparameters only. All fitted attributes are dropped. Plain parameter values are deep-copied, and any parameter that is itself an estimator is cloned recursively, so it comes back unfitted too.

open as a page

When does scikit-learn's fit_transform() differ from fit() then transform()?

level: middleimportance: should knowfreq 48%

basics

~20 s

Usually not at all: TransformerMixin's default fit_transform is literally fit(X, y).transform(X). Transformers override it when fitting and transforming together is cheaper, or when the embedding is defined only for the fitted samples — sklearn.manifold.TSNE offers fit_transform and no transform.

open as a page

In scikit-learn, when do you use silhouette_score versus adjusted_rand_score?

level: middleimportance: should knowfreq 32%

basics

~20 s

silhouette_score(X, labels) is internal: it scores cluster shape from the feature matrix alone, so it works without ground truth. adjusted_rand_score(labels_true, labels_pred) is external: it compares two labelings and requires true labels you usually do not have.

open as a page

In scikit-learn, how does make_scorer turn a metric into a scorer?

level: middleimportance: should knowfreq 46%

basics

~10 s

make_scorer wraps a metric function into a callable of the form scorer(estimator, X, y). greater_is_better=False negates the value so larger always means better, and response_method decides whether the scorer calls predict, predict_proba, or decision_function.

open as a page

How does scikit-learn's TimeSeriesSplit differ from KFold, and what is gap for?

level: middleimportance: should knowfreq 55%

basics

~20 s

TimeSeriesSplit never puts later rows in a training fold: each split trains on a prefix of the rows and tests on the block immediately after, with the training window growing each split. gap drops a fixed number of rows between the train end and the test start.

open as a page

What does FeatureUnion do in scikit-learn, and when is it the wrong tool?

level: middleimportance: should knowfreq 34%

basics

~20 s

FeatureUnion fits several transformers on the same input in parallel and horizontally concatenates their outputs into one feature matrix. It is the wrong tool when each transformer should see a different subset of columns — that is ColumnTransformer's job.

open as a page

How do scikit-learn's SimpleImputer, KNNImputer and IterativeImputer differ?

level: middleimportance: should knowfreq 55%

basics

~20 s

SimpleImputer learns one constant per column (mean, median, most frequent or a fixed value) and stores it in statistics_. KNNImputer fills from similar training rows, and IterativeImputer models each column from the others — both far more expensive.

open as a page

Why does SVC.predict_proba in scikit-learn require probability=True, and what does it cost?

level: middleimportance: should knowfreq 40%

basics

~20 s

An SVM produces a signed distance to the hyperplane, not a probability. Setting probability=True makes fit run Platt scaling — a sigmoid fitted by internal 5-fold cross-validation — which sharply increases fit time and can yield probabilities whose argmax disagrees with predict.

open as a page

A fraud model shows 0.97 ROC-AUC but poor live precision — which sklearn.metrics calls do you reach for?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Score the ranking with average_precision_score, then call precision_recall_curve to choose an operating threshold: it returns precision and recall arrays one element longer than thresholds. Rebuild confusion_matrix at that threshold instead of trusting predict()'s default 0.5.

open as a page

In scikit-learn, which CV splitter keeps all of one patient's rows in a single fold?

level: seniorimportance: should knowfreq 48%

basics

~20 s

GroupKFold, given a groups array of patient IDs, guarantees no group's rows appear in both the training and test side of a split. The groups array is passed at fit time — GroupKFold(n_splits=5).split(X, y, groups) or search.fit(X, y, groups=ids).

open as a page

In scikit-learn, how do you run nested cross-validation around a GridSearchCV?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Pass the unfitted search as the estimator to an outer cross-validation: cross_val_score(GridSearchCV(est, grid, cv=inner), X, y, cv=outer). Each outer fold tunes on its own training portion and is scored on data no tuning decision ever saw.

open as a page

showing 1–30 of 36