skip to content

Pipelines

A Pipeline chains transformers and a final estimator into one object you can fit, search over, and serialise as a unit. Interviewers care most about why it exists: keeping every transform inside the cross-validation fold so leakage cannot creep in.

on this pageshow

questions

6

In scikit-learn, what happens at each step when you call Pipeline.fit() then predict()?

level: juniorimportance: must knowfreq 72%

answer

  1. fit learns, predict only applies
  2. fit_transform down the chain, then fit
  3. transformers keep their learned state
  4. last step decides which methods exist
  5. named_steps and slicing reach inside

basics

~20 s

fit runs fit_transform on every step except the last, feeding each output into the next, then fit on the final estimator. predict runs transform only on those same intermediate steps — never fit again — and calls predict on the final estimator.

solid answer

~40 s

A `Pipeline` is a list of `(name, estimator)` pairs where every step but the last must be a transformer (it needs `fit` and `transform`). `pipe.fit(X, y)` walks the steps in order calling `fit_transform` and passing each result forward, then calls `fit` on the last step with the transformed matrix and `y`. `pipe.predict(X)` walks the same intermediate steps calling `transform` only, then `predict` on the last step — the transformers keep the state they learned during `fit`. The pipeline forwards whatever methods the final estimator has: `predict_proba`, `decision_function` and `score` exist only if the last step defines them, and if the last step is itself a transformer the pipeline exposes `transform` instead of `predict`. Fitted steps are reachable through `named_steps["clf"]` or by index, and slicing like `pipe[:-1]` gives a sub-pipeline of the preprocessing.

code

python · 19 lines
python
import numpy as np
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

pipe = Pipeline([
    ("impute", SimpleImputer(strategy="median")),
    ("scale", StandardScaler()),
    ("clf", LogisticRegression()),
])

X = np.array([[1.0, 2.0], [np.nan, 3.0], [4.0, 5.0], [2.0, np.nan]])
y = np.array([0, 1, 1, 0])

pipe.fit(X, y)                     # fit_transform on impute + scale, then fit on clf
print(pipe.predict(X))             # transform on impute + scale, then predict on clf
print(pipe.named_steps["clf"].coef_)
print(pipe[:-1].transform(X))      # slicing yields the preprocessing sub-pipeline

go deeper

for a junior

Be able to state the two passes out loud: fit learns on each step in order, predict only applies what was learned. Know that every step but the last must be a transformer and that named_steps gets you the fitted objects.

for a middle

Explain the details — fit_transform is one fused call per intermediate step, the pipeline forwards only the methods the last step defines, and slicing produces a sub-pipeline you can transform with for debugging.

for a senior

Show why the object matters operationally: one artifact carries every transform, so training and serving cannot drift, and any step is inspectable after the fact when a prediction looks wrong.

for a principal

Own the composition question — how deep to nest pipelines, when a preprocessing chain should be its own reusable unit shared across models, and what step-naming discipline keeps tuning and serving code stable as models are swapped.

## What a Pipeline is `Pipeline(steps=[("impute", SimpleImputer()), ("scale", StandardScaler()), ("clf", LogisticRegression())])` takes an ordered list of `(name, estimator)` tuples. The names are strings you choose; they must be unique and must not contain a double underscore, because that sequence is reserved as the parameter separator used when addressing a step's hyperparameters. `make_pipeline(...)` builds the same object with names generated from the lowercased class names. The contract on the steps: **every step except the last must be a transformer**, meaning it implements `fit` and `transform`. The last step can be anything — a classifier, a regressor, another transformer, or the sentinel string `"passthrough"`. Construction validates this, so passing a classifier in the middle fails immediately rather than at fit time. ## fit `pipe.fit(X, y)` iterates over the steps in order. For each intermediate step it calls `fit_transform(X_current, y)` — one call, because many transformers implement a fused version that is cheaper than `fit` followed by `transform` — and the returned array becomes the input to the next step. After the last transformer, the final estimator's `fit(X_transformed, y)` is called. The return value is the pipeline itself, so `pipe.fit(X, y).predict(X_test)` chains. Each step retains its learned state on the step object: the imputer's medians, the scaler's mean and scale, the classifier's coefficients. Nothing is thrown away, which is exactly what makes the second phase possible. ## predict `pipe.predict(X)` walks the same intermediate steps but calls `transform(X_current)` — never `fit` or `fit_transform`. The imputer fills with the medians it learned during training; the scaler applies the training mean and scale. Then the final estimator's `predict` runs on the transformed matrix. This asymmetry is the single most important thing about the object: **fit learns, predict applies**, and the pipeline enforces it for every step at once rather than leaving it to you to remember per transformer. ## Method forwarding The pipeline does not invent methods. It exposes: - `predict`, `predict_proba`, `predict_log_proba`, `decision_function` and `score` **if and only if** the final estimator has them. Asking a pipeline ending in `LinearRegression` for `predict_proba` fails, as it should. - `transform` if the final step is a transformer, in which case the pipeline is a pure preprocessing chain with no prediction at all. - `fit_transform`, which fits everything and returns the last step's transform output. - `inverse_transform`, if every step supports it. Because the pipeline satisfies the estimator contract itself, it can be nested inside another pipeline, dropped into `cross_val_score`, or handed to a search — it is just another estimator. ## Reaching inside Four ways to get at a step, all useful: - `pipe.named_steps["clf"]` (also attribute-style, `pipe.named_steps.clf`) returns the fitted step object. - `pipe["clf"]` and `pipe[-1]` index by name or position. - `pipe[:-1]` **slices**, returning a new pipeline of the first steps — handy for `pipe[:-1].transform(X)` to see what the model actually receives. - `pipe.steps` is the raw list of tuples. So `pipe.named_steps["clf"].coef_` gets the fitted coefficients, and `pipe.named_steps["scale"].mean_` gets the learned means. ## The passthrough sentinel and empty steps Any step may be set to `"passthrough"`, which makes it forward its input unchanged. That is mostly used for hyperparameter search — asking whether a step earns its place — but it also lets you keep a stable step layout across configurations, which keeps parameter names stable. ## Common first mistakes Calling `fit_transform` on the pipeline when the last step is a classifier fails, because a classifier has no `transform`. Calling `pipe.fit(X_train)` without `y` fails once a supervised final step needs labels. And expecting the intermediate steps to refit during `predict` — some people assume the scaler re-standardises the new batch using that batch's own mean — is the misconception that produces wildly wrong predictions on small serving batches: the pipeline deliberately does not do that, because the model was trained in the training data's units. ## Why it matters beyond convenience Because the whole chain is one estimator with one `fit`, the pipeline is the natural unit to cross-validate, to search over, and to serialise. Every transform that was applied during training is applied identically at prediction time, from the same object, with no second script to keep in sync.

  • What can you do with a pipeline whose last step is a transformer rather than a model?
    It becomes a pure preprocessing chain: the pipeline exposes `transform` and `fit_transform` instead of `predict`. That is useful as a reusable feature-building block you can nest as a single step inside a larger pipeline, or slice off with `pipe[:-1]` to inspect exactly what matrix the final estimator sees.
  • How do you read a fitted transformer's learned attributes out of a pipeline?
    Index into it: `pipe.named_steps["scale"].mean_`, or equivalently `pipe["scale"]` or `pipe[1]`. The step objects held in the pipeline are the same ones you passed in and they carry their fitted attributes, so anything documented on the transformer is reachable after `fit`.
  • Why does the pipeline refuse predict_proba for some final estimators?
    Because it only forwards methods the last step actually defines. A pipeline ending in `LinearRegression` or `SVC(probability=False)` has no probability method to delegate to, and scikit-learn surfaces that rather than fabricating one. The same rule governs `decision_function`, `score` and `transform`.
  • Does the scaler recompute statistics on the batch passed to predict?
    No. `predict` calls `transform` only, so the scaler applies the mean and scale learned during `fit`. Recomputing per batch would put serving data in different units than the model was trained in, and would make predictions depend on which other rows happened to be in the same request.

saying these in an interview costs you the question

  • Thinking transformers refit during predict
  • Putting a classifier in a non-final step
  • Expecting predict_proba regardless of the final estimator
  • Assuming the pipeline only wraps the model, not the transforms
  • Believing you cannot access individual fitted steps

context

open as a page

How do you address a Pipeline step's parameters in a scikit-learn GridSearchCV param_grid?

level: middleimportance: must knowfreq 68%

basics

~20 s

Use the step name, a double underscore, then the parameter name: "clf__C": [0.1, 1.0]. Nesting repeats the pattern for nested estimators, and using a step's bare name as the key replaces the whole step with another estimator or with "passthrough".

open as a page

In scikit-learn, why must a scaler live inside the Pipeline passed to cross_val_score?

level: middleimportance: must knowfreq 78%

basics

~20 s

cross_val_score refits whatever estimator you hand it on each training fold. A scaler fitted outside it has already seen every fold's held-out rows, so its means and variances encode validation data and the reported score comes out optimistically biased.

open as a page

What does FeatureUnion do in scikit-learn, and when is it the wrong tool?

level: middleimportance: should knowfreq 34%

basics

~20 s

FeatureUnion fits several transformers on the same input in parallel and horizontally concatenates their outputs into one feature matrix. It is the wrong tool when each transformer should see a different subset of columns — that is ColumnTransformer's job.

open as a page

When does Pipeline(memory=...) in scikit-learn actually save time, and what does it cost?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Setting memory on a Pipeline caches fitted transformers on disk, keyed by the transformer, its parameters and its input. It pays off during a search where many candidates share an identical, expensive prefix — and buys nothing when every candidate changes an early step.

open as a page

Which preprocessing belongs inside a persisted scikit-learn Pipeline versus upstream ETL?

level: principalimportance: should knowfreq 30%

basics

~20 s

Anything with state learned from training data — imputer statistics, encoder categories, scaler means, vectorizer vocabularies, selection masks — must live inside the fitted Pipeline so training and serving share one artifact. Stateless business joins and label construction stay upstream.

open as a page