skip to content

How do you address a Pipeline step's parameters in a scikit-learn GridSearchCV param_grid?

level: middleimportance: must knowfreq 68%

answer

  1. step name, separator, parameter name
  2. two underscores, not one
  3. the step itself is also a parameter
  4. list of dicts for per-estimator grids
  5. get_params().keys() lists every legal key

basics

~20 s

Use the step name, a double underscore, then the parameter name: "clf__C": [0.1, 1.0]. Nesting repeats the pattern for nested estimators, and using a step's bare name as the key replaces the whole step with another estimator or with "passthrough".

solid answer

~40 s

A `Pipeline` exposes its steps through `get_params`, which is why the double-underscore convention works: `{"pca__n_components": [5, 10], "clf__C": [0.1, 1.0]}` reaches into the steps named `pca` and `clf`. Nesting composes — a transformer inside a `FeatureUnion` inside a pipeline is `features__pca__n_components`. Because the steps themselves are parameters, a bare step name as the key swaps the object: `{"clf": [LogisticRegression(), RandomForestClassifier()]}` searches over estimators, and `{"pca": ["passthrough"]}` disables a step for those candidates. Passing a list of dicts lets each block carry the parameters that only apply to its own estimator. Explicit `Pipeline` step names are yours to choose (they must be unique and contain no `__`); `make_pipeline` generates them as the lowercased class name, so the key becomes `logisticregression__C`. After the search, `best_estimator_` is a pipeline refitted on all the data.

code

python · 28 lines
python
from sklearn.decomposition import PCA
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

pipe = Pipeline([
    ("scaler", StandardScaler()),
    ("pca", PCA()),
    ("clf", LogisticRegression(max_iter=1000)),
])

param_grid = [
    {
        "pca__n_components": [5, 10],
        "clf": [LogisticRegression(max_iter=1000)],
        "clf__C": [0.1, 1.0],
    },
    {
        "pca": ["passthrough"],
        "clf": [RandomForestClassifier(random_state=0)],
        "clf__n_estimators": [100, 300],
    },
]

search = GridSearchCV(pipe, param_grid, cv=5)
print(sorted(k for k in pipe.get_params() if "__" in k)[:5])

go deeper

for a junior

Memorise the spelling: stepname__paramname, two underscores, and the step name is whatever you wrote in the pipeline. Be able to write a small grid tuning one preprocessing parameter and one model parameter.

for a middle

Explain why it works — the pipeline exposes steps through get_params, and set_params splits on the separator and recurses — and show nesting plus swapping a whole step or setting it to "passthrough".

for a senior

Demonstrate practical grid design: a list of dicts so estimator-specific parameters do not multiply out, awareness that every candidate refits the whole pipeline per fold, and the habit of inspecting get_params().keys() rather than guessing at names.

for a principal

Frame it as search-budget strategy — which parameters genuinely deserve grid points, when swapping whole steps beats tuning one, and when to move from an exhaustive grid to randomised or successive-halving search so the compute goes where the variance is.

## Why double underscore Every scikit-learn estimator implements `get_params(deep=True)`, which returns not only its own hyperparameters but also, for any parameter that is itself an estimator, that estimator's parameters prefixed with `<name>__`. `Pipeline` stores its steps as the `steps` parameter and additionally exposes each step under its own name. So a pipeline with steps named `scaler`, `pca` and `clf` reports parameters including `scaler`, `pca`, `clf`, `scaler__with_mean`, `pca__n_components`, `clf__C`, and so on. `GridSearchCV` does not parse your grid specially. For each candidate it calls `clone(estimator).set_params(**candidate)`, and `set_params` splits each key on the first `__`, looks up the named sub-estimator, and recurses. That is the whole mechanism — the same one behind `pipe.set_params(clf__C=10)` on a plain pipeline. ## The three shapes of key **Reach into a step.** `"clf__C": [0.01, 0.1, 1, 10]` sets the classifier's `C`. This is the everyday case. **Reach through several levels.** If a step is itself a composite — a `FeatureUnion` named `features` holding a `PCA` named `pca` — the key is `features__pca__n_components`. There is no limit to the depth; each `__` descends one level. **Replace the step.** Because `clf` is itself a parameter, `"clf": [LogisticRegression(max_iter=1000), RandomForestClassifier()]` searches over which estimator occupies that slot. The same trick applied to a transformer step with the sentinel string `"passthrough"` makes the step a no-op for those candidates, which is how you ask "is this preprocessing step worth having at all?". Note the sentinel is `"passthrough"` for a Pipeline step; `"drop"` is what `FeatureUnion` and `ColumnTransformer` accept for removing a branch. ## Lists of dicts A single dict takes the Cartesian product of everything in it, which breaks when a parameter only makes sense for one candidate estimator — `n_estimators` is meaningless for `LogisticRegression`. `param_grid` therefore accepts a **list of dicts**, each explored independently: ``` [ {"clf": [LogisticRegression()], "clf__C": [0.1, 1]}, {"clf": [RandomForestClassifier()], "clf__n_estimators": [100, 300]} ] ``` Each block is a separate grid; the search unions their candidate lists. The same list-of-dicts form works for `RandomizedSearchCV`'s distributions. ## Where the names come from With explicit `Pipeline([("scaler", StandardScaler()), ...])` you choose the names. Two rules are enforced at construction: names must be unique, and a name may not contain `__`, because that character sequence is the parameter separator and would make lookups ambiguous. Violating either raises a `ValueError` from the pipeline's name validation. `make_pipeline(StandardScaler(), PCA(), LogisticRegression())` generates names automatically as the lowercased class name — `standardscaler`, `pca`, `logisticregression` — and disambiguates repeats by appending `-1`, `-2`. That is convenient for scripts and awkward for grids, because your keys become `logisticregression__C` and they change if you swap the class. For anything you intend to tune, prefer explicit short names. You can always print `pipe.get_params().keys()` to see every legal key, which is the fastest way to fix the error scikit-learn raises for a bad one: an `InvalidParameterError`/`ValueError` complaining that the parameter is invalid for the estimator `Pipeline`, usually caused by a mistyped step name, a single underscore, or a parameter that belongs to a different step. ## After the search `GridSearchCV` with the default `refit=True` fits the best candidate on the entire dataset and exposes it as `best_estimator_` — a complete `Pipeline`, so `search.predict(X_new)` runs preprocessing and prediction together and `search.best_estimator_.named_steps["clf"].coef_` reaches the fitted classifier. `best_params_` echoes back the winning dict in the same double-underscore spelling you supplied, which is handy for logging. ## Cost and the interaction with folds Every candidate refits **the whole pipeline** in every fold, because that is what makes the score honest. A grid of 20 candidates at `cv=5` is 100 full pipeline fits, including 100 fits of an expensive vectorizer or decomposition that may not even vary across candidates. That is precisely the situation the pipeline's `memory` argument exists to relieve, and it is the natural follow-up an interviewer will reach for once you have the naming right.

  • How would you search over whether a preprocessing step should exist at all?
    Use the step's bare name as a grid key and include the sentinel string: `{"pca": ["passthrough", PCA(n_components=10)]}`. Candidates with `"passthrough"` forward their input unchanged, so the search directly compares having the step against not having it, scored under the same folds as every other candidate.
  • What step names does make_pipeline generate, and why can that bite you in a grid?
    It lowercases the class name — `standardscaler`, `logisticregression` — appending `-1`, `-2` to disambiguate repeats. Your grid keys then encode the class, so swapping `LogisticRegression` for `SVC` silently invalidates every key that referenced it. Explicit `Pipeline` names like `clf` are stable and shorter, which is why tuned pipelines usually spell out their steps.
  • A grid raises an error saying the parameter is invalid for estimator Pipeline. How do you debug it?
    Print `pipe.get_params().keys()` and compare. Nearly always it is a mistyped or renamed step, a single underscore instead of two, or a parameter assigned to the wrong step. The error names the offending key, and the parameter list is the authoritative set of what `set_params` will accept.
  • What is best_estimator_ after searching over a pipeline?
    With the default `refit=True`, it is the winning pipeline refitted on the whole dataset — every transformer plus the final estimator, not just the classifier. Calling `search.predict` delegates to it, so preprocessing travels with the model, and `best_estimator_.named_steps[name]` reaches any fitted step for inspection.

saying these in an interview costs you the question

  • Using a single underscore between step and parameter
  • Naming a pipeline step something containing a double underscore
  • Putting incompatible estimator parameters in one dict
  • Believing you can only tune the final estimator
  • Thinking best_estimator_ is just the classifier

context