How do you address a Pipeline step's parameters in a scikit-learn GridSearchCV param_grid?
answer
- step name, separator, parameter name
- two underscores, not one
- the step itself is also a parameter
- list of dicts for per-estimator grids
- get_params().keys() lists every legal key
basics
~20 sUse the step name, a double underscore, then the parameter name: "clf__C": [0.1, 1.0]. Nesting repeats the pattern for nested estimators, and using a step's bare name as the key replaces the whole step with another estimator or with "passthrough".
solid answer
~40 sA `Pipeline` exposes its steps through `get_params`, which is why the double-underscore convention works: `{"pca__n_components": [5, 10], "clf__C": [0.1, 1.0]}` reaches into the steps named `pca` and `clf`. Nesting composes — a transformer inside a `FeatureUnion` inside a pipeline is `features__pca__n_components`. Because the steps themselves are parameters, a bare step name as the key swaps the object: `{"clf": [LogisticRegression(), RandomForestClassifier()]}` searches over estimators, and `{"pca": ["passthrough"]}` disables a step for those candidates. Passing a list of dicts lets each block carry the parameters that only apply to its own estimator. Explicit `Pipeline` step names are yours to choose (they must be unique and contain no `__`); `make_pipeline` generates them as the lowercased class name, so the key becomes `logisticregression__C`. After the search, `best_estimator_` is a pipeline refitted on all the data.
code
python · 28 linesfrom sklearn.decomposition import PCA
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
pipe = Pipeline([
("scaler", StandardScaler()),
("pca", PCA()),
("clf", LogisticRegression(max_iter=1000)),
])
param_grid = [
{
"pca__n_components": [5, 10],
"clf": [LogisticRegression(max_iter=1000)],
"clf__C": [0.1, 1.0],
},
{
"pca": ["passthrough"],
"clf": [RandomForestClassifier(random_state=0)],
"clf__n_estimators": [100, 300],
},
]
search = GridSearchCV(pipe, param_grid, cv=5)
print(sorted(k for k in pipe.get_params() if "__" in k)[:5])go deeper
Memorise the spelling: stepname__paramname, two underscores, and the step name is whatever you wrote in the pipeline. Be able to write a small grid tuning one preprocessing parameter and one model parameter.
Explain why it works — the pipeline exposes steps through get_params, and set_params splits on the separator and recurses — and show nesting plus swapping a whole step or setting it to "passthrough".
Demonstrate practical grid design: a list of dicts so estimator-specific parameters do not multiply out, awareness that every candidate refits the whole pipeline per fold, and the habit of inspecting get_params().keys() rather than guessing at names.
Frame it as search-budget strategy — which parameters genuinely deserve grid points, when swapping whole steps beats tuning one, and when to move from an exhaustive grid to randomised or successive-halving search so the compute goes where the variance is.
## Why double underscore Every scikit-learn estimator implements `get_params(deep=True)`, which returns not only its own hyperparameters but also, for any parameter that is itself an estimator, that estimator's parameters prefixed with `<name>__`. `Pipeline` stores its steps as the `steps` parameter and additionally exposes each step under its own name. So a pipeline with steps named `scaler`, `pca` and `clf` reports parameters including `scaler`, `pca`, `clf`, `scaler__with_mean`, `pca__n_components`, `clf__C`, and so on. `GridSearchCV` does not parse your grid specially. For each candidate it calls `clone(estimator).set_params(**candidate)`, and `set_params` splits each key on the first `__`, looks up the named sub-estimator, and recurses. That is the whole mechanism — the same one behind `pipe.set_params(clf__C=10)` on a plain pipeline. ## The three shapes of key **Reach into a step.** `"clf__C": [0.01, 0.1, 1, 10]` sets the classifier's `C`. This is the everyday case. **Reach through several levels.** If a step is itself a composite — a `FeatureUnion` named `features` holding a `PCA` named `pca` — the key is `features__pca__n_components`. There is no limit to the depth; each `__` descends one level. **Replace the step.** Because `clf` is itself a parameter, `"clf": [LogisticRegression(max_iter=1000), RandomForestClassifier()]` searches over which estimator occupies that slot. The same trick applied to a transformer step with the sentinel string `"passthrough"` makes the step a no-op for those candidates, which is how you ask "is this preprocessing step worth having at all?". Note the sentinel is `"passthrough"` for a Pipeline step; `"drop"` is what `FeatureUnion` and `ColumnTransformer` accept for removing a branch. ## Lists of dicts A single dict takes the Cartesian product of everything in it, which breaks when a parameter only makes sense for one candidate estimator — `n_estimators` is meaningless for `LogisticRegression`. `param_grid` therefore accepts a **list of dicts**, each explored independently: ``` [ {"clf": [LogisticRegression()], "clf__C": [0.1, 1]}, {"clf": [RandomForestClassifier()], "clf__n_estimators": [100, 300]} ] ``` Each block is a separate grid; the search unions their candidate lists. The same list-of-dicts form works for `RandomizedSearchCV`'s distributions. ## Where the names come from With explicit `Pipeline([("scaler", StandardScaler()), ...])` you choose the names. Two rules are enforced at construction: names must be unique, and a name may not contain `__`, because that character sequence is the parameter separator and would make lookups ambiguous. Violating either raises a `ValueError` from the pipeline's name validation. `make_pipeline(StandardScaler(), PCA(), LogisticRegression())` generates names automatically as the lowercased class name — `standardscaler`, `pca`, `logisticregression` — and disambiguates repeats by appending `-1`, `-2`. That is convenient for scripts and awkward for grids, because your keys become `logisticregression__C` and they change if you swap the class. For anything you intend to tune, prefer explicit short names. You can always print `pipe.get_params().keys()` to see every legal key, which is the fastest way to fix the error scikit-learn raises for a bad one: an `InvalidParameterError`/`ValueError` complaining that the parameter is invalid for the estimator `Pipeline`, usually caused by a mistyped step name, a single underscore, or a parameter that belongs to a different step. ## After the search `GridSearchCV` with the default `refit=True` fits the best candidate on the entire dataset and exposes it as `best_estimator_` — a complete `Pipeline`, so `search.predict(X_new)` runs preprocessing and prediction together and `search.best_estimator_.named_steps["clf"].coef_` reaches the fitted classifier. `best_params_` echoes back the winning dict in the same double-underscore spelling you supplied, which is handy for logging. ## Cost and the interaction with folds Every candidate refits **the whole pipeline** in every fold, because that is what makes the score honest. A grid of 20 candidates at `cv=5` is 100 full pipeline fits, including 100 fits of an expensive vectorizer or decomposition that may not even vary across candidates. That is precisely the situation the pipeline's `memory` argument exists to relieve, and it is the natural follow-up an interviewer will reach for once you have the naming right.
- How would you search over whether a preprocessing step should exist at all?Use the step's bare name as a grid key and include the sentinel string: `{"pca": ["passthrough", PCA(n_components=10)]}`. Candidates with `"passthrough"` forward their input unchanged, so the search directly compares having the step against not having it, scored under the same folds as every other candidate.
- What step names does make_pipeline generate, and why can that bite you in a grid?It lowercases the class name — `standardscaler`, `logisticregression` — appending `-1`, `-2` to disambiguate repeats. Your grid keys then encode the class, so swapping `LogisticRegression` for `SVC` silently invalidates every key that referenced it. Explicit `Pipeline` names like `clf` are stable and shorter, which is why tuned pipelines usually spell out their steps.
- A grid raises an error saying the parameter is invalid for estimator Pipeline. How do you debug it?Print `pipe.get_params().keys()` and compare. Nearly always it is a mistyped or renamed step, a single underscore instead of two, or a parameter assigned to the wrong step. The error names the offending key, and the parameter list is the authoritative set of what `set_params` will accept.
- What is best_estimator_ after searching over a pipeline?With the default `refit=True`, it is the winning pipeline refitted on the whole dataset — every transformer plus the final estimator, not just the classifier. Calling `search.predict` delegates to it, so preprocessing travels with the model, and `best_estimator_.named_steps[name]` reaches any fitted step for inspection.
saying these in an interview costs you the question
- Using a single underscore between step and parameter
- Naming a pipeline step something containing a double underscore
- Putting incompatible estimator parameters in one dict
- Believing you can only tune the final estimator
- Thinking best_estimator_ is just the classifier