When does Pipeline(memory=...) in scikit-learn actually save time, and what does it cost?
answer
- repeated identical work, not less work
- only the transformer steps are cached
- key covers transformer, params and data
- reuse across candidates within a fold
- disk grows until you clear it
basics
~20 sSetting memory on a Pipeline caches fitted transformers on disk, keyed by the transformer, its parameters and its input. It pays off during a search where many candidates share an identical, expensive prefix — and buys nothing when every candidate changes an early step.
solid answer
~40 s`Pipeline(steps, memory=...)` accepts a directory path or a `joblib.Memory` instance and caches the result of fitting each **transformer** step — the final estimator is never cached. The cache key is a joblib hash of the transformer, its parameters and the data it was given, so during a `GridSearchCV` that only varies `clf__C`, an expensive `PCA` or `TfidfVectorizer` is fitted once per fold and reused across every `C`. It buys nothing for a single `fit`, and nothing for candidates that change an early step's parameters, since those produce different keys. The costs are real: disk grows until you call `memory.clear()`, hashing large arrays takes time of its own, and because the cached transformers are clones, the instances you passed into the pipeline do not end up carrying the fitted attributes — inspect `search.best_estimator_.named_steps[...]` instead.
code
python · 18 linesfrom joblib import Memory
from sklearn.decomposition import PCA
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV
from sklearn.pipeline import Pipeline
memory = Memory(location="./sklearn_cache", verbose=0)
pipe = Pipeline(
[("pca", PCA(n_components=10)), ("clf", LogisticRegression(max_iter=1000))],
memory=memory,
verbose=True,
)
# PCA is fitted once per fold and reused for all four C values.
search = GridSearchCV(pipe, {"clf__C": [0.01, 0.1, 1.0, 10.0]}, cv=5)
memory.clear(warn=False)go deeper
Know that a Pipeline can take a memory argument that caches fitted transformers to disk so repeated identical work is skipped, and that it is a search-time optimisation rather than something you switch on by default.
Explain the mechanics — a joblib hash over the transformer, its parameters and the input decides a hit, only transformer steps are cached, and the win depends on candidates sharing an identical prefix.
Show operational judgment: decide when hashing a large matrix costs more than refitting a cheap scaler, manage the cache directory and clear it, and inspect fitted steps through the fitted pipeline rather than the instances you constructed.
Treat it as one lever among several for search cost. Weigh caching against shrinking the grid, randomised or halving search, and staged tuning, and set the team convention for where caches live and how determinism in custom transformers is guaranteed.
## The problem it solves A grid search refits the entire pipeline for every candidate in every fold. That is required for honest scoring, but it can be enormously wasteful. Consider a text pipeline: `TfidfVectorizer` → `TruncatedSVD` → `LogisticRegression`, searching eight values of the classifier's `C` with `cv=5`. That is 40 fits of the pipeline, and therefore 40 fits of the vectorizer and the decomposition — even though, within a given fold, all eight candidates feed the classifier exactly the same matrix. Only five distinct vectorizer fits were ever needed. ## What memory does `Pipeline(steps=[...], memory=cache)` where `cache` is a directory path or a `joblib.Memory` object wraps the fit of each **transformer** step in joblib's on-disk memoisation. Before fitting a transformer, joblib hashes the transformer object with its parameters together with the input data and the `y` it was given; if that hash is already on disk, the fitted transformer and its output are loaded instead of recomputed. Two scoping facts matter. First, only the intermediate transformer steps are cached — the final estimator is fitted every time, which is right, since it is usually the thing being varied. Second, the key includes the data, so different folds are different keys: caching gives reuse *across candidates within a fold*, not across folds. ## When it actually pays The benefit is proportional to how much of the pipeline prefix is *invariant* across candidates and how expensive that prefix is. - **Big win:** an expensive prefix with a search that only varies the last step. Vectorisation, decomposition or heavy feature extraction fitted once per fold instead of once per candidate. - **Partial win:** a grid that varies an early step over a few values and a late step over many. Candidates sharing the same early-step parameters share the cached prefix — for example, four `C` values under each of two `n_components` values means two prefix fits per fold, not eight. - **No win:** a single `fit` or `cross_val_score` with no repeated prefix; or a grid where every candidate perturbs the very first step, so no two keys collide. Cheap transformers such as `StandardScaler` are usually not worth caching at all — hashing the matrix can cost more than refitting. ## What it costs **Disk.** The cache stores fitted transformers and their outputs. A dense transformed matrix per fold per distinct prefix configuration adds up quickly on a large dataset, and nothing evicts it automatically. `Memory.clear()` empties it; in a CI or shared machine, point `location` at a scratch directory you control and clear it deliberately. **Hashing time.** Every cache lookup hashes the input array. For very large arrays that is real CPU time, paid on hits and misses alike, and it can erase the saving for a fast transformer. **Inspection.** The pipeline clones transformers when caching, so the transformer instances you constructed and passed in are not the objects that get fitted. Reading fitted attributes off the original instance therefore gives nothing useful — go through the fitted pipeline instead: `search.best_estimator_.named_steps["pca"].components_`. **Correctness of the key.** Memoisation is only safe if the fit is a pure function of (transformer, params, data). A custom transformer that reads a global, a file, or a random source without a fixed `random_state` can be handed a stale cached result. That is a genuine hazard with hand-written transformers and an argument for keeping them deterministic. ## Reading the win `Pipeline(..., verbose=True)` prints per-step timings so you can see cached steps collapse to near zero, and `GridSearchCV(..., verbose=...)` reports total fit time. Measure before and after rather than assuming — the answer depends entirely on the shape of your grid. ## The alternatives Caching is not the only lever, and often not the first one. Restructuring the search so the expensive prefix is fitted outside the tuned estimator would reintroduce leakage and is not acceptable. But you can shrink the grid, move to a randomised or halving search so fewer candidates reach the expensive configurations, subsample for a coarse pass, or split the search into two stages — tune the prefix first with a fixed cheap head, then tune the head. Caching composes with all of those; it is a way to stop paying twice for identical work, not a way to reduce the work itself.
- Is the final estimator ever cached?No — only the transformer steps are memoised. That is the sensible split, because the final estimator is normally the thing whose parameters the search is varying, so its fit would miss the cache every time anyway. Caching the prefix is where all the reuse lives.
- If the grid tunes an early transformer's parameters as well, does caching still help?Partially. The cache key includes the transformer's parameters, so each distinct early-step configuration gets its own entry, and every later-step candidate sharing that configuration reuses it. Four classifier values under two decomposition settings means two prefix fits per fold instead of eight — a real saving, just smaller than when the prefix is fixed.
- Why can inspecting the transformer instance you passed in show no fitted attributes when caching is on?Because the pipeline clones transformers when memoisation is enabled, so the object that actually gets fitted is a copy, and a cache hit restores a deserialised copy rather than mutating yours. Read fitted state off the fitted pipeline — `best_estimator_.named_steps[name]` — which is the correct habit regardless of caching.
- What can make a cached fit silently wrong?Memoisation assumes fitting is a pure function of the transformer, its parameters and the input. A custom transformer that reads a global, a clock, an external file, or an unseeded random source can be served a stale result whose key no longer describes what it would compute. Keep custom transformers deterministic and seed anything stochastic with `random_state`.
saying these in an interview costs you the question
- Caching speeds up any pipeline fit
- The cache reuses work across different folds
- The final estimator gets cached too
- The cache cleans itself up automatically
- Hashing the input array is free