skip to content

Which preprocessing belongs inside a persisted scikit-learn Pipeline versus upstream ETL?

level: principalimportance: should knowfreq 30%

answer

  1. learned state travels with the model
  2. two implementations will diverge silently
  3. the pickle stores class paths, not code
  4. pin the library version beside the artifact
  5. column names become part of the contract

basics

~20 s

Anything with state learned from training data — imputer statistics, encoder categories, scaler means, vectorizer vocabularies, selection masks — must live inside the fitted Pipeline so training and serving share one artifact. Stateless business joins and label construction stay upstream.

solid answer

~50 s

The dividing line is **learned state**. If a transform's behaviour depends on statistics estimated from the training set, it must be a step of the pipeline you persist, because otherwise serving code has to reproduce those numbers by hand and will eventually drift from them. That covers `SimpleImputer` medians, `StandardScaler` means, `OneHotEncoder` categories, vectorizer vocabularies and feature-selection masks. Stateless work — joins, deduplication, label definition, unit conversions with fixed constants — is cheaper and more debuggable upstream, though a row-wise deterministic transform can sit in a `FunctionTransformer` inside the pipeline for parity. The costs of the in-pipeline choice are real: `joblib.dump` writes a pickle whose custom transformer classes must be importable at the same module path when it loads, the scikit-learn version is not guaranteed compatible across releases (a mismatch raises `InconsistentVersionWarning` since 1.3), and if you fitted on a DataFrame the column names in `feature_names_in_` become part of the request contract.

code

python · 21 lines
python
import joblib
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import FunctionTransformer, StandardScaler


def log1p_features(X):          # module-level, so it pickles; a lambda would not
    return np.log1p(X)


pipe = Pipeline([
    ("log", FunctionTransformer(log1p_features, feature_names_out="one-to-one")),
    ("scale", StandardScaler()),
    ("clf", LogisticRegression(max_iter=1000)),
])

X = np.array([[1.0, 2.0], [3.0, 4.0], [5.0, 6.0], [7.0, 8.0]])
y = np.array([0, 1, 0, 1])
pipe.fit(X, y)
joblib.dump(pipe, "model.joblib")

go deeper

for a junior

Know the basic rule: whatever the model learned from training data — fill values, category lists, scaling numbers — must be saved with the model, not re-derived in the serving code.

for a middle

Be able to justify the split with the learned-state test, and know that persisting a pipeline means persisting a pickle whose custom classes must be importable and whose library version should be pinned.

for a senior

Show that you have debugged the failure modes: an unseen category at serve time, a reordered column, a version-mismatched pickle, a per-request transform that is too slow — and where each of those is fixed.

for a principal

Own the boundary as an architecture and ownership decision: what the model artifact is allowed to contain, what the feature platform owns, how divergence between two implementations is detected, and how versions of library, artifact and feature contract are pinned together.

## The question behind the question Every team eventually draws a line between "the data platform prepares features" and "the model object transforms features". Put the line in the wrong place and you get one of two failures: training/serving skew, where two implementations of the same transform diverge; or an unmaintainable model artifact that has swallowed business logic it should never have owned. ## The rule: learned state goes inside The test that survives contact with real systems is whether the transform has **parameters estimated from training data**. Inside the pipeline, always: - imputation statistics (a median, a most-frequent category) - scaling parameters (means, variances, ranges, quantiles) - categorical vocabularies — an encoder's known category set, and its policy for unseen values at serve time - text vectoriser vocabularies and document frequencies - learned projections and feature-selection masks - anything using `y`, which additionally must never be fitted outside a cross-validation fold The reason is not tidiness. These numbers *are* part of the model. A scaler mean recomputed at serve time from the serving batch puts inference in different units than training, and the failure is silent — no exception, just worse predictions, often worst on the smallest batches. Upstream, usually: - joins, aggregation windows and entity resolution that need the warehouse - label construction and any time-based cutoff logic - deduplication, hard validity filtering, PII handling - anything expensive that is shared across many models and should be computed once The grey zone is stateless row-wise maths — a log, a ratio, a clip. It leaks nothing wherever it lives. Putting it in a `FunctionTransformer` inside the pipeline buys parity for free; leaving it upstream buys debuggability. Choose per team, but choose once. ## What the in-pipeline choice costs **Pickle fragility.** `joblib.dump(pipe, "model.joblib")` serialises by reference for classes: the file records that a step is `myproject.features.RatioTransformer`, not the class's code. At load time that import path must exist and mean the same thing. Rename the module, restructure the package, or deploy a serving image that does not ship the training package, and the load fails. Practical consequences: define custom transformers in a small, stable, importable module that both training and serving depend on; never define one in a notebook or as a closure; and never pass a lambda to `FunctionTransformer` — module-level named functions pickle, lambdas do not. **Version coupling.** Scikit-learn does not promise pickle compatibility across versions. Since 1.3 it detects the mismatch and raises `InconsistentVersionWarning`, telling you which version wrote the file — a warning, not a guarantee of correctness. So the serving image should pin the exact scikit-learn (and numpy/scipy) versions the artifact was fitted with, and the pin belongs next to the artifact, not in a README. **Contract surface.** Fit on a DataFrame and estimators record `feature_names_in_`; predict on a frame whose columns differ in name or order and scikit-learn complains. Fit on a raw array and you get no such check — column order becomes an undocumented, silently-violable contract. Fitting on named columns is the safer default precisely because it makes the contract explicit, and `set_output(transform="pandas")` keeps names flowing through the intermediate steps for debugging. **Latency and shape.** A transform that is fine over a training matrix may be poor per request. A vectoriser is fast; an expensive aggregation is not. If a step is slow per row, that is an argument for precomputing it upstream and passing it in as a feature — accepting that you now own keeping the two implementations aligned, which is a cost you should name out loud rather than discover. ## Unseen data at serve time This is where the boundary is tested. An encoder meets a category it never saw; a numeric column arrives null when training had none. Those policies must be configured on the estimators inside the pipeline — the encoder's unknown-value handling, an imputer covering columns that were complete in training — because serving code that patches the input before the pipeline is a second implementation of the model's assumptions. Decide the policy at fit time, encode it in the object, and let one artifact answer for it. ## The organisational angle The pipeline boundary is also a team boundary. Everything inside it is owned by whoever trains the model and moves at model-release cadence. Everything outside is owned by the data platform and moves at its own. Learned state inside the artifact keeps the fast-changing, model-specific, leak-prone part in one versioned object that can be rolled back atomically; shared, expensive, cross-model computation upstream keeps it from being recomputed per model. When someone proposes moving a fitted statistic upstream "for performance", the question to ask is who will notice when the two copies of that number diverge — because nothing will raise when they do.

  • What actually breaks when you load a pipeline pickle under a different scikit-learn version?
    Nothing is guaranteed. Since 1.3 scikit-learn detects the mismatch and raises `InconsistentVersionWarning` naming the writing version, but loading may still succeed with subtly different behaviour, or fail outright if an internal attribute changed. Treat the library version as part of the artifact: pin it in the serving image and record it alongside the model file.
  • How do you put a custom row-wise transform into a pipeline without breaking serialisation?
    Wrap a module-level named function in `FunctionTransformer`, or write a small class in a stable module that both training and serving import. Pickle stores the import path, not the code, so lambdas and notebook-defined classes fail to load. Keep that module tiny and dependency-light so the serving image can ship it without the training stack.
  • Why prefer fitting on a DataFrame rather than a raw array for a pipeline you will serve?
    Fitting on named columns records `feature_names_in_`, so a request with renamed or reordered columns is caught instead of silently scored. With a bare array, column order is an undocumented contract that any upstream change can violate without error. `set_output(transform="pandas")` keeps names flowing through intermediate steps, which also makes debugging a mid-pipeline matrix far easier.
  • When is it legitimate to move a fitted statistic out of the pipeline and upstream?
    When the feature is genuinely shared across many models and recomputing it per model is prohibitive — a company-wide aggregate, say. Then it becomes a versioned upstream artifact with its own contract and its own drift monitoring. The cost you are accepting is two places that must agree; name it explicitly, and make the divergence detectable, since nothing will raise.

saying these in an interview costs you the question

  • Serving code recomputes the scaler mean per request
  • Custom transformers defined in a notebook or as lambdas
  • Assuming pickles load safely across scikit-learn versions
  • Treating column order as guaranteed by the caller
  • Patching unseen categories before the pipeline instead of configuring the encoder

context