skip to content

When does scikit-learn's fit_transform() differ from fit() then transform()?

level: middleimportance: should knowfreq 48%

answer

  1. one of them is composition of the others
  2. the mixin, not the estimator, supplies it
  3. some transformers only have the joint form
  4. no out-of-sample mapping
  5. training data only

basics

~20 s

Usually not at all: TransformerMixin's default fit_transform is literally fit(X, y).transform(X). Transformers override it when fitting and transforming together is cheaper, or when the embedding is defined only for the fitted samples — sklearn.manifold.TSNE offers fit_transform and no transform.

solid answer

~50 s

`fit_transform` comes from `sklearn.base.TransformerMixin`, whose default implementation is `self.fit(X, y, **fit_params).transform(X)` — fit on this data, then transform *the same* data. So for a `StandardScaler` the two spellings are equivalent, and choosing between them is style. Concrete transformers override the method for two reasons. The first is efficiency: an algorithm that already computes the transformed training data while fitting can return it instead of recomputing. The second is more fundamental — some methods have no out-of-sample mapping at all, so the joint form is the only form. `sklearn.manifold.TSNE` exposes `fit_transform` but no `transform`, because its embedding is optimised jointly for the samples it was given and cannot be applied to unseen points. The rule that follows: `fit_transform` belongs on training data only; new data goes through `transform`, which reuses the statistics learned at fit time.

go deeper

for a junior

Be able to say that fit_transform is fit followed by transform on the same data, and that only training data goes through it — everything else uses transform.

for a middle

Explain that TransformerMixin supplies the default implementation and forwards y, and give the two reasons an estimator overrides it: efficiency, or having no out-of-sample mapping.

for a senior

Show the operational consequence — a transformer without transform cannot sit in a fitted pipeline serving new data, and an overridden fit_transform that diverges from fit().transform() makes pipeline fitting differ from hand-wired fitting.

for a principal

Own the interface question: exposing only the joint form is how the library encodes 'this method has no out-of-sample map', and any in-house transformer should make that property explicit rather than faking a transform.

## The default implementation A scikit-learn transformer implements two methods, `fit` and `transform`. It does not have to implement a third — `fit_transform` arrives by inheriting `sklearn.base.TransformerMixin`, whose default body is essentially: - if `y` is `None`: `return self.fit(X, **fit_params).transform(X)` - otherwise: `return self.fit(X, y, **fit_params).transform(X)` That is the whole story for the majority of transformers. `StandardScaler().fit_transform(X)` and `StandardScaler().fit(X).transform(X)` produce the same array and do the same work. The convenience method exists because fitting and immediately transforming the training set is overwhelmingly the common case. Note the `y` handling: the mixin forwards `y` when you pass it, because some transformers are supervised — a feature selector that scores columns against the target needs `y` during `fit`, and its `transform` then needs only `X`. ## Reason one to override: efficiency Some algorithms compute the transformed training data as a by-product of fitting. Re-running `transform` afterwards would repeat that work. Such estimators override `fit_transform` to return the already-computed result, so the joint call is genuinely cheaper than the two-step spelling even though the output is identical. This is invisible to callers and is a pure optimisation. ## Reason two to override: there is no out-of-sample map This is the conceptually interesting case. Whether `transform` can exist at all depends on the method. A *parametric* transformer learns a small set of parameters from the training data and then applies a formula. `StandardScaler` learns `mean_` and `scale_`; applying them to any new row is trivial, so `transform` is well defined and works on data the estimator has never seen. Same for `PCA`, which learns `components_` and projects. A *non-parametric* embedding optimises the positions of the given points directly against each other. There is no formula to carry over. `sklearn.manifold.TSNE` is the canonical example: its output is a set of coordinates found by minimising a divergence between neighbourhood distributions over exactly those samples. Ask where a new point would land and the question has no answer without redoing the optimisation. So `TSNE` provides `fit_transform` and deliberately provides no `transform` at all. The practical consequence is that you cannot put such a transformer in the middle of a fitted pipeline and push new data through it — at predict time the pipeline calls `transform` on every intermediate step, and the method is not there. ## The rule about test data Because `fit_transform` fits on whatever you hand it, calling it on the test set re-fits the transformer on the test set. The test rows are then scaled by their own mean and standard deviation, or encoded against their own category set, instead of the training-time statistics. Nothing raises; the arrays come out the right shape and the score is simply wrong — often optimistically so, since information from the evaluation data has shaped the transform. The discipline is mechanical: `fit_transform` on training data, `transform` on validation, test and production data. The same asymmetry is why transformers must live inside the object being cross-validated rather than being applied to the whole dataset beforehand. ## The output type: set_output `TransformerMixin` carries one more piece of behaviour. Transformers accept `set_output(transform="pandas")` (or `"polars"`, or `"default"` for NumPy), which changes what both `transform` and `fit_transform` return — a DataFrame with column names taken from `get_feature_names_out()` rather than a bare array. It is configured per estimator instance and applies equally to both entry points, so switching between the one-step and two-step spellings never changes the container type you get back. ## Writing a custom transformer Do not implement `fit_transform` yourself unless you have one of the two reasons above. Inherit `TransformerMixin`, implement `fit` (returning `self`) and `transform` (guarded by `check_is_fitted`), and let the mixin compose them. If you *do* override it, keep the contract: the override must leave the estimator fitted exactly as `fit` would have, because callers legitimately use the estimator afterwards. A subtle trap when overriding: if your `fit_transform` and your `fit().transform()` path can diverge numerically, anything that fits a pipeline (which uses `fit_transform` on intermediate steps) will behave differently from code that fits the steps by hand. Keep them equivalent, or make the difference an explicit, documented property of the method. ## What an interviewer is listening for That you know the default is composition rather than a separate algorithm; that you can name a case where `transform` cannot exist; and that you never, under any circumstance, call `fit_transform` on data you intend to evaluate on.

  • Why does sklearn.manifold.TSNE expose fit_transform() but no transform()?
    Its embedding is optimised jointly over the samples it was given — coordinates are chosen so that pairwise neighbourhood structure is preserved among those points. Nothing is learned that could be applied to a new point, so there is no out-of-sample mapping to expose. Practically, that means a t-SNE step cannot sit inside a fitted pipeline that must transform unseen data at predict time.
  • What does TransformerMixin's default fit_transform do with the y argument?
    It forwards it: with `y` given it calls `self.fit(X, y, **fit_params).transform(X)`, and with `y=None` it calls `self.fit(X, **fit_params).transform(X)`. That supports supervised transformers — a feature selector that scores columns against the target needs `y` while fitting, even though `transform` afterwards takes only `X`.
  • What actually goes wrong if you call fit_transform on the test set?
    The transformer is re-fitted on the test data, so those rows are scaled or encoded against their own statistics instead of the training-time ones. Nothing raises and the shapes are right — the reported score is simply not a measure of generalisation, and typically flatters the model. New data must always go through `transform`.

saying these in an interview costs you the question

  • Thinks fit_transform is a distinct algorithm rather than composition
  • Calls fit_transform on the test set as well
  • Assumes every transformer also exposes a standalone transform
  • Thinks fit_transform returns self like fit does
  • Believes fit_transform is defined on BaseEstimator

context