skip to content

What does scikit-learn's StandardScaler learn in fit(), and why never fit it on test data?

level: middleimportance: must knowfreq 78%

answer

  1. fit learns, transform applies
  2. one statistic per column, stored
  3. trailing underscore attributes
  4. mean_ and scale_ frozen after fit
  5. test data never gets fit_transform

basics

~20 s

StandardScaler.fit() computes and stores each column's mean (mean_) and standard deviation (scale_). Test data must only go through transform(), because refitting replaces those statistics with test-set numbers the model would never have at serving time.

solid answer

~40 s

`StandardScaler` is a stateful transformer. `fit(X_train)` computes per-column statistics and stores them as `mean_`, `var_`, `scale_` and `n_samples_seen_`; `transform(X)` then applies `(x - mean_) / scale_` using those stored numbers. On held-out or production data you call `transform()` only — calling `fit_transform(X_test)` overwrites the learned statistics with the test set's own mean and standard deviation, which invalidates the evaluation and cannot be reproduced at serving time, where rows arrive one at a time. Two practical details: a constant column has zero variance, and scikit-learn substitutes `scale_ = 1.0` for it rather than dividing by zero, so it transforms to all zeros; and on a sparse matrix you must pass `StandardScaler(with_mean=False)`, because subtracting a mean would densify the matrix and otherwise raises.

code

python · 10 lines
python
import numpy as np
from sklearn.preprocessing import StandardScaler

X_train = np.array([[1.0, 5.0], [2.0, 5.0], [3.0, 5.0], [4.0, 5.0]])
X_test = np.array([[10.0, 5.0]])

scaler = StandardScaler().fit(X_train)
print(scaler.mean_)   # [2.5 5. ]
print(scaler.scale_)  # [1.11803399 1.  ]  <- zero-variance column gets 1.0
print(scaler.transform(X_test))

go deeper

for a junior

Know that the scaler is fitted on the training data and only applied to everything else, and be able to say which method is which: fit learns, transform applies, fit_transform does both on the same data.

for a middle

Be ready to name mean_ and scale_, write the (x - mean_) / scale_ formula, and explain why a single production row cannot be standardised without frozen statistics.

for a senior

Show you have debugged this: constant columns silently becoming zeros, sparse input needing with_mean=False, and scaler state that must be persisted alongside the model artifact or predictions drift.

for a principal

Own the position that fitted preprocessing state is part of the model artifact and must be versioned and served with it. Argue for a serving path that shares one code path with training rather than reimplementing the arithmetic.

## The transformer contract Every scikit-learn preprocessing object is stateful. `fit()` learns something from the data it is handed and stores it on the instance under a trailing-underscore attribute; `transform()` applies that stored state to any matrix with the same columns. `fit_transform()` is nothing but the two calls in sequence on the same data. The whole train/test discipline in preprocessing falls out of that split: whoever calls `fit` decides what statistics the model is standardised against forever after. ## What StandardScaler stores After `scaler.fit(X_train)` the instance carries: - `mean_` — one mean per column (unless `with_mean=False`, when it is `None`). - `var_` — one variance per column, and `scale_` — the square root of it, the actual divisor. - `n_samples_seen_` — how many rows contributed, per column, which is what makes `partial_fit()` able to update the statistics incrementally over minibatches. - `n_features_in_` and, for a DataFrame input, `feature_names_in_` — used to check that later `transform()` calls present the same columns in the same order. `transform()` computes `(x - mean_) / scale_` elementwise. `inverse_transform()` reverses it, which is how you get predictions back into original units when you have scaled a regression target with a separate scaler. ## Why fitting on test data is wrong Two separate harms, and candidates usually only name the first. The measurement harm: the test set exists to estimate performance on data the model has never influenced. If the scaler is fitted on the test set, information about the test distribution — its centre and spread — has been folded into the features the model consumes, so the score is optimistic and no longer an estimate of anything you will observe in production. The deployment harm, which is the more concrete one: at serving time you often score a single row. A single row has a mean equal to itself and a standard deviation of zero. There is no test-set-shaped batch to fit on, so a pipeline that only works when it can re-fit is a pipeline that cannot be deployed. The scaler must ship with frozen `mean_`/`scale_` learned once on training data, exactly as a model ships with frozen coefficients. The correct sequence is `scaler.fit(X_train)` then `scaler.transform(X_train)` and `scaler.transform(X_test)`; in practice you put the scaler in a `Pipeline` so `fit`/`predict` route the calls for you and there is no opportunity to get it wrong by hand. ## Edge cases interviewers probe **Constant columns.** Variance is zero, so the naive divisor is zero. scikit-learn replaces zero scales with `1.0`, so a constant feature standardises to all zeros instead of `NaN` or `inf`. No warning is raised, so a column that is constant only in your training split silently becomes a dead feature. **Sparse input.** Centring a sparse matrix would fill in every implicit zero and blow up memory, so `StandardScaler` raises on sparse input unless you pass `with_mean=False`, which divides by `scale_` without subtracting. When you want the structure of sparse data untouched, `MaxAbsScaler` is the natural choice — it divides by the maximum absolute value per column and leaves zeros as zeros. **Unseen ranges.** Standardising does not clip. A test value far outside the training range simply produces a large z-score, which is usually what you want. The sibling scalers behave differently in this respect: `MinMaxScaler` maps the training range onto `feature_range` (default `(0, 1)`) using `data_min_`/`data_max_`, so test values outside the training range land outside that interval unless you pass `clip=True`; `RobustScaler` centres on the median and divides by the interquartile range (`center_`, `scale_`, controlled by `quantile_range`, default the 25th–75th percentiles), which is why it is the usual pick when a few extreme values would otherwise dominate the mean and standard deviation. **`with_std=False`.** Centres without scaling; occasionally useful when a downstream method needs zero-mean input but the units are meaningful. ## What a strong answer sounds like Name the attributes, state the `(x - mean_) / scale_` formula, say plainly that the test set only gets `transform()`, and give the single-row serving argument rather than only the abstract 'leakage' word. If you add the zero-variance and sparse behaviours, you have demonstrated you have actually read what the object does rather than only used it.

  • What happens if one of your training columns is constant?
    Its variance is zero, so scikit-learn substitutes `scale_ = 1.0` for that column instead of dividing by zero. The column standardises to all zeros — no error, no warning. It becomes a dead feature, which matters if the column is constant only in this particular split.
  • Why does StandardScaler raise on a sparse matrix, and what do you pass to make it work?
    Subtracting a per-column mean would replace every implicit zero with a nonzero value and densify the matrix, so centring is refused. Pass `StandardScaler(with_mean=False)` to divide by `scale_` only, or use `MaxAbsScaler`, which divides by the per-column maximum absolute value and preserves the zeros.
  • How would you scale streaming data that does not fit in memory?
    Use `partial_fit()` in a loop over minibatches. `StandardScaler` tracks `n_samples_seen_` per column and updates `mean_` and `var_` incrementally, so the final statistics match a single `fit()` over the concatenated data. `MinMaxScaler` and `MaxAbsScaler` expose `partial_fit()` too.

The scaler is a ruler calibrated once on the training data. You measure everything else with that ruler; re-calibrating it against the test set means the two measurements are no longer comparable.

saying these in an interview costs you the question

  • Calling fit_transform on the test set as well
  • Thinking transform recomputes statistics each call
  • Believing scaling clips values into a fixed range
  • Assuming a zero-variance column raises an error
  • Saying scaling is unnecessary because the model handles it

context