skip to content

What does sklearn.base.clone() copy from an estimator, and what does it drop?

level: middleimportance: should knowfreq 45%

answer

  1. configuration, not learned state
  2. built from get_params, not from copy
  3. fitted attributes never survive
  4. nested estimators come back unfitted
  5. the round-trip identity check

basics

~20 s

clone() returns a new, unfitted estimator of the same class built from get_params(deep=False): hyperparameters only. All fitted attributes are dropped. Plain parameter values are deep-copied, and any parameter that is itself an estimator is cloned recursively, so it comes back unfitted too.

solid answer

~40 s

`sklearn.base.clone(estimator)` is "same configuration, no learned state". It reads `get_params(deep=False)`, recursively clones each value — an estimator-valued parameter is cloned, anything else is deep-copied — constructs a fresh object with `Klass(**params)`, and then verifies the new object hands back the identical parameter objects, raising `RuntimeError` if the constructor modified them. Fitted attributes such as `coef_` are never carried over. This is the primitive underneath everything that fits a model more than once: every cross-validation fold, every candidate in a hyperparameter search, and every step of a pipeline is fitted on a clone, which is why the estimator object you passed in stays unfitted afterwards. Estimators that need to preserve something across a clone can override `__sklearn_clone__`.

go deeper

for a junior

Know that clone gives you a fresh, unfitted estimator with the same settings, and that this is why the model you pass into cross-validation is still unfitted afterwards.

for a middle

Describe the mechanism: get_params(deep=False), recursive cloning of values, reconstruction through the constructor, and the identity check that raises RuntimeError on a modifying init.

for a senior

Use it diagnostically — reach for clone when a custom estimator misbehaves inside a search, and explain how per-fold cloning is what prevents state leaking between folds.

for a principal

Discuss the guarantee it encodes: fit-time isolation between repeated fits is a correctness property of the whole evaluation stack, and sklearn_clone overrides move that guarantee onto the estimator author.

## What clone is for Generic machine-learning machinery needs to fit the same *configuration* many times over different data: five folds of a cross-validation, two hundred candidates in a hyperparameter search, one hundred base learners in a bagging ensemble. It must never reuse a fitted object, because leftover state would leak between fits. It also cannot use the class name and a dict of arguments, because it was handed a configured *instance*. `sklearn.base.clone` bridges that gap. Given an instance, it produces a new instance of the same class with the same hyperparameters and no learned state — an unfitted template. ## What it copies, precisely The procedure is short enough to hold in your head: 1. Call `estimator.get_params(deep=False)` to obtain this estimator's own hyperparameters (not the nested `component__param` expansion). 2. Recurse into every value with `clone(value, safe=False)`. The `safe=False` flag means "if this value is not an estimator, do not complain" — such values are returned as a `copy.deepcopy`. 3. Construct `Klass(**cloned_params)`. 4. Call `get_params(deep=False)` on the *new* object and check, for each name, that the value it reports **is** the object that was passed in. Identity, not equality. Step 4 is the round-trip guard: a constructor that converts, copies or defaults an argument fails it and clone raises `RuntimeError` saying the constructor either does not set or modifies that parameter. ## What it drops Everything with a trailing underscore. `coef_`, `intercept_`, `classes_`, `n_features_in_`, `tree_` — none of it survives, because none of it appears in `get_params`. A clone of a fitted `LogisticRegression` is an unfitted `LogisticRegression`; calling `predict` on it raises `NotFittedError`. Any other attribute you happened to attach to the instance outside the constructor is gone as well: only declared hyperparameters make the crossing. ## Nested estimators are cloned, not shared Step 2 has a consequence people miss. If a hyperparameter's value is itself an estimator — the base learner of a bagging ensemble, the estimator inside a recursive feature eliminator, a step of a pipeline — that value is *cloned* rather than copied. It therefore comes back unfitted, even if you handed in a fitted one, and the outer clone does not share it with the original. A non-estimator value is deep-copied instead, so a NumPy array or list passed as a hyperparameter is not shared between the original and the clone either. Mutating one will not affect the other. ## Why your estimator stays unfitted This explains a behaviour that surprises newcomers. After `cross_val_score(model, X, y)` or a hyperparameter search's `fit`, the `model` object you passed in is still unfitted — every fold fitted a clone. Search objects expose the winning configuration as `best_estimator_` (a refit clone, when refitting is enabled), which is the object you actually deploy. Similarly, fitting a pipeline does not fit the transformer instances you constructed by hand; it fits copies held inside the pipeline. ## The failure mode to recognise `RuntimeError: Cannot clone object ..., as the constructor either does not set or modifies parameter ...` means precisely one thing: `__init__` did something other than `self.name = name`. It is raised at clone time, which is usually deep inside a cross-validation or a search, so the traceback points at library code and the cause is in your constructor. The fix is always to move the conversion or validation into `fit`. A related error is an `AttributeError` from `get_params` when the constructor stored an argument under a different name than the signature declares. ## Customising: __sklearn_clone__ Since scikit-learn 1.3 an estimator can override `__sklearn_clone__` to control what cloning produces. `clone` checks for it and delegates when present. It is a deliberately narrow escape hatch — a frozen or pre-trained estimator that should survive being cloned inside a pipeline is the motivating case — and overriding it means taking responsibility for the guarantee everything else depends on: that a clone carries no state from a previous fit. ## clone(..., safe=False) The `safe` keyword controls what happens when the argument is not an estimator at all. With the default `safe=True`, a non-estimator raises `TypeError`. With `safe=False`, it is deep-copied and returned. That is what the recursive step over parameter values uses, and it is occasionally useful when you hold a heterogeneous list of "things that may or may not be estimators". ## The one-line summary for an interview Clone is a *configuration* copy, never a *state* copy, and the identity check it performs on the round-trip is what forces estimator constructors to be inert.

  • Why is the estimator you pass to cross_val_score still unfitted when it returns?
    Because each fold fits a clone, not your object. The routine calls `clone` per split precisely so that no state leaks between folds and your input is left untouched. If you want a fitted model afterwards you fit it yourself on the full data, or take `best_estimator_` from a search that refits on the winning configuration.
  • Does clone deep-copy a NumPy array passed as a hyperparameter?
    Yes. Each parameter value goes through `clone(value, safe=False)`, which returns a `copy.deepcopy` for anything that is not an estimator. So arrays, lists and dicts are not shared between the original and the clone, and mutating one leaves the other unaffected. Estimator-valued parameters take the other branch and are cloned recursively, arriving unfitted.
  • How would you make an already-trained sub-estimator survive being cloned inside a larger model?
    Override `__sklearn_clone__` on that estimator — since scikit-learn 1.3, `clone` delegates to it when present, so the estimator decides what a clone of itself means. It is the sanctioned way to express a frozen or pre-trained component, and it comes with the obligation to be sure the retained state is genuinely independent of the data the outer model is being fitted on.

saying these in an interview costs you the question

  • Thinks clone copies the fitted coefficients too
  • Describes clone as just copy.deepcopy of the estimator
  • Expects the estimator handed to a search to come back fitted
  • Believes clone re-runs fit on the original data
  • Thinks a fitted sub-estimator passed as a parameter stays fitted

context