skip to content

What are scikit-learn estimator tags, and when do you implement __sklearn_tags__?

level: seniorimportance: nice to knowfreq 28%

answer

  1. declarative capability metadata
  2. read by the conformance checks, not by fit
  3. a dataclass since 1.6, a dict before
  4. call super() then mutate the fields
  5. describing is not enabling

basics

~20 s

Estimator tags are declarative metadata describing what an estimator can accept and do — sparse input, NaN tolerance, whether y is required, the estimator type. Since scikit-learn 1.6 they are a dataclass returned by sklearn_tags, which you override only when your estimator departs from the defaults.

solid answer

~40 s

Tags let generic code ask an estimator about its capabilities without trying an operation and catching the failure. They are consumed mainly by `check_estimator` and `parametrize_with_checks`, which use them to decide which of the common conformance checks apply, and by helpers such as `is_classifier` that read the estimator type. Since scikit-learn 1.6 the mechanism is `__sklearn_tags__`, returning a `sklearn.utils.Tags` dataclass with nested groups — `input_tags`, `target_tags`, `classifier_tags`, `regressor_tags`, `transformer_tags` — replacing the older `_more_tags`/`_get_tags` dict protocol, which has since been removed. You override it by calling `super().__sklearn_tags__()`, mutating the fields that differ, and returning the object; for example setting `tags.input_tags.allow_nan = True` on an estimator that handles missing values natively. Tags **describe**; they do not change behaviour, so setting one does not make the estimator capable of anything.

code

python · 8 lines
python
from sklearn.base import BaseEstimator, TransformerMixin

class NanTolerantCenterer(TransformerMixin, BaseEstimator):
    def __sklearn_tags__(self):
        tags = super().__sklearn_tags__()   # start from the defaults
        tags.input_tags.allow_nan = True    # we handle missing values
        tags.input_tags.sparse = True       # and sparse matrices
        return tags

go deeper

for a junior

Enough to know tags exist as metadata describing what an estimator accepts, and that a default estimator inheriting BaseEstimator gets sensible ones without any work from you.

for a middle

Be able to say tags are returned by sklearn_tags as a dataclass, that you override by calling super() and mutating fields, and that they describe capabilities rather than granting them.

for a senior

Recognise the 1.6 migration in the wild — the super() AttributeError, silently ignored _more_tags after an upgrade — and know that the conformance checks are the main consumer.

for a principal

Treat declared tags as an API contract for in-house estimators: a false claim removes the checks that would have caught the bug, so tag declarations belong in review alongside the code that implements the capability.

## What tags are for scikit-learn's generic machinery frequently needs to know something about an estimator before calling it. Does it accept sparse matrices? Does it cope with `NaN` in `X`, or must missing values be imputed first? Does `fit` require `y`? Is it a classifier, a regressor, a clusterer or a transformer? Is it deterministic, so a check may assert that two fits produce identical output? Discovering these by trial — call and catch the exception — is slow and unreliable. Tags make the answers declarative: a small block of metadata the estimator publishes about itself. ## The shape since 1.6 Before scikit-learn 1.6, tags were a dictionary produced by a `_more_tags` method and merged by `_get_tags`. In 1.6 that was replaced by `__sklearn_tags__`, returning a `sklearn.utils.Tags` dataclass, and the old protocol was subsequently removed — code that still defines `_more_tags` is simply ignored by current versions, which is a quiet way to lose behaviour during an upgrade. The dataclass groups related flags: - **`estimator_type`** — the kind of estimator: classifier, regressor, clusterer, transformer, or none. - **`target_tags`** — properties of `y`, notably `required` (does `fit` need a target at all) and whether multi-output is supported. - **`input_tags`** — properties of `X`: `sparse`, `allow_nan`, `categorical`, `string`, `positive_only`, `pairwise`, and the accepted array dimensionalities. - **`classifier_tags`** / **`regressor_tags`** — kind-specific expectations such as multi-label or multi-class support and whether the estimator is expected to beat a trivial baseline in the checks. - **`transformer_tags`** — transformer-specific expectations. - Top-level flags including `requires_fit`, `non_deterministic` and `array_api_support`. `BaseEstimator.__sklearn_tags__` returns sensible defaults, and each mixin refines them — `ClassifierMixin` sets the estimator type to classifier and marks `y` as required, `TransformerMixin` populates the transformer group. ## Who reads them The main consumer is the conformance suite. `check_estimator` from `sklearn.utils.estimator_checks` runs a long list of common checks, and tags decide which ones apply: an estimator that declares `input_tags.sparse = False` is not subjected to the sparse-input checks, and one that declares `allow_nan = True` is not failed for accepting `NaN`. `parametrize_with_checks` from the same module turns those checks into individual pytest cases. Meta-estimators and helper functions also consult them — `sklearn.base.is_classifier` and `is_regressor` read the estimator type, which in turn drives decisions such as whether a cross-validation helper stratifies by default. ## How you set them The pattern is fixed: take the inherited tags, mutate what differs, return them. Only override for capabilities that genuinely differ from the defaults. Most custom estimators never need to touch tags at all — that is why this is a senior-tier detail rather than everyday API. ## Tags describe, they do not enable The misconception worth guarding against: setting `input_tags.allow_nan = True` does not teach your estimator to handle missing values. It asserts that it already does. If the claim is false, the common checks will stop protecting you and real data will produce garbage or an obscure error deep inside the fit. Tags are a contract you are signing, not a switch you are flipping. Similarly, letting `NaN` through your input validation — for instance by passing `ensure_all_finite=False` to `check_array` inside `fit` — is the *implementation* half. The tag is the *declaration* half. Real support needs both, and they are independent: one without the other is either an undeclared capability or a false claim. ## The upgrade trap The 1.6 migration is where most people meet tags for the first time, usually through an error rather than the documentation: `AttributeError: 'super' object has no attribute '__sklearn_tags__'` This appears when a mixin's `__sklearn_tags__` calls `super()` and the chain does not reach `BaseEstimator`. The usual causes are declaring bases in the wrong order (`BaseEstimator` first instead of last), inheriting a mixin without inheriting `BaseEstimator` at all, or subclassing a third-party estimator base that was written for the old protocol and never migrated. The fix on your own code is to put mixins first and `BaseEstimator` last; on a third-party base it is to upgrade the dependency, or to implement `__sklearn_tags__` yourself and construct the tags rather than delegating upward. ## When this actually comes up in an interview Rarely as "list the tags" — nobody memorises the dataclass. It comes up as a maintenance story: someone upgraded scikit-learn, a custom estimator started raising an `AttributeError` about `__sklearn_tags__`, and the question is whether you can explain the protocol change, name the current mechanism, and say what `_more_tags` used to do. Knowing that tags are descriptive metadata consumed by the checks, and that the dict protocol became a dataclass protocol in 1.6, is the whole of the expected answer.

  • What replaced the _more_tags mechanism, and what happens to code that still defines it?
    scikit-learn 1.6 introduced `__sklearn_tags__`, returning a `sklearn.utils.Tags` dataclass instead of a dict, and the `_more_tags`/`_get_tags` protocol was removed in a subsequent release. Current versions simply ignore a `_more_tags` method, so an estimator that relied on it silently reverts to default tags after an upgrade — nothing raises, the declared capabilities just disappear.
  • How does check_estimator use tags?
    It runs the library's common conformance checks and consults the tags to decide which apply. An estimator declaring no sparse support is not subjected to the sparse-input checks; one declaring `allow_nan` is not failed for accepting missing values; the estimator type selects the classifier-, regressor- or transformer-specific checks. Tags therefore keep the suite from failing an estimator for capabilities it never claimed.
  • An estimator that worked on scikit-learn 1.5 now raises AttributeError about __sklearn_tags__ on super(). What is wrong?
    The tags super() chain does not reach `BaseEstimator`. Either the bases are declared in the wrong order — `BaseEstimator` must come last, after the mixins — or the class inherits a mixin without `BaseEstimator`, or it subclasses a third-party base still written for the pre-1.6 dict protocol. Fix your own ordering, upgrade the dependency, or implement `__sklearn_tags__` directly rather than delegating upward.

saying these in an interview costs you the question

  • Thinks setting a tag makes the estimator capable of that input
  • Still overrides _more_tags on a current scikit-learn version
  • Sets tags as plain class attributes instead of via __sklearn_tags__
  • Overrides __sklearn_tags__ without calling super() first
  • Believes fit validates data against the declared tags

context