In scikit-learn, what does a trailing underscore in coef_ or mean_ mean?
answer
- two kinds of attribute, one naming rule
- the suffix is not privacy
- set only inside fit
- how the library detects fitted state
- check_is_fitted scans for it
basics
~20 sA trailing underscore marks state learned during fit, such as coef_, mean_ or n_features_in_. It does not exist on a freshly constructed estimator, and check_is_fitted uses the presence of such attributes to decide whether the estimator has been fitted.
solid answer
~40 sscikit-learn splits an estimator's attributes into two disjoint groups. Names without a trailing underscore are the hyperparameters you passed to `__init__` — they exist from construction and are what `get_params` reports. Names ending in an underscore are *fitted* state, created only inside `fit`: `coef_`, `mean_`, `classes_`, `n_features_in_`. The convention is load-bearing, not cosmetic: `check_is_fitted(self)` decides an estimator is fitted by looking for any attribute that ends in an underscore and is not a dunder, and raises `NotFittedError` otherwise, which is how calling `predict` before `fit` produces a clear error. `fit` must also return `self` so chaining works, and by default it *resets* — a second `fit` discards the previous fitted attributes rather than continuing training.
go deeper
Recall the split: no underscore means a constructor hyperparameter, a trailing underscore means something learned during fit. Say that reading coef_ before fitting fails, and that fit returns self.
Explain the mechanism behind it — check_is_fitted scans for trailing-underscore attributes and raises NotFittedError, so the convention is what makes the guard work in a custom estimator.
Talk about the operational edges: fit resetting rather than accumulating, warm_start and partial_fit as the deliberate exceptions, and n_features_in_/feature_names_in_ catching schema drift between training and serving.
Own the consistency argument — a naming convention that generic code can rely on removes the need for per-estimator metadata, and any in-house estimator that ignores it becomes invisible to the whole ecosystem of wrappers built on it.
## Two disjoint groups of attributes Every scikit-learn estimator carries two kinds of state, and the naming tells you which is which at a glance. **Hyperparameters** are the arguments to `__init__`, stored under their own names with no underscore: `alpha`, `n_estimators`, `C`, `with_mean`. They exist the instant the object is constructed, they are what `get_params` reports and `set_params` writes, and they are the only thing that survives a `clone`. **Fitted attributes** are learned from data. They are created inside `fit` and their names end in a single trailing underscore: `coef_`, `intercept_`, `mean_`, `scale_`, `classes_`, `n_iter_`, `feature_importances_`. On a freshly constructed estimator they do not exist at all — touching one raises `AttributeError`. That separation is what lets generic library code reason about an object it has never seen. Configuration is introspectable and reproducible; learned state is disposable and rebuilt from data. ## The underscore is not "private" The common confusion is with Python's own convention, where a *leading* underscore signals "internal, do not touch". scikit-learn's marker is on the other end of the name, and it means the opposite: `coef_` is public API, the thing you are meant to read after fitting. A leading underscore in scikit-learn still means internal, so `self._cache` is private scratch, while `self.coef_` is a documented result. ## check_is_fitted and NotFittedError `sklearn.utils.validation.check_is_fitted(estimator)` is how estimators guard `predict`, `transform` and `score`. With no explicit attribute list, it scans the instance's `__dict__` for any name that ends in `_` and is not a dunder. If it finds none, it raises `sklearn.exceptions.NotFittedError` with a message telling you to call `fit` first. Two practical implications. First, this is why the convention must be followed in a custom estimator: if you store learned state as `self.mean` instead of `self.mean_`, `check_is_fitted` cannot see it and will insist the estimator is unfitted forever. Second, it explains a subtler bug — assigning *any* underscore-suffixed attribute in `__init__` makes the estimator look fitted from birth, so learned state must be created in `fit` and nowhere else. You can also pass explicit names, `check_is_fitted(self, "mean_")`, when you want a specific attribute checked. ## The fitted attributes almost every estimator sets Input validation during `fit` records two standard pieces of state for you: - `n_features_in_` — the number of columns of `X` seen during `fit`. At `predict` time the estimator compares the incoming width against it and raises a clear `ValueError` on a mismatch instead of producing nonsense. - `feature_names_in_` — an array of column names, set **only** when `X` had string feature names, i.e. when you fitted on a pandas DataFrame. Fit on a DataFrame and predict on a raw NumPy array and you get a warning about missing feature names; fit on an array and predict on a DataFrame gets one too. In a custom estimator, calling `validate_data(self, X)` from `sklearn.utils.validation` inside `fit` sets both, and `validate_data(self, X, reset=False)` inside `transform`/`predict` checks the new data against them rather than overwriting them. (Before scikit-learn 1.6 the same logic lived on the private `BaseEstimator._validate_data` method.) ## fit resets; it does not accumulate Calling `fit` a second time on the same object is a fresh fit. The previous fitted attributes are overwritten, and any state from the earlier run is discarded — estimators are expected to behave as if they had just been cloned. This is deliberate: cross-validation and hyperparameter search reuse configurations constantly, and silent accumulation would make results depend on call history. There are two explicit opt-outs. Some iterative estimators accept `warm_start=True`, which tells `fit` to continue from the existing solution instead of resetting — useful for adding trees to a forest or extra iterations to a linear model. And a separate method, `partial_fit`, exists on estimators that support genuine incremental learning over mini-batches; it deliberately does not reset, and for classifiers usually needs the full `classes` list on the first call because it cannot infer the label set from one batch. ## fit returns self `fit` must end with `return self`. It is what makes `model.fit(X, y).predict(X_test)` and `scaler.fit(X).transform(X)` work, and generic library code relies on it. Returning `None`, or returning the transformed data, breaks pipelines and searches in confusing ways. ## Writing it correctly In a custom estimator the shape is fixed: `__init__` assigns hyperparameters only; `fit` validates input, computes results into trailing-underscore attributes, and returns `self`; `predict`/`transform` start with `check_is_fitted(self)` and then use only the fitted attributes and the hyperparameters. Follow that and the object becomes usable by every generic tool in the library without further work.
- What exception does calling predict() before fit() raise, and what produces it?`sklearn.exceptions.NotFittedError`, raised by `check_is_fitted` at the top of `predict`. With no attribute list given, it scans the instance for any non-dunder attribute whose name ends in an underscore; finding none, it reports that the estimator is not fitted and tells you to call `fit` first. That is precisely why a custom estimator must store learned state with the trailing-underscore suffix.
- Does calling fit() twice on the same estimator continue training from where it stopped?No. By default `fit` resets: the fitted attributes are recomputed from scratch and the earlier solution is discarded, so a refit behaves like fitting a clone. Two opt-ins exist — `warm_start=True` on estimators that support it continues from the current solution, and `partial_fit` performs genuine incremental learning over mini-batches without resetting.
- What are n_features_in_ and feature_names_in_, and when is the second one absent?Both are set during `fit` by input validation. `n_features_in_` is the column count seen at fit time, checked again at predict time so a width mismatch raises a clear error. `feature_names_in_` holds the column names, and is set only when `X` carried string feature names — that is, when you fitted on a DataFrame. Fitting on a NumPy array leaves it unset, and mixing the two between fit and predict triggers a warning.
saying these in an interview costs you the question
- Thinks the trailing underscore means the attribute is private
- Sets learned attributes in __init__ rather than in fit
- Expects fit to continue training by default on a second call
- Thinks check_is_fitted reads a boolean flag you must set yourself
- Forgets to return self from fit