In scikit-learn, how does make_scorer turn a metric into a scorer?
answer
- metric takes arrays, scorer takes a model
- one direction only: bigger wins
- errors arrive with a minus sign
- which response the estimator is asked for
- extra kwargs ride along to the metric
basics
~10 smake_scorer wraps a metric function into a callable of the form scorer(estimator, X, y). greater_is_better=False negates the value so larger always means better, and response_method decides whether the scorer calls predict, predict_proba, or decision_function.
solid answer
~40 sA scikit-learn *metric* has signature `metric(y_true, y_pred)`; a *scorer* has signature `scorer(estimator, X, y)` and is what search and cross-validation utilities actually call. `make_scorer(score_func, *, response_method='predict', greater_is_better=True, **kwargs)` bridges them: it calls the chosen response method on the estimator, feeds the result to your metric, and returns the number. Any extra keyword arguments are forwarded to the metric, so `make_scorer(fbeta_score, beta=2)` fixes beta once. Two conventions matter. First, scorers are always **maximised**, so error metrics must be negated — `greater_is_better=False` does that, and it is exactly why the built-in strings are `'neg_mean_squared_error'`, `'neg_root_mean_squared_error'` and `'neg_log_loss'`. Values in `cv_results_` are therefore negative, and you flip the sign before reporting. Second, `response_method` takes `'predict'`, `'predict_proba'`, `'decision_function'`, or a tuple tried in order; the older `needs_proba` and `needs_threshold` flags were removed in 1.6. `get_scorer_names()` lists the built-ins.
code
python · 13 linesimport numpy as np
from sklearn.metrics import fbeta_score, log_loss, make_scorer, get_scorer
f2_scorer = make_scorer(fbeta_score, beta=2, average="macro")
def mean_abs_cost(y_true, y_pred):
return np.abs(y_true - y_pred).mean()
cost_scorer = make_scorer(mean_abs_cost, greater_is_better=False)
ll_scorer = make_scorer(log_loss, greater_is_better=False,
response_method="predict_proba")
print(get_scorer("neg_root_mean_squared_error"))go deeper
Know that model-selection tools want a scorer, not a raw metric, and that make_scorer is the adapter. Recognise the neg_ prefix on built-in names as a sign convention rather than an error.
Explain the scorer signature, why greater_is_better=False negates the value, how extra kwargs are forwarded to the metric, and which response_method a probability metric needs.
Show that a mis-signed custom scorer silently selects the worst model, and be ready to write a scorer directly against the (estimator, X, y) protocol when the objective involves more than predictions.
Own what the organisation optimises for: whether a single business-aligned scorer is standardised, how guardrail metrics are tracked alongside it, and how that choice is kept stable so model comparisons across teams remain meaningful.
## Two different callable shapes Scikit-learn distinguishes metrics from scorers, and the distinction is the whole reason `make_scorer` exists. A **metric** lives in `sklearn.metrics` and takes ground truth and predictions: `accuracy_score(y_true, y_pred)`, `roc_auc_score(y_true, y_score)`, `mean_absolute_error(y_true, y_pred)`. It knows nothing about models. A **scorer** takes a fitted estimator and data: `scorer(estimator, X, y)`. That is the shape every model-evaluation utility calls, because those utilities hold the estimator and the data, and must decide themselves whether to ask it for labels, probabilities or margins. `make_scorer(score_func, *, response_method='predict', greater_is_better=True, **kwargs)` converts the first shape into the second. ## The greater-is-better convention Everything in scikit-learn that consumes a scorer assumes **higher is better**, so it can pick the maximum without knowing which metric it holds. Error metrics break that assumption, and rather than carry a per-metric direction flag, scikit-learn negates them. Hence `greater_is_better=False`, which wraps your metric so it returns `-value`. This is the origin of the `neg_` prefix on built-in scorer strings: `'neg_mean_squared_error'`, `'neg_mean_absolute_error'`, `'neg_root_mean_squared_error'`, `'neg_log_loss'`, `'neg_brier_score'`. It surprises people the first time they see a mean squared error of `-14.2` in cross-validation output. Nothing is wrong; multiply by -1 to report it. The corollary is that any custom loss you wrap must be negated too, or the search will confidently select the *worst* model — a failure with no error message, only a bad model. A related trap: `davies_bouldin_score` is a clustering metric where lower is better, so wrapping it also needs `greater_is_better=False`. ## Choosing the response method Some metrics take predicted labels, some take continuous scores. `response_method` says which to fetch: - `'predict'` (the default) for label metrics such as `accuracy_score`, `f1_score`, `matthews_corrcoef`. - `'predict_proba'` for probability metrics such as `log_loss` and `brier_score_loss`. - `'decision_function'` for margin-based ranking metrics on estimators that expose it. - A list or tuple such as `('decision_function', 'predict_proba')`, tried in order, which is how ranking scorers work across estimators that expose only one of the two. This parameter arrived in 1.4 and replaced the older boolean flags `needs_proba` and `needs_threshold`, which were deprecated then and **removed in 1.6**. Code written against those flags raises a `TypeError` on modern scikit-learn, and that version seam is a fair interview question in itself. Say which version you are targeting when you answer. ## Forwarding metric parameters Extra keyword arguments to `make_scorer` are passed straight through to the metric on every call. That is how you pin `beta` on `fbeta_score`, `average='macro'` on `f1_score`, `pos_label` on `precision_score`, or `zero_division` on any of them. It keeps the scorer a zero-argument-at-call-time object while still configuring the metric. ## Built-in scorers Most of the time you do not need `make_scorer` at all: passing a string like `'f1_macro'`, `'roc_auc_ovr'`, `'average_precision'` or `'r2'` selects a pre-built scorer that already knows the correct response method. `get_scorer_names()` returns the full list and `get_scorer(name)` returns the callable, which is useful for inspecting one. Preferring the strings also eliminates the whole class of bugs where a hand-rolled evaluation calls `predict` and feeds hard labels into a ranking metric. ## Beyond make_scorer When the score you want cannot be expressed as `f(y_true, y_pred)` — say it needs the fitted model's sparsity, its inference latency, or a business cost table keyed on extra columns — you write the scorer callable directly as a function of `(estimator, X, y)` and pass that instead. `make_scorer` is a convenience for the common case, not the only way in. Multiple scorers can also be evaluated at once by passing a dict of name-to-scorer, which is how you track a headline metric and a guardrail metric in the same run.
- You wrap a custom cost function with make_scorer and forget greater_is_better=False. What goes wrong?The search maximises your cost, so it selects the worst candidate in the grid and reports it as best. Nothing raises and `cv_results_` looks ordinary — the only tell is that the winning hyperparameters are implausible and the reported score is high for a quantity you wanted low. Negation is not cosmetic; it encodes the optimisation direction.
- Why does cross-validation report mean squared error as a negative number?Because scorers are maximised by convention, so scikit-learn ships error metrics pre-negated under the `neg_` prefix. `'neg_mean_squared_error'` returns -MSE, which makes the largest value the best model without the framework needing to know the metric's direction. Flip the sign when you report the figure to anyone.
- How would you score something that make_scorer cannot express, such as a penalty on model size?Write the scorer directly as a callable of `(estimator, X, y)` and pass it wherever a scorer is accepted. Inside, you can call any response method, inspect fitted attributes such as coefficient sparsity, or combine a metric with a latency measurement. make_scorer only covers the `f(y_true, y_pred)` case; the scorer protocol itself is the general interface.
- What replaced needs_proba and needs_threshold, and when?`response_method`, introduced in scikit-learn 1.4. It accepts `'predict'`, `'predict_proba'`, `'decision_function'`, or a tuple tried in order, which is strictly more expressive than two booleans and makes the fallback behaviour explicit. The old flags were deprecated in 1.4 and removed in 1.6, so on 1.9 passing them raises a TypeError.
saying these in an interview costs you the question
- Wrapping a loss without greater_is_better=False
- Treating a negative neg_mean_squared_error value as a bug
- Still passing needs_proba to make_scorer on current scikit-learn
- Assuming a scorer has the same signature as a metric
- Forgetting response_method so a probability metric receives hard labels