skip to content

What does it mean for a probability scoring rule like the Brier score to be proper?

level: middleimportance: should knowfreq 41%

answer

  1. about incentives, not about accuracy
  2. what report maximises your expected reward?
  3. differentiate the expected score in p
  4. the optimum sits exactly at your belief
  5. absolute error would pay you to say 1.0

basics

~20 s

A scoring rule is proper when a forecaster gets their best expected score by reporting the probability they genuinely believe, rather than by shading it. Log-loss and the Brier score are strictly proper; thresholded accuracy is not.

solid answer

~50 s

A scoring rule turns a predicted probability plus the realised outcome into a number. It is **proper** if, whatever a forecaster privately believes the probability `q` to be, their expected score is optimised by reporting exactly `q`; **strictly proper** if `q` is the unique optimum. Check it for Brier: expected penalty under belief `q` is `q*(1-p)^2 + (1-q)*p^2`, whose derivative in `p` is `2p - 2q`, zero only at `p = q`. Log-loss gives the same answer: `-q*log(p) - (1-q)*log(1-p)` is minimised at `p = q`. Improper rules pay you to lie. Mean absolute error on probabilities is linear in `p`, so it is minimised at `p = 1` whenever `q > 0.5` and rewards fake certainty. Accuracy at a fixed threshold is flat in `p` on each side of the cut, so it is indifferent between an honest 0.55 and a bluffed 0.99.

go deeper

for a junior

Recall the one-sentence definition: you get your best expected score by reporting the probability you actually believe, and both log-loss and Brier have that property.

for a middle

Be ready to prove it. Write the expected Brier penalty under a belief q, differentiate to get 2p - 2q, and show the optimum sits at p = q. Then name an improper rule and say why it fails.

for a senior

Demonstrate why it matters in practice: a headline metric that is improper quietly pushes teams to extremise their probabilities, so you check properness before adopting any new scoring metric.

for a principal

Own the incentive design. Choosing what the org scores decides what people optimise, so a probability metric must be strictly proper, and any business-facing summary built on top of it must not reintroduce a reward for false certainty.

## Scoring rules, and what "proper" adds A **scoring rule** is any function that takes a forecast probability `p` and the realised outcome `y` and returns a number. Log-loss uses `-log(p)` when `y = 1` and `-log(1 - p)` when `y = 0`. Brier uses `(p - y)^2`. Both are negatively oriented, so a forecaster wants the number small. Being a scoring rule is a low bar; almost any formula qualifies. **Properness** is the property that makes a scoring rule usable as an incentive. Suppose the forecaster's honest belief is that the event happens with probability `q`. If they report `p`, their **expected** score, taken over the outcome drawn with probability `q`, is `E[S] = q * S(p, y=1) + (1 - q) * S(p, y=0)` The rule is **proper** if this expected score is optimised at `p = q` for every possible `q`, and **strictly proper** if `p = q` is the only optimum. In words: the best thing you can do, in expectation, is tell the truth. There is no shading, sharpening or hedging strategy that beats honesty. ## Verifying it for the two scores For the Brier score: `E[S] = q*(1 - p)^2 + (1 - q)*p^2` Differentiate with respect to `p`: `-2q(1 - p) + 2(1 - q)p = 2p - 2q`. Set to zero and `p = q`; the second derivative is 2, so it is a strict minimum. Brier is strictly proper. For log-loss: `E[S] = -q*log(p) - (1 - q)*log(1 - p)` Differentiate: `-q/p + (1 - q)/(1 - p)`. Setting it to zero gives `q(1 - p) = (1 - q)p`, hence `p = q`. Strictly proper again. (At the optimum, the value of that expectation is the constant a perfectly honest forecaster cannot beat, which is why an achievable log-loss floor exists at all.) ## What an improper rule looks like Two instructive failures: **Mean absolute error on probabilities.** Score a row with `|p - y|`. Expected score under belief `q` is `q(1 - p) + (1 - q)p = q + p(1 - 2q)`, which is **linear** in `p`. A linear function on `[0, 1]` is minimised at an endpoint: `p = 1` when `q > 0.5`, `p = 0` when `q < 0.5`. So this rule pays a forecaster who believes 0.6 to report 1.0. It rewards manufactured certainty, and any model tuned against it will drift to the extremes. **Accuracy at a fixed threshold.** Convert `p` to a label at 0.5 and count hits. Expected accuracy depends on `p` only through which side of 0.5 it falls, so it is completely flat within each side. An honest 0.55 and a bluffed 0.99 score identically, and if the threshold is somewhere else the rule can actively prefer a distorted probability. Accuracy is a fine summary of decisions, but it carries no information about whether the probabilities themselves are honest. ## Why practitioners care 1. **Model selection cannot be gamed by sharpening.** Under a proper rule, a model does not improve its expected score by pushing predictions toward 0 and 1 for show. If sharpening does improve the score, that is evidence the model was genuinely underconfident, not evidence of gaming. 2. **Leaderboards and forecasting tournaments are safe.** Public competitions score probabilities with strictly proper rules precisely so that entrants have no incentive to submit anything other than their best estimate. 3. **It separates a metric problem from a model problem.** If your headline metric is improper and the model keeps drifting to extremes, the metric is the cause. Ruling that out first saves a lot of wasted modelling. ## Two things properness does *not* give you - **It does not make two proper rules agree.** Properness is a statement about one forecaster's own optimum, not about how two imperfect models are ordered. Log-loss and Brier are both strictly proper and can still rank two models differently, because they weight extreme-probability errors differently. - **It does not make the score interpretable on its own.** A proper score still has no absolute meaning; it has to be read against a reference forecaster on the same data. ## Interview register Define it in one sentence: the expected score is optimised by reporting your true belief. Then show the Brier derivative `2p - 2q` as the one-line proof, and name one improper rule and why it fails. That combination, definition plus proof plus counterexample, is what separates a candidate who has read the term from one who understands it.

  • Why is mean absolute error on predicted probabilities not a proper scoring rule?
    Its expected penalty under a belief `q` is `q + p(1 - 2q)`, which is linear in the reported `p`. A linear function is minimised at an endpoint, so the rule pays a forecaster who believes 0.6 to report 1.0 and one who believes 0.4 to report 0. It systematically rewards certainty the forecaster does not have.
  • If log-loss and Brier are both strictly proper, why can they crown different winners?
    Properness constrains each rule against one forecaster's own beliefs: honesty is optimal. It says nothing about how two imperfect models compare. Log-loss weights a confident miss without bound while Brier caps it at 1, so models that differ mainly in the extreme-probability rows can be ordered differently by the two rules.
  • Would shrinking every prediction toward the base rate improve a proper score?
    Only if the model was overconfident. Under a proper rule the expected score is optimised at the true conditional probability, so uniform shrinkage helps exactly to the extent the predictions were too extreme and hurts an already-honest forecaster. If shrinkage reliably improves the score, that is a finding about the model, not a trick.

A proper scoring rule is a bet designed so that bluffing never pays: your best move is to state the odds you actually believe.

saying these in an interview costs you the question

  • Says proper just means the score is a good metric
  • Thinks two proper rules must rank models identically
  • Claims accuracy is proper because it punishes wrong answers
  • Confuses properness with the score being bounded
  • Cannot name a single improper scoring rule

context