What does it mean for a probability scoring rule like the Brier score to be proper?
answer
- about incentives, not about accuracy
- what report maximises your expected reward?
- differentiate the expected score in p
- the optimum sits exactly at your belief
- absolute error would pay you to say 1.0
basics
~20 sA scoring rule is proper when a forecaster gets their best expected score by reporting the probability they genuinely believe, rather than by shading it. Log-loss and the Brier score are strictly proper; thresholded accuracy is not.
solid answer
~50 sA scoring rule turns a predicted probability plus the realised outcome into a number. It is **proper** if, whatever a forecaster privately believes the probability `q` to be, their expected score is optimised by reporting exactly `q`; **strictly proper** if `q` is the unique optimum. Check it for Brier: expected penalty under belief `q` is `q*(1-p)^2 + (1-q)*p^2`, whose derivative in `p` is `2p - 2q`, zero only at `p = q`. Log-loss gives the same answer: `-q*log(p) - (1-q)*log(1-p)` is minimised at `p = q`. Improper rules pay you to lie. Mean absolute error on probabilities is linear in `p`, so it is minimised at `p = 1` whenever `q > 0.5` and rewards fake certainty. Accuracy at a fixed threshold is flat in `p` on each side of the cut, so it is indifferent between an honest 0.55 and a bluffed 0.99.
go deeper
Recall the one-sentence definition: you get your best expected score by reporting the probability you actually believe, and both log-loss and Brier have that property.
Be ready to prove it. Write the expected Brier penalty under a belief q, differentiate to get 2p - 2q, and show the optimum sits at p = q. Then name an improper rule and say why it fails.
Demonstrate why it matters in practice: a headline metric that is improper quietly pushes teams to extremise their probabilities, so you check properness before adopting any new scoring metric.
Own the incentive design. Choosing what the org scores decides what people optimise, so a probability metric must be strictly proper, and any business-facing summary built on top of it must not reintroduce a reward for false certainty.
## Scoring rules, and what "proper" adds A **scoring rule** is any function that takes a forecast probability `p` and the realised outcome `y` and returns a number. Log-loss uses `-log(p)` when `y = 1` and `-log(1 - p)` when `y = 0`. Brier uses `(p - y)^2`. Both are negatively oriented, so a forecaster wants the number small. Being a scoring rule is a low bar; almost any formula qualifies. **Properness** is the property that makes a scoring rule usable as an incentive. Suppose the forecaster's honest belief is that the event happens with probability `q`. If they report `p`, their **expected** score, taken over the outcome drawn with probability `q`, is `E[S] = q * S(p, y=1) + (1 - q) * S(p, y=0)` The rule is **proper** if this expected score is optimised at `p = q` for every possible `q`, and **strictly proper** if `p = q` is the only optimum. In words: the best thing you can do, in expectation, is tell the truth. There is no shading, sharpening or hedging strategy that beats honesty. ## Verifying it for the two scores For the Brier score: `E[S] = q*(1 - p)^2 + (1 - q)*p^2` Differentiate with respect to `p`: `-2q(1 - p) + 2(1 - q)p = 2p - 2q`. Set to zero and `p = q`; the second derivative is 2, so it is a strict minimum. Brier is strictly proper. For log-loss: `E[S] = -q*log(p) - (1 - q)*log(1 - p)` Differentiate: `-q/p + (1 - q)/(1 - p)`. Setting it to zero gives `q(1 - p) = (1 - q)p`, hence `p = q`. Strictly proper again. (At the optimum, the value of that expectation is the constant a perfectly honest forecaster cannot beat, which is why an achievable log-loss floor exists at all.) ## What an improper rule looks like Two instructive failures: **Mean absolute error on probabilities.** Score a row with `|p - y|`. Expected score under belief `q` is `q(1 - p) + (1 - q)p = q + p(1 - 2q)`, which is **linear** in `p`. A linear function on `[0, 1]` is minimised at an endpoint: `p = 1` when `q > 0.5`, `p = 0` when `q < 0.5`. So this rule pays a forecaster who believes 0.6 to report 1.0. It rewards manufactured certainty, and any model tuned against it will drift to the extremes. **Accuracy at a fixed threshold.** Convert `p` to a label at 0.5 and count hits. Expected accuracy depends on `p` only through which side of 0.5 it falls, so it is completely flat within each side. An honest 0.55 and a bluffed 0.99 score identically, and if the threshold is somewhere else the rule can actively prefer a distorted probability. Accuracy is a fine summary of decisions, but it carries no information about whether the probabilities themselves are honest. ## Why practitioners care 1. **Model selection cannot be gamed by sharpening.** Under a proper rule, a model does not improve its expected score by pushing predictions toward 0 and 1 for show. If sharpening does improve the score, that is evidence the model was genuinely underconfident, not evidence of gaming. 2. **Leaderboards and forecasting tournaments are safe.** Public competitions score probabilities with strictly proper rules precisely so that entrants have no incentive to submit anything other than their best estimate. 3. **It separates a metric problem from a model problem.** If your headline metric is improper and the model keeps drifting to extremes, the metric is the cause. Ruling that out first saves a lot of wasted modelling. ## Two things properness does *not* give you - **It does not make two proper rules agree.** Properness is a statement about one forecaster's own optimum, not about how two imperfect models are ordered. Log-loss and Brier are both strictly proper and can still rank two models differently, because they weight extreme-probability errors differently. - **It does not make the score interpretable on its own.** A proper score still has no absolute meaning; it has to be read against a reference forecaster on the same data. ## Interview register Define it in one sentence: the expected score is optimised by reporting your true belief. Then show the Brier derivative `2p - 2q` as the one-line proof, and name one improper rule and why it fails. That combination, definition plus proof plus counterexample, is what separates a candidate who has read the term from one who understands it.
- Why is mean absolute error on predicted probabilities not a proper scoring rule?Its expected penalty under a belief `q` is `q + p(1 - 2q)`, which is linear in the reported `p`. A linear function is minimised at an endpoint, so the rule pays a forecaster who believes 0.6 to report 1.0 and one who believes 0.4 to report 0. It systematically rewards certainty the forecaster does not have.
- If log-loss and Brier are both strictly proper, why can they crown different winners?Properness constrains each rule against one forecaster's own beliefs: honesty is optimal. It says nothing about how two imperfect models compare. Log-loss weights a confident miss without bound while Brier caps it at 1, so models that differ mainly in the extreme-probability rows can be ordered differently by the two rules.
- Would shrinking every prediction toward the base rate improve a proper score?Only if the model was overconfident. Under a proper rule the expected score is optimised at the true conditional probability, so uniform shrinkage helps exactly to the extent the predictions were too extreme and hurts an already-honest forecaster. If shrinkage reliably improves the score, that is a finding about the model, not a trick.
A proper scoring rule is a bet designed so that bluffing never pays: your best move is to state the odds you actually believe.
saying these in an interview costs you the question
- Says proper just means the score is a good metric
- Thinks two proper rules must rank models identically
- Claims accuracy is proper because it punishes wrong answers
- Confuses properness with the score being bounded
- Cannot name a single improper scoring rule