skip to content

Log-Loss and Brier Score

Log-loss punishes a confident wrong probability without mercy while the Brier score stays quadratic, and both are proper, so honest probabilities beat hedged ones. A scoring-rule staple.

on this pageshow

questions

4

How do log-loss and the Brier score differ in punishing a confident wrong probability?

level: middleimportance: must knowfreq 68%

answer

  1. one squares the error, one takes a log
  2. how bad can one single row get?
  3. capped at 1 versus no cap at all
  4. log of zero is the problem
  5. the tail rows decide log-loss

basics

~20 s

Log-loss charges minus the log of the probability the model gave to the outcome that actually happened, so a confident miss costs unboundedly much. The Brier score squares the probability error, so no single row can cost more than 1.

solid answer

~50 s

Both are averaged per-row penalties on predicted probabilities, lower is better, and 0 is perfect. Log-loss is `-(1/n) * sum[ y*log(p) + (1-y)*log(1-p) ]`; Brier is `(1/n) * sum (p - y)^2`. The difference is the shape of the penalty as the probability assigned to the truth goes to zero. Give the truth 0.01 and log-loss charges `-ln(0.01) = 4.61` for that row; Brier charges `0.99^2 = 0.98`, and it can never charge more than 1. So an election-night model that put 99% on the side that lost takes an enormous, effectively unbounded hit on log-loss and a bounded one on Brier. Practically: log-loss is dominated by a handful of overconfident rows and is undefined at an exact 0 or 1, so it needs clipping; Brier is bounded, needs no clipping, and is more stable across resamples.

code

python · 20 lines
python
import math

actual = [1, 1, 0, 0, 1, 0, 1, 0]
hedged = [0.7, 0.6, 0.4, 0.3, 0.7, 0.4, 0.6, 0.6]
bold   = [0.95, 0.9, 0.1, 0.05, 0.9, 0.1, 0.9, 0.99]

def log_loss(p, y):
    return -sum(t * math.log(q) + (1 - t) * math.log(1 - q)
                for q, t in zip(p, y)) / len(y)

def brier(p, y):
    return sum((q - t) ** 2 for q, t in zip(p, y)) / len(y)

for name, p in (("hedged", hedged), ("bold", bold)):
    print(name, round(log_loss(p, actual), 4), round(brier(p, actual), 4))

# hedged 0.5037 0.1588
# bold 0.6543 0.1294
# the last row (0.99 on an outcome that did not happen) costs
# bold 4.6052 of log-loss but only 0.9801 of Brier

go deeper

for a junior

Be ready to write both formulas from memory and say that lower is better with 0 perfect. Know that both score probabilities, not hard yes/no labels.

for a middle

An interviewer expects you to explain the penalty shapes: minus log of the probability given to the truth versus its squared error, and why only one of them can go to infinity. Have the clipping fix ready.

for a senior

Show that you have operated these numbers: log-loss on a real hold-out is often decided by a few overconfident rows, so you inspect the worst rows, fix the clipping epsilon, and expect more resample noise than Brier gives you.

for a principal

Own the consequence for how the org compares models: an unbounded metric makes leaderboards fragile and sensitive to a clipping constant nobody agreed on, so the metric definition, including the epsilon, has to be written down and frozen.

## The two formulas Both scores take a set of predicted probabilities `p_i` for a binary outcome `y_i` in {0, 1}, charge each row a penalty, and average. Both are negatively oriented: lower is better, and 0 is a perfect set of forecasts. - **Log-loss** (also written as the average negative log-likelihood of the observed labels): `-(1/n) * sum_i [ y_i*log(p_i) + (1 - y_i)*log(1 - p_i) ]`. Only one of the two terms is alive on any row: when `y = 1` the row costs `-log(p)`, when `y = 0` it costs `-log(1 - p)`. In both cases the row costs **minus the log of the probability the model gave to what actually happened**. - **Brier score**: `(1/n) * sum_i (p_i - y_i)^2`. This is simply the mean squared error of the probabilities against the 0/1 outcomes. ## The penalty curves Write `t` for the probability the model gave the true outcome. Then the row costs `-log(t)` under log-loss and `(1 - t)^2` under Brier. Reading off a few values: | probability given to the truth | log-loss row | Brier row | |---|---|---| | 0.9 | 0.105 | 0.01 | | 0.5 | 0.693 | 0.25 | | 0.1 | 2.303 | 0.81 | | 0.01 | 4.605 | 0.980 | | 0.001 | 6.908 | 0.998 | | 0 | infinite | 1 | That last line is the whole story. Brier's per-row penalty is capped at 1 no matter how wrong the model was; log-loss grows without limit and blows up at an exact 0. An election-night model that put 99% on the side that lost contributes 4.61 to the log-loss sum from that one race, more than forty times what a mildly wrong 0.9-on-a-miss row contributes, while under Brier the same race contributes 0.98 and a row can never contribute more than a row that was merely certain-and-wrong. A consequence people miss: because log-loss is unbounded, a small number of rows can decide the whole number. On a hold-out of a few thousand rows, ten badly overconfident predictions can move average log-loss more than the other rows combined, which also makes log-loss noisier across bootstrap resamples than Brier. ## The two rules can disagree Because the penalties weight extremes so differently, two models scored on the same data can rank differently under the two rules. A model that hedges everything toward the middle avoids the log-loss cliff but accumulates mediocre squared errors on every row; a model that is sharp and mostly right but occasionally catastrophically overconfident does the opposite. That is not a bug in either metric and neither one is "broken" when it happens; it is a signal that the models differ mostly in the tails. ## Zeros, ones and clipping `log(0)` is undefined, so a predicted probability of exactly 0 on a parcel that did arrive late makes log-loss infinite and the whole evaluation useless. The standard fix is to clip predictions into `[eps, 1 - eps]` before scoring, commonly with `eps = 1e-15`, which caps the worst row at `-ln(1e-15) = 34.54`. Be aware of what that means: the clip is a knob on the metric. A larger epsilon quietly forgives overconfidence, a smaller one amplifies it, and two teams using different epsilons are not reporting the same number. Fix the epsilon, write it down, and never change it mid-comparison. Brier needs no clipping at all, since `(p - y)^2` is perfectly well defined at 0 and 1. ## Details worth having ready - **Base of the logarithm.** Natural log is the convention; using log base 2 multiplies every value by `1/ln(2) = 1.4427` and changes the units, not the ranking of models. - **Multiclass.** Log-loss generalises directly: each row costs minus the log of the probability assigned to the true class, so the other classes matter only through the constraint that the probabilities sum to one. Brier generalises as the sum of squared errors across the one-hot vector. - **Scale.** Neither number is a percentage. A log-loss of 0.31 is not "31% wrong", and both scores must be read against a reference forecaster rather than against an absolute standard. ## What to say in an interview State both formulas, state that both are averaged per-row penalties where lower is better, and then make the one point that distinguishes them: log-loss punishes confident mistakes without bound and needs clipping, Brier caps each row at 1 and does not. Finish by noting that the two can therefore disagree about which of two models is better, and that the disagreement localises to the extreme-probability rows.

  • What happens to log-loss if the model outputs exactly 0 for an event that then occurs?
    That row costs `-log(0)`, so the average is infinite and the evaluation is destroyed. The fix is to clip predictions into `[eps, 1 - eps]` before scoring; with `eps = 1e-15` the worst row costs 34.54 instead. Treat the epsilon as part of the metric definition: changing it changes everyone's score. Brier has no such problem, since `(p - y)^2` is finite at 0 and 1.
  • Does reporting log-loss in bits instead of nats change which model wins?
    No. Switching from natural log to log base 2 multiplies every row, and therefore the average, by the same constant `1/ln(2) = 1.4427`. Rankings, relative improvements and skill ratios are all unchanged; only the units of the printed number change. Just make sure everyone comparing numbers is using the same base.
  • How does log-loss extend to a five-class problem?
    Each row costs minus the log of the probability the model assigned to that row's true class, and you average over rows. The probabilities for the wrong classes enter only through the constraint that all five sum to one, so pushing mass onto a wrong class is punished indirectly by the mass it takes away from the right one.

Brier fines you like a parking ticket with a maximum penalty; log-loss fines you like compound interest on a bet you swore was safe.

saying these in an interview costs you the question

  • Says log-loss is bounded between 0 and 1
  • Thinks a lower Brier score always implies a lower log-loss
  • Reads log-loss as a percentage or an accuracy
  • Forgets that an exact 0 or 1 makes log-loss infinite
  • Claims the Brier score is undefined at probabilities of 0 or 1
  • Says the two scores are the same metric on a different scale

context

open as a page

What does it mean for a probability scoring rule like the Brier score to be proper?

level: middleimportance: should knowfreq 41%

basics

~20 s

A scoring rule is proper when a forecaster gets their best expected score by reporting the probability they genuinely believe, rather than by shading it. Log-loss and the Brier score are strictly proper; thresholded accuracy is not.

open as a page

Is a log-loss of 0.31 good when the positive class occurs 12% of the time?

level: seniorimportance: should knowfreq 47%

basics

~20 s

On its own the number means nothing. Compare it with the constant forecaster that predicts 0.12 on every row, which scores about 0.367. A log-loss of 0.31 is therefore roughly a 15% improvement on that reference: real, but modest.

open as a page

How would you decide whether log-loss or the Brier score is your team's headline probability metric?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Choose by the decision the probabilities feed. Log-loss when an overconfident mistake is expensive and you want extremes punished hard; the Brier score when you need a bounded, stakeholder-legible number that no single row can dominate.

open as a page