Is a log-loss of 0.31 good when the positive class occurs 12% of the time?
answer
- compared with what, exactly?
- the do-nothing forecaster on the same rows
- a formula in the base rate alone
- 0.12 and 0.88, both terms
- report loss over reference loss
basics
~20 sOn its own the number means nothing. Compare it with the constant forecaster that predicts 0.12 on every row, which scores about 0.367. A log-loss of 0.31 is therefore roughly a 15% improvement on that reference: real, but modest.
solid answer
~50 sLog-loss has no absolute scale, so the only honest reading is against a reference forecaster on the same data. The natural reference is the constant base-rate forecaster: predict 0.12 every time. Its log-loss is `-(0.12*ln(0.12) + 0.88*ln(0.88)) = 0.367`. So 0.31 beats it, with a skill score of `1 - 0.31/0.367 = 0.155`, meaning the model removes about 15% of the reference's loss. That is a real but unremarkable model, and I would want the baseline printed next to the metric in every report. Two traps follow. First, log-loss falls automatically as the base rate gets more extreme, so a smaller number on a rarer target is not a better model. Second, comparing this month's 0.31 to last month's number is meaningless if the base rate drifted; recompute the baseline for each period and compare the skill, not the raw loss.
go deeper
Remember that log-loss has no absolute scale and that the simplest yardstick is a model that always predicts the overall positive rate. Know that lower is better and 0 is perfect.
Be able to compute the reference: the constant base-rate forecaster's log-loss is minus b log b minus (1 - b) log(1 - b), which is 0.367 at a 12% rate. Then express the model as a ratio against it.
Show you have been burned by this in monitoring: base-rate drift moves the reference, so a falling log-loss can hide a degrading model. Track skill against a rolling baseline and against the incumbent, not the raw number.
Set the reporting standard for the org: no proper score ships without its evaluation base rate, its reference loss and a skill score, so that no team can present a smaller number on an easier problem as progress.
## Why the raw number cannot be judged Log-loss is a mean penalty in log units. It has a floor of 0 for perfect forecasts and no ceiling, and crucially the *achievable* value depends entirely on how predictable the target is and how common it is. A 0.31 might be outstanding on one dataset and worse than doing nothing on another. So the first thing to do with any reported log-loss is ask: compared with what? ## The constant base-rate forecaster The standard reference is the dumbest non-degenerate model available: ignore every feature and predict the overall positive rate `b` on every row. Because that forecast is the same everywhere, its expected log-loss is `-(b*log(b) + (1 - b)*log(1 - b))` With `b = 0.12`: - `ln(0.12) = -2.1203`, so the positive rows contribute `0.12 * 2.1203 = 0.2544` - `ln(0.88) = -0.1278`, so the negative rows contribute `0.88 * 0.1278 = 0.1125` - total: **0.367** So the model at 0.31 is better than the do-nothing forecaster, but not by much. A convenient way to express this is a **skill score**: `1 - (model loss) / (reference loss)`, here `1 - 0.31/0.367 = 0.155`. Read it as "the model removes 15.5% of the reference's loss". Zero means no skill, 1 means perfect, and negative means the features are actively hurting you. The same discipline applies to the Brier score. The constant base-rate forecaster's Brier score is `b*(1 - b)^2 + (1 - b)*b^2 = b*(1 - b)`, which at `b = 0.12` is **0.1056**. Any Brier score above that is a model doing worse than a single number. ## Trap 1: the base rate moves the yardstick The reference loss shrinks as the base rate gets more extreme. At a 12% rate the floor-setting reference is 0.367; at a 9% rate it is 0.303; at a 1% rate it is 0.056. So a model reporting a log-loss of 0.10 on a 1% target is *worse than useless* even though 0.10 looks like a much smaller number than 0.31. Never compare log-losses across datasets or across classes with different prevalences. ## Trap 2: drift makes month-over-month comparisons lie Suppose last month the model reported 0.31 at a 12% base rate and this month it reports 0.26 at a 9% base rate. The raw number improved, but the reference also fell from 0.367 to 0.303. Skill went from `1 - 0.31/0.367 = 0.155` to `1 - 0.26/0.303 = 0.14`. The model actually got slightly *worse* at the job; all that happened is that the world got easier to predict. This is one of the most common ways an evaluation dashboard tells a comfortable lie, and it is why monitoring should track skill against a rolling baseline rather than the raw loss. ## Other references worth having The constant base-rate forecaster is the floor, not the bar. Two more useful comparisons: - **The incumbent.** If a model, a rule set or a human process is already in production, its score on the same rows is the number that decides whether the new model ships. - **A deliberately simple model.** A single-feature or shallow model gives a sense of how much of the achievable signal is cheap. If your ensemble's skill barely exceeds it, the complexity is not paying for itself. Also note the reference must be computed on the *same* evaluation rows as the model, with the base rate measured on those rows, not on the training set. Computing the baseline from a different prevalence quietly biases the comparison. ## How to report it Never ship a bare proper score. Report the model's loss, the base rate of the evaluation set, the reference loss implied by that base rate, and the skill score. Four numbers, one line, and any reader can tell in a second whether the model is worth anything. That habit also protects you in the review where someone asks whether 0.31 is good, because the answer is already on the slide.
- What Brier score would that same constant 0.12 forecaster achieve?`b*(1 - b) = 0.12 * 0.88 = 0.1056`. The algebra is the same idea as for log-loss: with a single constant forecast, the squared error is `(1 - b)^2` on positive rows and `b^2` on negative ones, and averaging with weights `b` and `1 - b` collapses to `b*(1 - b)`. Any model scoring above 0.1056 is beaten by a single number.
- The monthly log-loss fell from 0.31 to 0.26; has the model improved?Not necessarily. If the base rate fell from 12% to 9% over the same period, the constant-forecaster reference fell from 0.367 to 0.303, so skill went from 0.155 to 0.14. The target got easier while the model got slightly worse. Always recompute the reference for each evaluation window and compare skill scores rather than raw losses.
- Can a model's skill score against the base-rate reference come out negative?Yes, and it is a useful alarm. A negative skill score means the model's log-loss exceeds that of predicting the base rate on every row, which usually signals badly overconfident probabilities, a train/serve prevalence mismatch, or scores that were never mapped to probabilities at all.
saying these in an interview costs you the question
- Judges a log-loss value as good or bad in isolation
- Compares log-loss across datasets with different base rates
- Thinks a smaller log-loss on a rarer event means a better model
- Uses the training-set base rate as the reference
- Treats 0.31 as meaning 31% of predictions were wrong