What does RMSLE measure that RMSE does not when scoring skewed positive order values?
answer
- the error is taken in log space
- ratios rather than currency gaps
- the plus one admits exact zeros
- big orders stop dominating the score
- negatives cannot be scored at all
basics
~20 sRMSLE is the root mean squared error of log(1 + value), so it scores the ratio between prediction and actual, not the currency gap. Being twice too high on a $20 order costs what it costs on a $2,000 one.
solid answer
~50 s`RMSLE = sqrt(mean((log(1 + pred) - log(1 + actual))^2))`. A difference of logs is a log ratio, so the metric measures proportional error: predicting $40 for a $20 order and $4,000 for a $2,000 order score almost identically, whereas RMSE would treat the second miss as a hundred times worse. On a right-skewed order-value distribution that matters, because plain RMSE is dominated by a handful of very large baskets and a model can win on it while being hopeless on the typical order. The `+ 1` offset lets exact zeros through; negative values cannot be scored at all, so the metric is only for non-negative targets. Two cautions: for the same absolute miss, under-prediction scores slightly worse than over-prediction, and the value itself is in log units, so report a currency error alongside it for anyone who has to act on the number.
go deeper
Know its shape: RMSLE is the root mean squared error computed on log(1 + value) instead of the raw value, and it only works when both targets and predictions are non-negative.
Explain that a difference of logs is a ratio, so the metric scores proportional error. Be able to say why that stops a handful of very large orders from deciding the whole score.
Justify the choice from the cost of being wrong: if a 2x miss hurts equally at every order size, a ratio metric is aligned. Say how you would still surface an absolute error in currency for the people acting on it.
Weigh whether a scale-free number should steer the organisation at all. Ratio errors flatten the tail, which is the wrong call when a small number of large orders carry most of the revenue at risk.
## Definition `RMSLE = sqrt( (1/n) * sum( (log(1 + pred_i) - log(1 + actual_i))^2 ) )` It is the same root-mean-square construction as RMSE, applied to the logarithm of one plus each value rather than to the value. Everything interesting follows from one identity: a difference of logarithms is the logarithm of a ratio, `log(a) - log(b) = log(a/b)`. So RMSLE is a root-mean-square *ratio* error, while RMSE is a root-mean-square *absolute* error. ## What that buys you on skewed positive targets Order values on an e-commerce site typically span two or three orders of magnitude, from a single $8 accessory to a $5,000 appliance order, with a long thin upper tail. Under RMSE: - Missing a $2,000 order by $2,000 contributes 100 times the squared error of missing a $20 order by $200. - Both are the same *relative* mistake: 100% off. - The consequence is that the score is decided almost entirely by the largest orders, and a model that predicts the tail passably while being useless on the body of the distribution can look excellent. Under RMSLE those two errors contribute almost equally, because `log(1 + 4000) - log(1 + 2000)` is about 0.693 and `log(1 + 40) - log(1 + 20)` is about 0.669 - both the log of roughly 2. Every order gets a vote of comparable weight regardless of size. That is the right behaviour when the business cost of being wrong is proportional: a 2x error in a recommended reorder quantity or a delivery-fee estimate hurts the same at any basket size. A second, related effect: taking logs compresses the upper tail, so a single enormous outlying order cannot dominate the metric the way it dominates RMSE. This is a property of the metric only - what you do to the target before fitting is a separate decision. ## The `+ 1` and the sign restriction The offset exists so that a value of exactly zero is representable: `log(1 + 0) = 0`. Zero-value rows are therefore fine. Negative values are not - the logarithm of a non-positive number is undefined, so a target column containing refunds, adjustments or net amounts that can go below zero rules the metric out entirely. Predictions must also be non-negative for the same reason, which is a constraint on the model, not just on the data. The offset also means the metric is not purely a ratio measure at the bottom of the range. For values that are small relative to 1, adding one dominates, so the distinction between an actual of 0.2 and 0.4 is heavily damped. On dollars this rarely matters; on a target measured in small fractional units it can matter a lot, and rescaling the units changes the score - a genuine awkwardness of the metric. ## The asymmetry For the same absolute miss on the same actual, under-prediction scores slightly worse. With an actual of 1,000, predicting 900 gives a log difference of about 0.105, while predicting 1,100 gives about 0.095. The effect is mild at small relative errors and grows as the miss grows, because a prediction can be arbitrarily many times too high but at most one times too low. If your problem has a strong preference in one direction, do not rely on this weak tilt to express it - it is a side effect of the log, not a designed cost model. ## Reporting RMSLE's value lives in log space. An RMSLE of 0.35 corresponds to a typical multiplicative error of roughly `exp(0.35)`, about 1.42, so "typically around 40% off in ratio terms" - but no stakeholder derives that unaided, and it is not dollars. Use RMSLE for model selection when proportional error is the thing you care about, and put an absolute error in the target's own currency next to it in any report a human acts on. The pair is honest: one number chooses the model, the other tells the business what to expect. ## When not to use it If the largest orders carry most of the revenue at risk, a scale-free metric is actively wrong: it flattens exactly the rows that matter, and a model tuned on it will happily trade accuracy on $5,000 baskets for accuracy on $8 ones. Choose the ratio metric when every order counts equally, and an absolute-error metric when dollars count equally.
- Why is RMSLE awkward to report to a business stakeholder?Its value is in log units, so 0.35 is neither dollars nor a percentage and nobody can act on it directly. It is fine as a selection criterion when proportional error is what matters, but pair it with an absolute error in the target's own currency for any report that drives a decision.
- What breaks if some target values are negative?The metric cannot be computed at all, since the logarithm of a non-positive number is undefined. The plus-one offset covers exact zeros only. A column of net amounts containing refunds or adjustments therefore rules RMSLE out, and predictions must stay non-negative for the same reason.
- When would a ratio metric be the wrong choice on order values?When the large orders carry most of the revenue at risk. A scale-free metric deliberately flattens the tail, so a model tuned on it will trade accuracy on the few $5,000 baskets for accuracy on many $8 ones. If dollars are what count, score in dollars.
Working in logs turns "twice as much" into a fixed step on the ruler, so a $20 order and a $2,000 order are measured against the same yardstick rather than against the size of their own price tag.
saying these in an interview costs you the question
- Describes RMSLE as RMSE on a rescaled target
- Claims RMSLE can handle negative target values
- Quotes an RMSLE of 0.4 as if it were a currency amount
- Assumes RMSLE penalises over and under prediction equally
- Uses a ratio metric when the largest orders carry the risk