Why does exponentiating a log-price model's prediction land on the median, not the mean?
answer
- exp is not a linear operator
- averaging happens between log and exp
- monotone maps carry quantiles, not means
- Jensen's inequality points one way
- multiply by exp(sigma squared over two)
basics
~20 sExponentiating is a convex map, so the average of log-scale predictions does not map back to the average price. Under symmetric log-scale errors it returns the conditional median; with normal log errors of standard deviation s, multiply by exp(s^2 / 2) to recover the mean.
solid answer
~50 sWhen you train on `log(price)`, the model estimates the conditional mean of the log, not of the price. Exponentiating undoes the transform on a single point, but not on an average: because exp is convex, `E[exp(Z)] >= exp(E[Z])` by Jensen's inequality, so `exp(prediction)` sits below the conditional mean of price. What it actually estimates, when log-scale residuals are symmetric, is the conditional median. If those residuals are roughly normal with standard deviation s, the mean is recovered by multiplying by `exp(s^2 / 2)`. Duan's smearing estimator is the distribution-free alternative: multiply by the average of `exp(residual)` over the training residuals, which needs no normality assumption. The first decision is which quantity you actually want. For 'what is a typical apartment worth' the median is arguably the better answer; for 'what will this portfolio of 40,000 listings total' you need the mean, and the uncorrected sum will be systematically short.
code
python · 15 linesimport math, random, statistics
random.seed(0)
sigma = 0.5 # spread of the log-scale error
# true world: log(price) = 12 + e, e ~ Normal(0, sigma^2)
prices = [math.exp(12 + random.gauss(0, sigma)) for _ in range(200000)]
log_scale_mean = statistics.fmean(math.log(p) for p in prices)
naive = math.exp(log_scale_mean) # what exp(prediction) gives you
corrected = naive * math.exp(sigma ** 2 / 2) # log-normal mean correction
print(round(naive)) # 163023 -> the MEDIAN price
print(round(corrected)) # 184730 -> the MEAN estimate
print(round(statistics.fmean(prices))) # 184690 -> actual mean, matches
print(round(statistics.median(prices))) # 162928 -> actual median, matchesgo deeper
Recall that a model trained on the log of a target does not predict the target directly, and that you must exponentiate to get back — and that this step is not free of consequences.
Explain that exp is convex, so exponentiating an average of logs lands on the median and understates the mean, and be able to state the exp(s^2 / 2) correction and where s comes from.
Show that you check the assumptions before correcting: residual normality and constant spread, out-of-sample estimation of s, and the smearing alternative when the shape is unknown.
Own the framing decision: whether the business needs a median or a mean, whether a log target plus correction beats modelling the original scale with a log-link or gamma objective, and what the team's convention should be.
## Why anyone logs the target Apartment sale prices are right-skewed and multiplicative: an error of 20,000 on a 100,000 flat is a disaster, the same error on a 2,000,000 penthouse is noise. Fitting squared error on the raw price makes the expensive rows dominate the loss, because their residuals are simply larger. Fitting squared error on `log(price)` changes the objective to something closer to relative error, gives the cheap and expensive segments comparable influence, and often makes the residual spread roughly constant across the price range. That is a real modelling gain, and it is why the log target is so common in pricing, demand and claim-size work. ## What the model then predicts A squared-error model trained on `log(price)` estimates `E[log(price) | x]`. Note carefully what that is not: it is not `log(E[price | x])`. The two differ, and the gap is the whole subject here. Write the fitted model as `log(y) = f(x) + e`. Exponentiating gives `y = exp(f(x)) * exp(e)`. Taking the conditional expectation: ``` E[y | x] = exp(f(x)) * E[exp(e)] ``` and `E[exp(e)] > 1` for any non-degenerate error, because exp is convex and `E[exp(e)] >= exp(E[e]) = exp(0) = 1` (Jensen's inequality). So `exp(f(x))` is systematically **below** the conditional mean. What it does hit is the conditional median. If e is symmetric about zero, then the median of `f(x) + e` is `f(x)`, and because exp is strictly increasing it carries medians through unchanged: the median of y is `exp(f(x))`. Monotone transforms preserve quantiles; they do not preserve means. That single sentence is the whole mechanism. ## Correcting back to the mean **Parametric correction.** If the log-scale residuals are normal with standard deviation s, then `exp(e)` is log-normal with mean `exp(s^2 / 2)`, so: ``` mean estimate = exp(f(x)) * exp(s^2 / 2) ``` With s = 0.5 the factor is `exp(0.125)` = 1.133, a systematic 13% shortfall if you skip it. With s = 0.3 it is about 1.046. Estimate s from the residual standard deviation of the log-scale fit, measured out of sample rather than on the training residuals the model has already shrunk. **Distribution-free correction.** Duan's smearing estimator drops the normality assumption. Take the residuals `e_i` from the log-scale fit, and multiply the back-transformed prediction by their empirical mean of exponentials: ``` smearing factor = average of exp(e_i) over the fitted residuals mean estimate = exp(f(x)) * smearing factor ``` Both corrections assume the residual spread does not vary with x. If it does — cheap flats predicted far less precisely than expensive ones, say — a single global factor is wrong, and you need a factor per segment or a model that targets the mean on the original scale directly. ## When the bias actually hurts **Summing predictions.** This is the big one. A 10% per-row shortfall is invisible in a scatter plot and catastrophic in a total: predicted portfolio value, expected revenue, expected claim reserve. Any figure produced by adding predictions together needs the mean, and therefore needs the correction. **Comparing to a raw-scale baseline.** A log-target model back-transformed without correction will look worse than a raw-scale model on any mean-based error measure, purely because it is aiming at a different quantity. **Reporting a typical value.** If the deliverable is 'what does a flat like this go for', the median is defensible and arguably preferable — the untouched back-transform is then correct, provided you say which quantity you are reporting. ## The alternative to transforming at all If you need the conditional mean of a positive skewed target and you dislike the correction bookkeeping, model the original scale with a loss that suits multiplicative error rather than logging the target. Generalised linear models with a log link, or gradient-boosting trained with a Poisson or gamma objective, estimate `E[y | x]` directly on the original scale, with no back-transform step and no smearing factor. The tradeoff is that the log-target route is simpler and every squared-error learner can do it. ## What a strong answer sounds like Name the quantity the log-scale model estimates, state that exp carries medians and not means, give either correction factor with its assumption, and close by asking what the prediction is for. Candidates who say 'log and exp cancel out' have missed that averaging happens between the two operations, which is precisely where the bias enters.
- Why does this bias matter far more when you sum predictions than when you inspect one?Because the shortfall is systematic, not random. Every row is biased downward by roughly the same factor, so errors do not cancel when you add them. A 13% per-row gap on 40,000 listings becomes a 13% miss on the portfolio total — a budget or reserve figure that is wrong in a single, consistent direction.
- What does Duan's smearing estimator give you that the exp(s^2 / 2) factor does not?Freedom from the normality assumption. Smearing multiplies the back-transformed prediction by the average of exp(residual) taken over the observed log-scale residuals, so it works whatever shape those residuals have. The exp(s^2 / 2) factor is exact only when the log-scale errors are normal and their spread is constant across the range.
- When is the uncorrected back-transform the answer you actually want?When the deliverable is a typical value rather than an expected total — a suggested listing price, a benchmark rent, a headline 'flats like this go for'. Those are median questions, and exp(prediction) estimates the median directly. The requirement is only that you state which quantity you are reporting.
- How would you avoid the back-transform question entirely?Model the original scale with a loss built for multiplicative, positive targets — a log-link generalised linear model, or gradient boosting with a Poisson or gamma objective. Those estimate the conditional mean on the original scale directly, so there is no transform to undo and no smearing factor to maintain.
saying these in an interview costs you the question
- Says log and exp cancel, so back-transformed predictions are unbiased
- Applies exp(s^2 / 2) without checking residuals are near-normal and evenly spread
- Reports log-scale error and describes it as the price error
- Treats the shortfall as a model bug rather than a scale artefact
- Sums uncorrected back-transformed predictions into a revenue total
- Estimates the correction from training residuals the model has already overfitted