Before shipping a regression model, what baseline must it beat and why?
answer
- a number needs a reference point
- predict one constant for everyone
- which constant depends on the metric
- median for MAE, mean for squared error
- also beat today's rule of thumb
basics
~20 sTwo: the best constant prediction — the target's median for MAE, its mean for squared error — and the rule of thumb the business already runs. Both on the same held-out rows, they turn a bare error number into a decision.
solid answer
~50 sAn error number means nothing on its own; an MAE of 4.2 minutes could be excellent or embarrassing. So I score baselines on exactly the same held-out rows and the same metric. The first is the best constant: predict one number for everyone, the training median if I am scoring MAE and the training mean if I am scoring squared error, because those are the constants each metric is minimised by. The second is the incumbent — the rule already in use, such as `quote this store's median basket size` or a distance-over-average-speed estimate. If the model does not clearly beat both, there is nothing to ship: the rule is cheaper, already trusted, and has no training pipeline to operate. I also run the baselines inside each cohort, because a model can beat one global constant while losing to a per-store constant.
go deeper
Be ready to name a trivial baseline unprompted and to say which constant goes with which metric: median for mean absolute error, mean for squared error. Know that both model and baseline are scored on the same held-out rows.
Explain why each constant is optimal in a line of reasoning, and go beyond the global constant to the per-group constant. Be able to say what it means when a model beats one and loses to the other.
Show that you treat the existing business rule as the real competitor, translate the model-versus-baseline gap into operational cost, and check that the gap is bigger than resampling noise before calling it a win.
Own the shipping bar itself: what lift over the incumbent justifies the ongoing cost of a trained model, when the honest recommendation is to keep the rule, and how baselines stay in the evaluation report permanently rather than appearing once.
## Why a baseline exists at all A regression error has units and no reference point. `MAE = 4.2 minutes` on a delivery-time model is meaningless until you know what four point two minutes is being compared against. Every evaluation of a regression model is implicitly a comparison, and a baseline is the act of making that comparison explicit and cheap. ## Baseline one: the best constant Throw away every feature and predict a single number for every row. Which number is best depends on the metric, and getting this right matters: - **Mean absolute error is minimised by the median.** Nudge the constant `c` upward by a small step and the total absolute error changes by `(number of targets below c) - (number above c)`. That is zero exactly when half the points lie on each side, which is the median. - **Mean squared error, and therefore RMSE, is minimised by the mean.** The derivative of `sum (y - c)^2` with respect to `c` is `-2 * sum (y - c)`, which vanishes when `c` equals the average of the targets. Using the wrong constant makes the baseline artificially weak and flatters your model. On a right-skewed target such as basket size or trip duration, the mean sits above the median, so a mean constant scored under MAE is worse than it needs to be, and your model's apparent lift is inflated. Compute the constant on the **training** data and apply it to the test rows, exactly as you would a model. Taking the median of the test targets leaks information the model never had and makes the baseline unbeatable-looking in the wrong direction. ## Baseline two: the conditional constant A far stronger and still learning-free baseline is one constant per group: `predict this store's median basket size`, `predict this route's historical mean duration`, `predict this category's typical demand`. It captures level differences across the population without fitting anything, and it is often uncomfortably hard to beat. Where the target has a time index, the same role is played by `predict the last observed value` or `predict the same weekday last week`. This baseline is diagnostic as well as competitive. If the model beats the global constant but loses to a per-group constant, it has learned the overall level of the target and not the group structure — usually a feature problem, not a model-capacity problem. ## Baseline three: the incumbent The rule of thumb the business already uses: distance divided by an assumed average speed, a planner's manual estimate, last quarter's number plus ten percent, a spreadsheet someone maintains. This is the true counterfactual, because it is what happens if you ship nothing. It costs nothing to run, is already trusted, and needs no monitoring, retraining or on-call rotation. A model that beats a constant but not the incumbent has no business case, however good its metrics look in isolation. ## How to score baselines honestly Same held-out rows, same metric, same units, same row-exclusion rules, reported side by side in one table. Two extra disciplines: - **Check the gap survives noise.** Resample the test set and recompute both numbers, or compare across folds. A two-percent gap that moves by five percent under resampling is not a result. - **Translate the gap.** Minutes of ETA error become late deliveries, refunds and support contacts; units of demand error become stockouts or overstock. A four-percent improvement can be decisive at high volume and pointless at low volume, and only the translation tells you which. ## Baselines belong inside slices too Run the constant baseline within each cohort and report the model-versus-baseline gap per cohort, not only globally. Two things fall out. First, you find cohorts where the model adds nothing and the constant should simply be used. Second, you separate a hard cohort from a broken one: if a cohort's model error is triple the average but its constant baseline is triple too, the target is intrinsically noisier there and the model's lift is normal. ## What interviewers listen for That you reach for a baseline before quoting an error at all; that you know median goes with absolute error and mean with squared error, and can justify it in one line; that you name the incumbent rule as a baseline, not just the statistical constant; and that you score everything on the same held-out rows rather than comparing a test-set model against a training-set baseline.
- Which constant minimises MAE, and what goes wrong if you use the mean instead?The median. Shifting the constant changes total absolute error by the count of points below it minus the count above, which is zero at the median. Using the mean on a skewed target gives a needlessly bad baseline, so the model looks better than it is — you have measured yourself against a straw man.
- The model beats the global constant but loses to a per-store constant in a third of stores. What now?The global baseline was too weak. The model has learned the overall level but not store-to-store differences, so I would add store identity or store-level aggregate features, or fall back to the per-store constant where it wins, and make the per-store comparison the ship gate rather than the global one.
- How much better than the baseline is enough to ship?Translate the gap into the decision it feeds — minutes of error into late deliveries and refunds, demand error into stockouts — and compare that against the cost of building, serving and monitoring the model. Also require the gap to be larger than the run-to-run variation you see when resampling the test set.
Saying a runner finished in four hours means nothing until you hear that walking it takes six and the current champion takes three.
saying these in an interview costs you the question
- Presents an RMSE with nothing to compare it against
- Uses the mean as the baseline while scoring MAE
- Computes the baseline from the test set's own targets
- Scores the baseline on training rows and the model on test rows
- Ignores the heuristic the business already runs today