skip to content

Why can a regression model's good overall MAE hide a cohort where it is badly broken?

level: middleimportance: must knowfreq 68%

answer

  1. an average hides the minority
  2. weighted by row count
  3. a 3% cohort barely moves it
  4. compute the metric inside each cohort
  5. worst slice reported beside the headline

basics

~20 s

Overall MAE is a row-weighted average, so a small cohort can run three times worse and barely move it: a cohort holding 3% of rows at triple the error lifts overall MAE by only about 6%. Cut the metric by cohort.

solid answer

~50 s

MAE and RMSE are averages over rows, so they are dominated by whichever cohort supplies most of the rows. On a delivery-ETA model, if one metro is 3% of deliveries and its MAE is three times the rest, the overall MAE comes out only about 6% higher than it would be if that metro were healthy — well inside the swing I would tolerate between two retrains of the same pipeline. So I never report a single number. I compute the same metric inside each business cohort — city, customer tenure band, order size, time of day — alongside the row count for each, and I put the worst cohorts next to the headline. That turns `the model is fine` into `the model is fine except in one metro, which is 3% of volume and most of the complaints`, which is a sentence the business can act on.

go deeper

for a junior

Be ready to say that MAE and RMSE are averages over rows, and that a small group of rows can be much worse without changing the average. Know that the fix is to recompute the same metric inside each group.

for a middle

You should be able to do the arithmetic on the spot: a 3% cohort at triple the error moves the overall MAE by roughly 6%. Name the cohorts you would cut by for the domain in the question and say why each one.

for a senior

Show the reporting discipline: a standing slice table with counts, worst cohorts beside the headline, and a release gate that blocks a per-cohort regression even when the overall metric improves. Talk about how you prioritise the cohorts you find.

for a principal

Own the question of what the organisation is accountable for: whether the shipping bar is the average request or the worst cohort, who owns each cohort's number, and how you stop the slice suite from growing into a dashboard nobody reads.

## The metric is an average, and averages hide minorities Every headline regression metric — mean absolute error, root mean squared error, mean absolute percentage error — is computed by walking every row of the evaluation set, turning each row into a per-row error, and averaging. The number you report is therefore **weighted by how many rows each part of the population contributes**. A part of the population that supplies few rows can be arbitrarily broken and still barely move the headline. The arithmetic is worth doing once. Suppose the bulk of the test set has `MAE = m`, and one cohort holding 3% of the rows has `MAE = 3m`. Then `overall = 0.97 * m + 0.03 * 3m = 1.06 * m` Six percent. A model that is three times worse for one in every thirty-three requests shows up in the release report as six percent worse — the same order as the noise between two retrains of an unchanged pipeline. Halve the cohort to 1.5% of rows and the headline moves 3%; take it to 0.5% and it moves 1%, which is invisible. ## What a slice is here A slice (or cohort) is any subset of the evaluation rows defined by something you can name **before** you look at the errors: a city or region, a customer-tenure band, a product category, an order-size decile, weekday versus weekend, new versus returning, acquisition channel, delivery partner. A slice can also be defined by the model's own output — the rows where the prediction is under ten minutes, ten to thirty minutes, and so on — which is the tabular form of a residual-versus-fitted inspection. Slicing is not a statistical trick. It is the recognition that one model is being asked to serve heterogeneous sub-populations, and that its error is almost never uniform across them. ## Why error concentrates in particular cohorts - **Training volume.** A cohort with few training rows — newly launched products, a market opened last quarter — is fit mostly by whatever the rest of the data implies about it. - **A different generating process.** A metro with tolls, ferries and dense one-way systems behaves unlike the suburban rows that dominate the fit. - **Feature availability.** A cohort where a strong feature is systematically missing falls back on weaker signal. - **Target scale.** Cohorts whose targets are naturally larger carry naturally larger absolute error even when relative accuracy matches. - **Label quality.** Some cohorts have noisier or later-arriving ground truth. Only the first three are model defects. The last two mean the cohort is intrinsically harder, which is why a slice report should carry the slice's own baseline next to the model's number. ## What to compute and report For each slice: the row count, the metric, ideally an uncertainty interval, and the slice's own trivial baseline. Then report the headline plus the worst few slices, and make both a release gate — for example, no named cohort may regress by more than a stated percentage even if the overall metric improves. Two weightings answer different questions and it is often worth showing both: - The **row-weighted** overall metric answers *what error should a randomly chosen request expect?* - An **unweighted average across cohorts** answers *what error should a randomly chosen cohort expect?* This is the right frame when every cohort has an owner — every city has a manager — and each one's experience counts equally regardless of volume. ## Prioritising what you find A weak cohort is not automatically the top of the backlog. Rank by `slice error excess * slice volume * unit cost of an error`, and add the direction of travel: a 1% cohort growing 20% a month is a different proposition from a 1% cohort that is shrinking. A small cohort with triple the error can still be the top complaint driver if errors there are expensive, and can equally be a rounding error if they are not. ## The comparison this protects The failure it prevents is the silent regression: candidate model B has a lower overall RMSE than the incumbent A, ships, and three cohorts get materially worse while the volume-dominant cohort improves a little. Only a per-slice comparison of A against B surfaces that before release. ## What interviewers listen for That you know the metric is row-weighted and can do the arithmetic that shows how little a small cohort moves it; that you can name concrete cohorts for the domain in front of you rather than saying `slice the data`; that you always carry the row count next to a slice metric; and that you keep this separate from the fairness question — asking whether a model is weakest somewhere is a different exercise, with different machinery, from asking whether it treats protected groups unequally.

  • The worst cohort is also the smallest — how do you decide whether it is worth fixing?
    Rank by excess error times volume times the cost of an error there, then look at the trend. A cohort at 0.5% of volume can still top the list if mistakes are expensive or if it is doubling every quarter. A stable, cheap, tiny cohort gets documented as a known limitation instead of a fix.
  • How do you choose which dimensions to cut by?
    Take the dimensions the business already manages by — geography, tenure, product line, channel, partner — plus the ones where you suspect thin data, such as newly launched items, plus bands of the model's own prediction. Fix that list before you look at any errors, so you are reporting a standing view rather than mining for a story.
  • Should the headline metric ever be re-weighted so small cohorts count more?
    Rather than re-weighting the headline, report both: the row-weighted number, which is the experience of a random request, and an unweighted average over cohorts, which is the experience of a random cohort. Re-weighting a single number silently changes what it means, and readers will keep interpreting it as the first one.

A national average temperature tells you nothing about the one town that is freezing. You only see it in the per-town numbers.

saying these in an interview costs you the question

  • Reports one overall RMSE and calls the model validated
  • Assumes a low average means every segment is healthy
  • Quotes a slice error without its row count
  • Treats the worst cohort as top priority regardless of volume or cost
  • Compares two model versions on the headline metric alone

context