skip to content

How do covariate shift, label shift and concept drift differ in what they break?

level: middleimportance: must knowfreq 54%

answer

  1. which factor of the joint moved
  2. p(y|x) times p(x), or p(x|y) times p(y)
  3. inputs moved versus rule moved
  4. only one of the three makes the model wrong
  5. prevalence change leaves the ROC curve alone

basics

~20 s

Covariate shift moves the input distribution p(x) while p(y|x) holds. Label shift moves the class balance p(y) while p(x|y) holds. Concept drift moves p(y|x) itself, which is the only one that makes the model wrong.

solid answer

~40 s

All three are ways the joint distribution `p(x, y)` changes between training and deployment; they differ in which factor moves. **Covariate shift**: `p(x)` changes but `p(y|x)` is stable — a lender widens its marketing and sees a different applicant mix, while the link between attributes and risk is unchanged. The mapping is still right, but the old held-out score was computed under the old input mix. **Label shift**: `p(y)` changes while `p(x|y)` holds — prevalence rises from 2% to 15%, so the same symptoms carry higher posterior risk. Ranking survives, but probabilities are miscalibrated and precision moves with prevalence. **Concept drift**: `p(y|x)` itself changes — a rewritten content policy relabels identical-looking posts. That is the serious one: no reweighting recovers it, because the target function moved. You need new labels.

go deeper

for a junior

Be able to name the three and give one concrete example of each: a changed customer mix, a changed base rate, a changed policy. Getting the names attached to the right examples is most of the credit at this level.

for a middle

Say which factor of the joint distribution moves in each case and what that implies for the fitted model versus the reported metric. The point interviewers listen for is that only concept drift makes the learned mapping wrong.

for a senior

Show how you would tell them apart with the data you actually have, and match each diagnosis to a different response: reweighting and coverage for covariate shift, prior and threshold correction for label shift, new labels for drift.

for a principal

Own the question of how much labelling budget and how much history a changing domain justifies. Decide the refit cadence and the point at which historical labels are retired rather than downweighted, and be able to defend that call on cost.

## One joint distribution, two factorisations Everything supervised learning does concerns a joint distribution `p(x, y)` over inputs and labels. It factorises two ways: - `p(x, y) = p(y | x) * p(x)` — the **discriminative** view: a marginal over inputs, and the labelling rule. - `p(x, y) = p(x | y) * p(y)` — the **generative** view: a class prior, and what members of each class look like. The three named failures are just "which factor moved, and which stayed put". ## Covariate shift: p(x) moves, p(y|x) holds The input population changes; the rule linking inputs to outcomes does not. Canonical case: a consumer-credit scorecard fitted on 2019 applicants is scoring 2021 applicants after the lender widened its marketing. The new applicants skew younger, thinner-file and lower-income. But an applicant with a given set of attributes still defaults at the same rate as an identical 2019 applicant would have — the risk relationship is unchanged, only who shows up is different. What this does: - The learned function is **not wrong**. If you had fitted the true `p(y | x)` perfectly, nothing would need fixing. - Real models are misspecified and fitted on finite data, so they are accurate where training density was high and shaky where it was thin. Shifting mass into the thin regions surfaces exactly those errors. - The **old held-out score no longer describes the new population**, because it averaged errors under the old input mix. Where the two populations overlap, you can re-weight the old held-out rows by the ratio of new-to-old input density to re-estimate the score under the new mix; where the new population reaches inputs the old data never covered, nothing can be re-weighted, because there is no old evidence there at all. ## Label shift: p(y) moves, p(x|y) holds The class balance changes; what each class looks like does not. This is natural when the label causes the observation rather than the other way round. Canonical case: a diagnostic model trained when a condition affected 2% of patients is deployed during an outbreak when it affects 15%. A patient who genuinely has the condition presents the same way as before — `p(x | y)` is stable — but the base rate is seven and a half times higher. What this does: - **Ranking survives.** Because the per-class feature distributions are unchanged, the true-positive and false-positive rates at any fixed threshold are unchanged, so the ROC curve and ROC-AUC are the same. - **Calibrated probabilities do not survive.** A model that learned the 2% prior systematically under-predicts risk at 15%. If the model outputs calibrated probabilities, you can correct it by multiplying each class probability by the ratio of new prior to old prior and renormalising — a cheap fix, provided you can estimate the new prior. - **Precision-flavoured metrics move on their own.** Precision, positive predictive value and the precision-recall curve all depend on prevalence, so a rising base rate raises precision at a fixed threshold with no change in the model whatsoever. Reporting that as an improvement is a classic error. ## Concept drift: p(y|x) moves The labelling rule itself changed. The same input now maps to a different outcome distribution. Canonical case: a content-moderation model where the platform rewrites its policy mid-year. The posts look identical; what counts as a violation is different. Historical labels now describe a rule that no longer exists. What this does: - The model is **wrong**, not merely evaluated on the wrong population. There is no reweighting scheme that repairs it, because there is no weighting of the old inputs that reproduces the new labelling rule. - Old labelled data becomes actively misleading for training: it teaches the superseded rule. - Recovery requires **new labels under the new rule**, and the practical question becomes how quickly you can obtain them and how much history you can still reuse. Drift also comes in gradual forms — fraud tactics evolving, language use changing — where the rule erodes rather than flips on a date. The date-stamped policy rewrite is the easy case precisely because you know exactly when the old labels stopped being valid. ## Diagnosing which one you have A useful ordering in an interview: 1. Did the **inputs** move? Compare the input distributions of the old and new populations, feature by feature and jointly. 2. Did the **outcome rate** move once you condition on inputs? If overall outcome rates moved but the rate within comparable input strata did not, you are looking at label or covariate shift, not drift. 3. Did the **rule** move? If comparable cases now get different outcomes, it is concept drift — and that usually correlates with a datable external event: a policy change, a pricing change, a regulatory change. The distinction matters because the responses differ sharply. Covariate shift argues for reweighting or for collecting data in the newly reached region. Label shift argues for a prior correction and for re-choosing thresholds. Concept drift argues for new labels and a refit — and for treating the historical file as evidence about a world that no longer exists.

  • Under pure label shift, why does ROC-AUC stay put while precision moves?
    ROC-AUC is built from true-positive and false-positive rates, which are computed within each class. Label shift leaves the per-class feature distributions `p(x|y)` unchanged, so those rates and the whole ROC curve are unchanged. Precision mixes the classes — it asks what share of flagged cases are truly positive — so it rises purely because positives became more common.
  • If covariate shift leaves p(y|x) intact, why do models degrade under it at all?
    Because a fitted model is not the true `p(y|x)`. It is an approximation whose errors were minimised under the training input density, so it is sharpest where training data was dense and crude where it was thin. Shifting mass into thin regions exposes the approximation error, and any region the training data never covered is pure extrapolation.
  • Can you have covariate shift and concept drift at the same time?
    Yes, and it is common. A lender that widens its marketing also changes its underwriting policy: the applicant mix moves and the outcome rule moves together. Disentangling them needs labelled data from the new period, because the only way to tell whether comparable applicants now behave differently is to observe their outcomes.
  • Is a falling model score by itself evidence of concept drift?
    No. A score can fall because the input mix moved into regions the model fits poorly, because the class balance moved and the metric is prevalence-sensitive, or because the labelling rule changed. Only the last is drift. Check whether comparable cases still get comparable outcomes before you conclude the relationship moved.

saying these in an interview costs you the question

  • Uses drift as a blanket word for any distribution change
  • Claims retraining on old labels fixes concept drift
  • Thinks covariate shift always requires a new model
  • Reports rising precision after a prevalence rise as an improvement
  • Assumes reweighting can repair a changed labelling rule
  • Says label shift changes ROC-AUC

context