A 2019 credit scorecard now scores a much wider 2021 applicant pool — do you still trust its held-out AUC?
answer
- the number describes the old population
- name the shift, then choose the fix
- reweighting cannot invent missing coverage
- extreme weights destroy effective sample size
- the labels are not mature yet
basics
~20 sNot as a statement about 2021. The 2019 held-out AUC describes the 2019 applicant mix. Re-estimate it under the new mix where the two populations overlap, and cap exposure where 2019 gave no coverage at all.
solid answer
~50 sThe number is not wrong; it answers a question about the 2019 population, and the deployed question is about 2021. First I name the shift I believe I have: marketing widened, so the applicant mix moved while the risk relationship plausibly held — covariate shift, under which the mapping may still be sound but the score must be re-estimated. Where the two populations overlap I reweight the old held-out rows toward the 2021 input distribution, with weights read off a classifier trained to tell 2019 rows from 2021 rows. Two limits matter more than the technique: extreme weights collapse the effective sample size, and applicants outside anything 2019 covered cannot be reweighted at all. Credit outcomes also mature over twelve to eighteen months, so no labelled 2021 holdout exists yet. So I ship with capped exposure on the new segments, keep a randomised holdout policy so unbiased outcomes accrue, and write the assumption down beside the number.
go deeper
Recall that a held-out score belongs to the population it was computed on, so a changed applicant mix means the old number no longer describes today's decisions. Saying that clearly is already a good answer at this level.
Explain why the model can remain valid while its score goes stale under a moved input distribution, and describe reweighting the old evaluation rows toward the new mix as the way to re-estimate performance.
Show that you check coverage and effective sample size before trusting a reweighted number, and that you plan how unbiased outcome data will be collected while labels are still maturing.
Own the launch call under acknowledged uncertainty: how much exposure the new segments get, what randomised holdout is worth its cost, and how the assumption behind the number is documented and revisited when mature labels arrive.
## Restate what the old number means A held-out AUC of, say, 0.82 computed on 2019 applicants is an estimate of one specific quantity: how well the model ranks risk **among applicants drawn from the 2019 population**. It is not a property of the model in isolation. Once the population changes, the estimate has not become inaccurate — it has become an accurate answer to a superseded question. So the honest response is not "yes" or "no" but a re-statement: *that number describes the old applicant mix; here is what I can and cannot say about the new one.* ## Name the shift before choosing a fix Marketing widened, so more people with different profiles applied. The plausible reading is that the input distribution moved while the relationship between an applicant's attributes and their default risk held — an applicant with a given profile is still as risky as an identical 2019 applicant would have been. That is covariate shift, and under it the fitted mapping is not invalidated. But this reading is an assumption, not an observation, and it is the assumption the whole plan rests on. It fails if underwriting policy, pricing, or the macroeconomic environment also changed, because then the outcome rule itself moved and the historical file describes a rule no longer in force. So state the assumption explicitly, and name what would falsify it. ## Re-estimating the score under the new mix Where the two populations overlap, the old held-out rows can be reweighted toward the new input distribution. Each old row gets a weight equal to the ratio of its density under the 2021 applicant distribution to its density under 2019, and the score is recomputed as a weighted average. You do not need the two densities separately: train a classifier to distinguish 2019 rows from 2021 rows and the odds it produces give the ratio directly. Three caveats deserve to be volunteered before the interviewer asks: 1. **Support.** Reweighting can only redistribute evidence you already have. If 2021 brings applicants in a region 2019 never populated, their weight has nothing to attach to, and the reweighted score is silent about exactly the segment you are most worried about. 2. **Variance.** When a few old rows carry enormous weight, the weighted average is effectively computed on a handful of observations. The reweighted estimate can be unbiased and useless at the same time, so report its effective sample size alongside it. 3. **Assumption dependence.** The whole procedure presumes the risk relationship held. If it did not, reweighting produces a confident estimate of the wrong thing. ## The constraint that dominates: label delay Credit outcomes are slow. Whether a 2021 applicant defaults is only known after twelve to eighteen months of loan performance. So the evidence that would settle the question — a labelled holdout drawn from the 2021 population — cannot exist yet, no matter what budget you assign. This reframes the problem from a measurement exercise into a decision under acknowledged uncertainty. The available moves are about how much you risk while the evidence accrues: - **Cap exposure** where the model is extrapolating: lower limits, tighter cutoffs, or manual review for applicants in the regions 2019 did not cover. - **Keep a randomised holdout policy** so a slice of decisions is made in a way that produces unbiased outcome data later. Otherwise the only 2021 outcomes you ever observe are for applicants the model approved, and the future evaluation set is selected by the model itself. - **Use early proxies with care.** Early-stage delinquency arrives sooner than final default and is genuinely informative, but it is a different target with a different base rate, and treating it as the real label quietly changes the question again. - **Re-fit once mature labels exist** on the new population, rather than tuning on the old file forever. ## What a strong answer sounds like A weak answer is binary: "no, retrain." Retrain on what? There are no mature 2021 labels. The strong answer sequences it: state what the old number does and does not describe; name the shift you believe you have and what would disprove it; re-estimate under the new mix where overlap permits, reporting how thin that estimate is; identify the region where no evidence exists and control risk there directly; and design today's decisions so that in a year you own unbiased outcome data instead of a self-selected sample. The underlying discipline is the same one the i.i.d. assumption always demands: a score is a claim with conditions attached, and the leader's job is to keep those conditions attached to the number wherever it travels.
- How would you estimate the reweighting factors without knowing either input density?Train a classifier to tell 2019 rows from 2021 rows using the features alone. Its predicted odds for a row are proportional to the ratio of the two input densities, which is exactly the weight required. Its overall separability is also a useful summary: if the two years are trivially distinguishable, the populations have moved a long way apart.
- Why can a reweighted estimate be unbiased and still worthless?Because the weights determine how much of the sample actually contributes. If a handful of old rows carry most of the weight, the weighted average is computed on effectively a few observations and its variance is enormous. Always report the effective sample size implied by the weights next to the reweighted score.
- Why does approving only what the model likes contaminate next year's evaluation?Because outcomes are observed only for approved applicants, so the future labelled set is selected by the model's own decisions and carries no information about the rejected region. Holding back a small randomised slice of decisions is what preserves an unbiased evaluation sample, and it is worth the cost precisely because that region is where the model is least tested.
- What would convince you the risk relationship itself changed rather than just the applicant mix?Comparable applicants behaving differently: within input strata the two populations share, mature outcome rates in 2021 diverging from 2019. Corroborating evidence would be a datable external change — an underwriting policy revision, a pricing change, a macroeconomic shock — that plausibly rewrote the relationship rather than merely who applied.
saying these in an interview costs you the question
- Answers just retrain without asking whether labels exist yet
- Treats the old held-out score as a property of the model
- Reweights across regions the old data never covered
- Ignores the effective sample size implied by extreme weights
- Assumes the relationship held without saying it is an assumption
- Plans a future evaluation only on applicants the model approved