Your conformal calibration split was collected before a pricing change — does the 90% guarantee still hold?
answer
- one assumption, and it is not iid
- any ordering equally likely
- the rank argument needs a stable population
- a known intervention voids the quantile
- recalibrate, or reweight if only inputs moved
basics
~20 sNo. Conformal coverage rests on exchangeability between the calibration rows and the rows you predict on. A pricing change makes post-change traffic a different population, so the stored quantile is stale and realised coverage can drift below 90%.
solid answer
~50 sConformal's only assumption is that the calibration scores and the new score are exchangeable — their joint distribution is unchanged by reordering. The coverage argument is a rank argument: the new score is equally likely to land in any position among the `n+1`. A pricing change breaks that. If it shifts the inputs, the mix of easy and hard cases changes; if it shifts the relationship between inputs and outcome, the residuals themselves come from a different distribution. Either way the stored quantile no longer describes the errors you are about to make. The remedies depend on what moved: recalibrate on post-change rows as soon as enough labels accumulate; use weighted conformal if only the covariate distribution moved and you can estimate the likelihood ratio. Treat the calibration split as perishable and re-run it after any known intervention.
go deeper
Remember that the interval was built from past errors. If the world changed after those errors were recorded, the interval describes the old world and its promise no longer applies.
Define exchangeability precisely and connect it to the rank argument: the new score must be equally likely to land in any position among the calibration scores. Explain why a systematic change breaks that, unlike a reshuffle of rows.
Show the operational instinct: a known intervention invalidates the calibration split by construction, so recalibration is scheduled with the release. Distinguish covariate shift, where reweighting can help, from a changed outcome relationship, where only new labels do.
Own the policy that an uncertainty claim carries an expiry. Decide what the product does during the label-delay window after an intervention, who is allowed to publish a coverage number, and whether the organisation prefers conservative widening or withdrawing the interval.
## What exchangeability actually requires A sequence of random variables is **exchangeable** if their joint distribution is invariant to permutation — every ordering is equally probable. Independent and identically distributed data is exchangeable, but exchangeability is weaker: it permits dependence, as long as the dependence is symmetric across positions. Split conformal needs exactly one thing: that the `n` calibration nonconformity scores plus the score of the new test row are jointly exchangeable. Then the test score is equally likely to occupy any of the `n+1` ranks, and the probability it exceeds the `ceil((n+1)(1-alpha))`-th smallest is at most `alpha`. That single sentence is the whole guarantee. ## Why a pricing change kills it After a pricing change, the post-change rows are not interchangeable with pre-change rows. Two distinct things can move: - **Covariate shift.** The inputs change distribution — different basket sizes, a different customer mix, different geographies responding to the new price. The model itself may still be right about `P(Y | X)`, but the *mix* of easy and hard cases has moved, so the population of residuals you now face is not the population you calibrated on. - **Concept shift.** The relationship between inputs and outcome changes — demand responds differently, the same features now imply a different outcome. Residuals are drawn from an outright different distribution. In both cases the rank argument dissolves. The test score is no longer equally likely to fall in any position among the `n+1`, so nothing bounds the miscoverage. Note the asymmetry that catches people out: the guarantee does not degrade gracefully in proportion to the size of the shift. It is simply void, and whether realised coverage lands at 88% or 61% is an empirical question, not something the theory bounds. ## Ordinary time-ordering is a shift too Even with no intervention, data collected over time is often not exchangeable: seasonality, gradual behaviour change, a model that was itself retrained. A calibration split taken from last quarter and frozen indefinitely is a slow-motion version of the same failure. This is why conformal deployments usually re-derive the quantile on a rolling window rather than once at launch. ## Remedies, matched to the cause **Recalibrate.** The direct fix. Once enough post-change labelled rows exist, recompute the quantile on them alone. This restores an exact guarantee with respect to the new population, which is why it is the first thing to reach for after any deliberate change to the product or the model. The catch is the label delay: if outcomes arrive days later, you have a window where you are flying on a stale quantile and should say so. **Weighted conformal.** If you are confident that only the covariates moved and `P(Y | X)` is stable, you can restore validity by weighting each calibration score by the likelihood ratio of the new covariate distribution to the old, and taking a weighted quantile. This needs that ratio to be estimable — feasible when you know the intervention, hopeless when you do not. It cannot repair a change in `P(Y | X)`, because no reweighting of old labels recovers a relationship that has changed. **Adaptive schemes.** Online conformal methods track realised coverage and adjust the effective alpha up or down, trading the exact finite-sample guarantee for a long-run average coverage target that survives gradual drift. Useful in a streaming setting; not a substitute for recalibrating after a known discrete intervention. **Pool with caution.** Mixing pre- and post-change rows to reach a usable sample size yields a quantile for a population that exists nowhere. It is sometimes the least bad option under label scarcity, but it should be an explicit, temporary, documented choice — not a default. ## Small post-change samples At `alpha = 0.10` you need at least 9 rows for the rank `ceil((n+1)(1-alpha))` to exist at all, but a quantile from 40 rows is valid yet volatile: the coverage you actually realise with that one fixed split scatters widely around 90%. Options while the sample fills up are to deliberately conservatively widen, to keep the old interval while labelling it provisional, or to fall back to a point prediction with the uncertainty claim withdrawn. Silently continuing to publish a "90%" interval built on the pre-change split is the one option that is indefensible. ## The habit to demonstrate Tie recalibration to the change-management process. Any deliberate intervention — a price change, a new upstream feature, a model retrain, a new market — invalidates the calibration split by construction, and you know it happened without needing to detect anything. Schedule the recalibration alongside the release, and state the coverage claim with the date range it was calibrated on.
- Is exchangeability the same as assuming the data are independent and identically distributed?No, it is weaker. Independent and identically distributed data is always exchangeable, but exchangeability only requires that the joint distribution of the scores be unchanged by reordering, which permits symmetric dependence between rows. Conformal exploits exactly that weaker condition, which is why it works for data where a full independence assumption would be uncomfortable.
- Which kind of shift can weighted conformal repair, and which can it not?It repairs pure covariate shift: the inputs move but the conditional relationship between inputs and outcome is stable, so reweighting calibration scores by the likelihood ratio of the new to old input distribution restores validity. It cannot repair a change in that conditional relationship — no reweighting of old labels recovers an outcome rule that has changed; only new labelled data does.
- Only 40 post-change labelled rows exist. What do you publish?A quantile from 40 rows is technically valid but the realised coverage with one such split scatters widely. Publish something honest: either widen deliberately and say the interval is conservative, or withhold the interval and ship the point prediction while the sample fills. What you cannot do is keep printing a 90% claim computed on the pre-change split.
saying these in an interview costs you the question
- Says exchangeability just means the data are independent and identically distributed
- Assumes coverage degrades smoothly in proportion to the shift
- Believes a large calibration set protects against distribution change
- Reweights calibration scores to fix a changed outcome relationship
- Keeps publishing the guarantee because nothing in the code errored