Your conformal intervals hit 90% coverage overall but only 74% on the oldest patient cohort — why?
answer
- the promise is an average over everyone
- one global quantile, one width for all
- noisier subgroup eats the error budget
- per-group calibration or input-varying width
- exact per-input coverage is provably impossible
basics
~20 sConformal prediction guarantees marginal coverage: 90% averaged over the whole population, not inside every subgroup. One global quantile under-covers groups whose errors are larger and over-covers the easy ones. Group-wise calibration or an input-adaptive score is the fix.
solid answer
~50 sThe guarantee is marginal — the 90% is averaged over the draw of the test point across the entire population. With an absolute-residual score you take a single quantile and hand every prediction the same half-width, so any cohort whose outcomes are genuinely noisier will have residuals concentrated in the upper tail and will absorb more than its share of the miscoverage; the easy cohorts silently run at 95%. Nothing is broken and nothing was implemented wrong. Three practical responses: calibrate separately within each pre-declared group so each gets its own quantile; normalize the score by a predicted error scale so the width tracks difficulty; or conformalize a predicted band whose width already varies with the input. And note the theory: exact coverage conditional on every possible input is impossible distribution-free without infinitely wide intervals, so audit the segments you care about rather than hoping for universal conditional validity.
go deeper
Learn the distinction in words: an average over the whole population is not a promise about each slice of it. If a headline metric is 90% and a subgroup is at 74%, that is arithmetic, not a bug.
Explain the mechanism: a single quantile from absolute residuals yields one constant width, so noisier inputs fall outside more often. Know that swapping the nonconformity score changes adaptivity without breaking validity.
Demonstrate the operational habit of auditing coverage per segment before launch, and name a concrete fix you have reasoned through — group-wise calibration, a normalized score, or conformalizing a predicted band — with its data cost.
Own the decision of which conditioning events the organisation guarantees. Conditional coverage everywhere is impossible, so the real deliverable is a declared list of protected segments, the calibration data budget they require, and what is merely monitored.
## Marginal means averaged When split conformal promises `P(Y in C(X)) >= 1 - alpha`, the probability is over the joint draw of the calibration set and the new pair `(X, Y)`. Read plainly: out of all future cases, about 90% are covered. It says nothing about whether that 90% is spread evenly. A model can hit exactly 90% overall while covering 95% of one cohort and 74% of another, and the arithmetic is perfectly consistent if the under-covered cohort is a minority of the traffic. **Conditional coverage** is the stronger property you actually wanted: `P(Y in C(X) | X = x) >= 1 - alpha` for every input, or at least `P(Y in C(X) | group = g) >= 1 - alpha` for each group `g`. Marginal coverage does not imply either. ## Why the oldest cohort loses With the absolute-residual score you compute one number `q` and issue `f(x) +/- q` to everyone. That is a constant-width interval. Suppose outcomes for the oldest cohort are intrinsically more variable — more comorbidities, sparser training data, more measurement noise. Their absolute residuals sit disproportionately in the upper tail of the calibration score distribution, above the global `q`. Because they are a minority of the calibration rows, the global quantile is dominated by the easier majority. The result is a band that is comfortably too wide for young patients and too narrow for old ones, averaging out to the promised 90%. The same phenomenon appears whenever the difficulty of the prediction varies with the input — small versus large order values, dense versus sparse regions of feature space, a rare category with fifty training rows. ## Fix 1 — group-conditional (Mondrian) calibration Partition the calibration set by a group label that is computable at prediction time, then take a separate quantile within each partition. Each group now carries its own finite-sample guarantee. The costs are concrete: every group needs enough calibration rows for its own quantile to exist and be stable (at `alpha = 0.10` a group with 12 rows technically has a quantile, but a wildly variable one), and the groups must be declared in advance from features you will have at inference. You cannot slice retrospectively on whatever subgroup someone complains about and still claim a guarantee there. ## Fix 2 — normalize the score Keep one global quantile but divide the residual by a predicted error scale: `s = |y - f(x)| / sigma(x)`, where `sigma` is a second model trained to predict how large the error at `x` typically is. The final interval becomes `f(x) +/- q * sigma(x)`, so width tracks difficulty. Validity is untouched — you may swap the score freely — and adaptivity improves to the extent that `sigma` is any good. If `sigma` is garbage, you get valid intervals with arbitrary widths, which is a genuine failure mode worth watching for. ## Fix 3 — conformalized quantile regression Instead of conformalizing a point, conformalize a band. Train two models to predict a low and a high quantile of the outcome — say the 5th and 95th percentiles — giving a raw band `[lo(x), hi(x)]` whose width already varies with the input. The nonconformity score becomes the signed distance outside that band, `E = max(lo(x) - y, y - hi(x))`, which is negative when the truth landed inside. Take the usual `ceil((n+1)(1-alpha))` quantile `Q` of those scores and output `[lo(x) - Q, hi(x) + Q]`. If the base quantile models were systematically too narrow, `Q` is positive and inflates them; if they were too wide, `Q` is negative and shrinks them. Coverage is guaranteed marginally either way, and the width adapts because the underlying band did. Concretely, a delivery-time model that would have emitted a bare point of 19 minutes instead emits 15-25 minutes for a hard address and 18-20 for an easy one, with the same overall 90% promise. That difference is what makes an interval usable in a product. ## The impossibility you should know about Exact conditional coverage at every `x`, distribution-free and in finite samples, is provably unattainable: any procedure that guarantees it for all continuous distributions must produce intervals of infinite expected width at almost every point. So the honest posture is not "achieve conditional coverage" but "choose which conditioning events matter, guarantee those, and measure the rest". ## What to do before shipping Compute realised coverage per segment on a fresh held-out sample, not just overall. Report the worst segment alongside the headline number. Decide in advance which segments are protected — clinical cohort, geography, customer tier — and calibrate within them. And resist the reflex of lowering `alpha` globally to rescue one cohort: that widens intervals for everyone and buys the failing group far less than a group-specific quantile would.
- How would you make the guarantee hold specifically for that cohort?Split the calibration set by cohort and take a separate `ceil((n+1)(1-alpha))` quantile inside each one. Each group then carries its own finite-sample guarantee. Two constraints: the group must be defined from features available at prediction time, and each group needs enough calibration rows for its quantile to be stable, so very small groups may have to be merged or handled with a normalized score instead.
- Why not just lower alpha until the worst cohort reaches 90%?Because a global alpha change widens every interval, including the cohorts already over-covered. You pay usefulness across the whole population to fix a minority, and you still have no guarantee for the cohort — you only observed that it happened to reach 90% on one sample. Group-conditional calibration targets the problem directly and comes with an actual guarantee.
- Can you audit coverage on any subgroup you like after the fact?You can measure it, but measuring is not guaranteeing. Slicing after seeing the results invites multiplicity: with enough slices some will look under-covered by chance. Declare the protected segments in advance, calibrate within them, and treat exploratory slices as signals to investigate rather than as violated promises.
saying these in an interview costs you the question
- Reads 90% marginal coverage as 90% for every subgroup
- Concludes the conformal procedure was implemented wrong
- Fixes one cohort by widening intervals for everyone
- Thinks exact per-input conditional coverage is achievable distribution-free
- Defines calibration groups using the label rather than features