How does split conformal prediction turn a trained regression model into 90% prediction intervals?
answer
- hold out rows the model never saw
- score how surprised the model was
- sort the scores, take a high rank
- rank is ceil((n+1)(1-alpha)), not n(1-alpha)
- same half-width added to every prediction
basics
~20 sSplit conformal holds out a calibration set the model never trained on, scores every row by its absolute residual, and takes the residual sitting at the 90% rank. Each new prediction then gets that one value added and subtracted.
solid answer
~50 sSplit the labelled data into a proper training set and a calibration set. Fit the model on the training portion only, then compute a nonconformity score on each of the `n` calibration rows — for a plain regressor the natural score is the absolute residual `|y - f(x)|`. Sort those scores and take the `k`-th smallest, where `k = ceil((n+1)(1-alpha))`; call it `q`. For any new input the interval is `f(x) +/- q`. Provided the calibration rows and the new row are exchangeable, that interval contains the true value with probability at least `1 - alpha` — in finite samples, for any model, with no assumption about the shape of the error distribution. The costs are real: the guarantee is marginal rather than per-input, the width is the same for every input, and you spent labelled data on calibration.
code
python · 15 linesimport math, random
random.seed(7)
# calibration rows the model never trained on: (true value, model's point prediction)
cal = [(random.gauss(20, 4), 19.0) for _ in range(500)]
scores = sorted(abs(y - yhat) for y, yhat in cal) # nonconformity score
alpha = 0.10
n = len(scores)
k = math.ceil((n + 1) * (1 - alpha)) # rank to take
q = scores[k - 1]
point = 19.0
print("rank k =", k, "| q =", round(q, 2)) # rank k = 451 | q = 6.89
print("interval =", (round(point - q, 1), round(point + q, 1))) # (12.1, 25.9)go deeper
Recall the shape of it: train on one part of the data, measure errors on a held-out part, and use those errors to put a band around future predictions. Know that 90% refers to how often the truth lands inside.
Be ready to write the recipe out: proper training split, nonconformity score, the ceil((n+1)(1-alpha)) rank, and the interval f(x) plus or minus q. Explain why the calibration rows must be unseen by the model.
Show that you know what the guarantee buys and what it does not: finite-sample and distribution-free, but marginal and constant-width with an absolute-residual score. Talk about how large a calibration split you would carve out and why.
Own the framing that conformal is a wrapper, not a modelling improvement. Argue when a distribution-free interval is worth spending labelled data on, versus when a cheaper parametric band or an unquantified point estimate is enough for the decision at hand.
## The problem it solves A regression model returns one number. A decision-maker usually needs a range plus a promise about how often the truth falls inside it. Split conformal prediction produces that range from any point predictor at all — a linear model, a gradient-boosted ensemble, a nearest-neighbour rule — without assuming the errors are Gaussian, symmetric, or homoscedastic, and without retraining anything. ## The recipe 1. **Split.** Partition the labelled data into a *proper training set* and a *calibration set* of size `n`. The calibration rows must never touch model fitting. 2. **Fit.** Train the model `f` on the proper training set alone. 3. **Score.** For each calibration row compute a **nonconformity score** — a number that is large when the model was surprised. For regression the default is the absolute residual `s_i = |y_i - f(x_i)|`. 4. **Quantile.** Sort the scores ascending and take `q = s_(k)` where `k = ceil((n+1)(1-alpha))`. 5. **Predict.** For a new input `x`, output `[f(x) - q, f(x) + q]`. With `n = 500` and `alpha = 0.10`, `k = ceil(501 * 0.9) = ceil(450.9) = 451`, so `q` is the 451st smallest of the 500 absolute residuals — not the 450th. That `+1` is the finite-sample correction, and it matters most when `n` is small. ## Why it works The argument is a rank argument, not a distributional one. Imagine the test row's own nonconformity score `s_new` dropped into the pile of `n` calibration scores. If all `n+1` scores are **exchangeable** — their joint distribution is unchanged by any reordering — then `s_new` is equally likely to occupy any of the `n+1` positions. The interval fails exactly when `s_new` exceeds `q`, and the probability of landing above the `k`-th of `n+1` positions is at most `alpha` by construction. Nothing in that reasoning needed a normality assumption, an asymptotic argument, or a correctly specified model. Exchangeability is weaker than independent and identically distributed: iid implies exchangeable, but not the reverse. It is the only assumption in play, and it is the one that breaks in practice. ## What the guarantee does and does not say - **Distribution-free and finite-sample.** Valid at `n = 50`, not just asymptotically. - **Model-agnostic.** A badly fitted model still yields *valid* intervals — they simply become wide. Interval width is therefore a clean read on model quality, while coverage is not. - **Marginal.** The probability is averaged over the random draw of the calibration set *and* the random test point. It is not a per-input statement, and it is not automatically true within a subgroup. - **Slightly conservative.** With continuous scores and no ties, coverage sits between `1-alpha` and `1-alpha + 1/(n+1)`. - **Random in realisation.** Once you fix one calibration set, the coverage you actually experience fluctuates around `1-alpha`; the spread shrinks as `n` grows, which is the practical reason to prefer several hundred calibration rows over a few dozen. The mechanics only *require* `n >= 9` at `alpha = 0.10`, since otherwise `ceil((n+1)(1-alpha))` exceeds `n` and no finite quantile exists — that degenerate case corresponds to an infinite-width interval. ## The score is the design decision Everything interesting lives in the choice of nonconformity score. The absolute residual gives one global `q`, so every interval has the identical width `2q` — cheap, valid, and blind to the fact that some inputs are far harder to predict than others. A **normalized** score `|y - f(x)| / sigma(x)`, where `sigma(x)` is a second model predicting the typical error magnitude at `x`, restores width that varies with the input. **Conformalized quantile regression** goes further by conformalizing a predicted band rather than a point. Swapping the score never breaks validity — only the usefulness of the interval changes. ## Variants worth naming **Full (transductive) conformal** refits the model for every candidate value of the new label, which recovers the data spent on calibration but multiplies compute by the size of the candidate grid. **Jackknife+** and **CV+** reuse all the data through leave-one-out or fold-wise residuals, buying back sample efficiency at the cost of a weaker worst-case guarantee of `1 - 2*alpha`. Split conformal stays the default in practice because it costs exactly one extra model evaluation per calibration row. ## Common way to get it wrong Computing the residual quantile on the training data. In-sample residuals are optimistically small for any model with capacity, so the interval comes out too narrow and the guarantee evaporates — the rank argument requires that the calibration rows were never seen during fitting.
- Why the rank ceil((n+1)(1-alpha)) instead of the plain empirical 90th percentile of the residuals?Because the guarantee comes from ranking the test row's score among all `n+1` exchangeable scores, not among the `n` calibration scores alone. The `+1` is the finite-sample correction that keeps coverage at or above `1-alpha`. With 500 rows it moves the rank from 450 to 451 and changes almost nothing; with 20 rows it is the difference between valid and slightly under-covering.
- Does the coverage guarantee survive a badly fitted model?Yes. Coverage depends only on exchangeability and the ranking argument, not on model quality. A weak model produces large calibration residuals, so `q` is large and the intervals come out wide but still cover 90% of the time. That is why width, not coverage, is the diagnostic for how good the model is.
- How much calibration data do you actually need?At `alpha = 0.10` the rank only exists once `n >= 9`, but that is a floor, not a target. Coverage is guaranteed on average over calibration draws; with one fixed small split the realised coverage swings widely. A few hundred rows keeps the quantile stable, and you need more if you plan to calibrate per subgroup.
saying these in an interview costs you the question
- Computes the residual quantile on the training data
- Says the interval is 90% likely for each individual input
- Claims conformal requires normally distributed residuals
- Believes conformal makes the point prediction more accurate
- Uses the plain n(1-alpha) empirical quantile with no correction