skip to content

How does split conformal prediction turn a trained regression model into 90% prediction intervals?

level: middleimportance: must knowfreq 40%

answer

  1. hold out rows the model never saw
  2. score how surprised the model was
  3. sort the scores, take a high rank
  4. rank is ceil((n+1)(1-alpha)), not n(1-alpha)
  5. same half-width added to every prediction

basics

~20 s

Split conformal holds out a calibration set the model never trained on, scores every row by its absolute residual, and takes the residual sitting at the 90% rank. Each new prediction then gets that one value added and subtracted.

solid answer

~50 s

Split the labelled data into a proper training set and a calibration set. Fit the model on the training portion only, then compute a nonconformity score on each of the `n` calibration rows — for a plain regressor the natural score is the absolute residual `|y - f(x)|`. Sort those scores and take the `k`-th smallest, where `k = ceil((n+1)(1-alpha))`; call it `q`. For any new input the interval is `f(x) +/- q`. Provided the calibration rows and the new row are exchangeable, that interval contains the true value with probability at least `1 - alpha` — in finite samples, for any model, with no assumption about the shape of the error distribution. The costs are real: the guarantee is marginal rather than per-input, the width is the same for every input, and you spent labelled data on calibration.

code

python · 15 lines
python
import math, random
random.seed(7)

# calibration rows the model never trained on: (true value, model's point prediction)
cal = [(random.gauss(20, 4), 19.0) for _ in range(500)]

scores = sorted(abs(y - yhat) for y, yhat in cal)   # nonconformity score
alpha = 0.10
n = len(scores)
k = math.ceil((n + 1) * (1 - alpha))                # rank to take
q = scores[k - 1]

point = 19.0
print("rank k =", k, "| q =", round(q, 2))          # rank k = 451 | q = 6.89
print("interval =", (round(point - q, 1), round(point + q, 1)))  # (12.1, 25.9)

go deeper

for a junior

Recall the shape of it: train on one part of the data, measure errors on a held-out part, and use those errors to put a band around future predictions. Know that 90% refers to how often the truth lands inside.

for a middle

Be ready to write the recipe out: proper training split, nonconformity score, the ceil((n+1)(1-alpha)) rank, and the interval f(x) plus or minus q. Explain why the calibration rows must be unseen by the model.

for a senior

Show that you know what the guarantee buys and what it does not: finite-sample and distribution-free, but marginal and constant-width with an absolute-residual score. Talk about how large a calibration split you would carve out and why.

for a principal

Own the framing that conformal is a wrapper, not a modelling improvement. Argue when a distribution-free interval is worth spending labelled data on, versus when a cheaper parametric band or an unquantified point estimate is enough for the decision at hand.

## The problem it solves A regression model returns one number. A decision-maker usually needs a range plus a promise about how often the truth falls inside it. Split conformal prediction produces that range from any point predictor at all — a linear model, a gradient-boosted ensemble, a nearest-neighbour rule — without assuming the errors are Gaussian, symmetric, or homoscedastic, and without retraining anything. ## The recipe 1. **Split.** Partition the labelled data into a *proper training set* and a *calibration set* of size `n`. The calibration rows must never touch model fitting. 2. **Fit.** Train the model `f` on the proper training set alone. 3. **Score.** For each calibration row compute a **nonconformity score** — a number that is large when the model was surprised. For regression the default is the absolute residual `s_i = |y_i - f(x_i)|`. 4. **Quantile.** Sort the scores ascending and take `q = s_(k)` where `k = ceil((n+1)(1-alpha))`. 5. **Predict.** For a new input `x`, output `[f(x) - q, f(x) + q]`. With `n = 500` and `alpha = 0.10`, `k = ceil(501 * 0.9) = ceil(450.9) = 451`, so `q` is the 451st smallest of the 500 absolute residuals — not the 450th. That `+1` is the finite-sample correction, and it matters most when `n` is small. ## Why it works The argument is a rank argument, not a distributional one. Imagine the test row's own nonconformity score `s_new` dropped into the pile of `n` calibration scores. If all `n+1` scores are **exchangeable** — their joint distribution is unchanged by any reordering — then `s_new` is equally likely to occupy any of the `n+1` positions. The interval fails exactly when `s_new` exceeds `q`, and the probability of landing above the `k`-th of `n+1` positions is at most `alpha` by construction. Nothing in that reasoning needed a normality assumption, an asymptotic argument, or a correctly specified model. Exchangeability is weaker than independent and identically distributed: iid implies exchangeable, but not the reverse. It is the only assumption in play, and it is the one that breaks in practice. ## What the guarantee does and does not say - **Distribution-free and finite-sample.** Valid at `n = 50`, not just asymptotically. - **Model-agnostic.** A badly fitted model still yields *valid* intervals — they simply become wide. Interval width is therefore a clean read on model quality, while coverage is not. - **Marginal.** The probability is averaged over the random draw of the calibration set *and* the random test point. It is not a per-input statement, and it is not automatically true within a subgroup. - **Slightly conservative.** With continuous scores and no ties, coverage sits between `1-alpha` and `1-alpha + 1/(n+1)`. - **Random in realisation.** Once you fix one calibration set, the coverage you actually experience fluctuates around `1-alpha`; the spread shrinks as `n` grows, which is the practical reason to prefer several hundred calibration rows over a few dozen. The mechanics only *require* `n >= 9` at `alpha = 0.10`, since otherwise `ceil((n+1)(1-alpha))` exceeds `n` and no finite quantile exists — that degenerate case corresponds to an infinite-width interval. ## The score is the design decision Everything interesting lives in the choice of nonconformity score. The absolute residual gives one global `q`, so every interval has the identical width `2q` — cheap, valid, and blind to the fact that some inputs are far harder to predict than others. A **normalized** score `|y - f(x)| / sigma(x)`, where `sigma(x)` is a second model predicting the typical error magnitude at `x`, restores width that varies with the input. **Conformalized quantile regression** goes further by conformalizing a predicted band rather than a point. Swapping the score never breaks validity — only the usefulness of the interval changes. ## Variants worth naming **Full (transductive) conformal** refits the model for every candidate value of the new label, which recovers the data spent on calibration but multiplies compute by the size of the candidate grid. **Jackknife+** and **CV+** reuse all the data through leave-one-out or fold-wise residuals, buying back sample efficiency at the cost of a weaker worst-case guarantee of `1 - 2*alpha`. Split conformal stays the default in practice because it costs exactly one extra model evaluation per calibration row. ## Common way to get it wrong Computing the residual quantile on the training data. In-sample residuals are optimistically small for any model with capacity, so the interval comes out too narrow and the guarantee evaporates — the rank argument requires that the calibration rows were never seen during fitting.

  • Why the rank ceil((n+1)(1-alpha)) instead of the plain empirical 90th percentile of the residuals?
    Because the guarantee comes from ranking the test row's score among all `n+1` exchangeable scores, not among the `n` calibration scores alone. The `+1` is the finite-sample correction that keeps coverage at or above `1-alpha`. With 500 rows it moves the rank from 450 to 451 and changes almost nothing; with 20 rows it is the difference between valid and slightly under-covering.
  • Does the coverage guarantee survive a badly fitted model?
    Yes. Coverage depends only on exchangeability and the ranking argument, not on model quality. A weak model produces large calibration residuals, so `q` is large and the intervals come out wide but still cover 90% of the time. That is why width, not coverage, is the diagnostic for how good the model is.
  • How much calibration data do you actually need?
    At `alpha = 0.10` the rank only exists once `n >= 9`, but that is a floor, not a target. Coverage is guaranteed on average over calibration draws; with one fixed small split the realised coverage swings widely. A few hundred rows keeps the quantile stable, and you need more if you plan to calibrate per subgroup.

saying these in an interview costs you the question

  • Computes the residual quantile on the training data
  • Says the interval is 90% likely for each individual input
  • Claims conformal requires normally distributed residuals
  • Believes conformal makes the point prediction more accurate
  • Uses the plain n(1-alpha) empirical quantile with no correction

context