How do you put a percentile band on a classifier's accuracy using bootstrap replicates?
answer
- resample rows you scored, not the model
- recompute the metric each replicate
- sort, then take 2.5% and 97.5%
- width comes from n, not from B
- training variability is not in there
basics
~20 sResample the evaluation rows with replacement, recompute accuracy on each resample, and take the 2.5th and 97.5th percentiles of those values as a 95% band. It captures evaluation-set sampling noise, with the trained model held fixed.
solid answer
~50 sHold the trained model fixed and resample the *evaluation* rows. Draw 500 documents with replacement from the 500 you scored, recompute accuracy on that draw, repeat 2,000 times, sort the 2,000 accuracies and read off the 2.5th and 97.5th percentiles - that pair is a 95% percentile band. On 500 documents at 0.85 accuracy the band comes out roughly three points either side, which is usually the single most useful thing you can add to a headline metric. Two things it does not do: it does not capture how much the model would move if you retrained it on different training data, because the model never changes across replicates; and more replicates stabilise the endpoints without narrowing the band - width is set by how many evaluation rows you have, not by how many times you resample them.
go deeper
Know that a lone metric hides sampling noise, and that resampling the scored rows with replacement and reading off the 2.5th and 97.5th percentiles gives a defensible interval around it.
Explain that the model stays frozen while the evaluation rows are resampled, that the metric is recomputed inside each replicate, and that band width is set by evaluation-set size rather than by the number of replicates.
Diagnose when the band is misleading: rare positives producing degenerate replicates, clustered rows needing group-level resampling, or a band drawn around a metric measured on the rows that chose the model.
Set the standard for how uncertainty is reported and acted on - what interval accompanies every headline metric, how large an evaluation set a launch decision requires, and when a band that overlaps the incumbent means the change simply is not decidable yet.
## The recipe You have a trained classifier and a set of evaluation rows it has never seen - say 500 documents with sentiment labels - and a point accuracy of 0.85. A single number invites over-reading, so put a band on it: 1. Draw 500 rows **with replacement** from the 500 evaluation rows. 2. Recompute accuracy using the predictions already made for those rows. 3. Repeat B times (2,000 is a comfortable default). 4. Sort the B accuracies. The 2.5th and 97.5th percentiles are the endpoints of a 95% **percentile band**. The model is fitted once and never refitted; you are resampling *scored rows*, which makes the whole thing cheap - no retraining, just re-averaging predictions that already exist. ## What the band means It answers: *if I had drawn a different 500 documents from the same population and scored this same fixed model on them, how much would the reported accuracy have wobbled?* That is evaluation-set sampling noise. On 500 rows at 0.85, the standard error of accuracy is about `sqrt(0.85 * 0.15 / 500) = 0.016`, so the band lands near [0.82, 0.88] - about three points either side. If your product decision turns on a two-point difference, the band tells you the evaluation set is too small to support it. Interpretation discipline matters. The band is a frequentist coverage statement about the procedure: bands built this way contain the fixed underlying accuracy about 95% of the time. It is not "a 95% probability that the true accuracy is inside this particular interval". ## What the band does NOT include **Training variability.** The model is frozen. If you retrained on a different sample, the model itself would change - possibly by more than the evaluation noise. To capture that you have to resample the *training* rows, refit the model on each replicate, and score it on the rows that replicate left out. That version measures training instability plus evaluation noise together, but it costs B full refits instead of B cheap averages, and it estimates the error of the *procedure* rather than of one specific trained model. Decide which question you are answering before you pick. **Selection effects.** If the same rows chose the model - picked the threshold, the features, the hyperparameters - then the point accuracy is already optimistic and a band around it is a band around an inflated number. Bands do not repair a contaminated estimate; they only quantify noise around whatever you measured. **Distribution shift.** Every replicate is drawn from your evaluation set, so the band describes noise under the assumption that production data looks like that set. It says nothing about next quarter's documents. ## More replicates do not narrow the band This is the most common misunderstanding. B controls **Monte Carlo** error - how precisely you have located the percentiles of the resampling distribution. Going from 500 to 20,000 replicates makes the endpoints reproducible to more decimal places; it does not change the underlying width, which is governed by the evaluation-set size `n` and the metric's variance. Want a tighter band? Score more rows. A useful rule of thumb: a few hundred replicates suffice for a central estimate, one to two thousand for stable 2.5%/97.5% endpoints, more if you are reading extreme tails. ## Practical wrinkles **Discreteness.** With few evaluation rows, accuracy can only take a handful of values, and the percentile band inherits that lumpiness - it may look artificially narrow or step-shaped. Report the sample size next to the band so the reader can judge. **Rare positives.** If the positive class is scarce, some resamples will contain very few positives, or none. Metrics like recall or AUC become unstable or undefined in those replicates. Either stratify the resample so each replicate keeps the class counts of the original evaluation set, or report explicitly how many replicates were degenerate. Silently dropping them biases the band. **Grouped rows.** If evaluation rows are not independent - several documents per author, repeated measurements per customer - resample the *groups*, not the rows. Row-level resampling on clustered data pretends you have more independent evidence than you do and yields a band that is too narrow. **Any metric works.** The same machinery applies to F1, precision at a fixed threshold, AUC, mean absolute error - anything computable from a set of scored rows. Recompute the metric inside each replicate rather than averaging per-row values, because most useful metrics are not simple row averages. ## How to report it Write it as "accuracy 0.85, 95% percentile band [0.82, 0.88] from 2,000 bootstrap resamples of 500 held-out documents". That single line tells a reader the estimate, the uncertainty, the source of the uncertainty and the sample size behind it - and it makes clear the band is about evaluation noise, not about model retraining.
- Does a wider band mean the model is worse?No. Width is driven by how many rows you scored and how variable the metric is, not by model quality. A strong model measured on 100 rows will show a wider band than a mediocre one measured on 50,000. Compare bands only between numbers computed on evaluation sets of comparable size.
- How would you also capture how much the model itself moves?Resample the training rows instead, refit the model on each replicate, and score it on the rows that replicate left out. The resulting spread mixes training instability with evaluation noise and describes the modelling procedure rather than one frozen model. It costs a full refit per replicate, so it is orders of magnitude more expensive.
- The positive class is only 3% of the evaluation rows - what changes?Some replicates will contain almost no positives, making recall noisy and threshold-dependent metrics unstable or undefined. Stratify the resample so each replicate preserves the original class counts, or report how many replicates were degenerate. Never quietly discard them - that truncates the low tail and makes the band look tighter than it is.
saying these in an interview costs you the question
- Reports a single accuracy number with no uncertainty
- Thinks raising the replicate count narrows the band
- Bootstraps rows the model was trained or tuned on
- Says there is a 95% chance the truth lies inside
- Refits the model on every replicate without needing to
- Resamples rows when data is clustered by author or customer