Why is bootstrap left-out error pessimistic, and what does the .632 estimator correct?
answer
- two biases, opposite directions
- each fit trains on only 63% distinct rows
- blend training error with left-out error
- 0.632 is the in-bag probability
- zero training error breaks the blend
basics
~20 sEach bootstrap replicate trains on only about 63% distinct rows, so its models are weaker and the error on left-out rows comes out too high. The .632 estimator blends that pessimistic error with the optimistic training error, 0.632 to 0.368.
solid answer
~50 sA bootstrap replicate contains about 63.2% of the distinct rows, so every fit is effectively trained on a smaller sample than the model you will ship. Learning curves slope downward, so those models predict worse and the error measured on the left-out rows is biased **upward** - pessimistic. The apparent error on the training rows is biased the other way, downward. Efron's .632 estimator splits the difference: `Err_632 = 0.368 * apparent_error + 0.632 * leave_one_out_bootstrap_error`, with 0.632 chosen because it is the probability a given row is in-bag. The catch is that any model that memorises its training data has an apparent error of zero - a 1-nearest-neighbour classifier is the canonical case - so the blend is dragged toward zero and becomes wildly optimistic. The .632+ variant fixes that by making the weight adaptive to how much the model is actually overfitting.
go deeper
Know that error measured on the rows a resample left out tends to come out slightly too high, and that a correction exists because each fit only trains on about two-thirds of the distinct rows.
Explain both biases and their directions, and state the blend: 0.368 times the training error plus 0.632 times the left-out error, with 0.632 being the probability a row is in the resample.
Show you know where it fails. Walk through a 1-nearest-neighbour classifier whose training error is zero, why that drags the blend toward optimism, and how an adaptive weight based on the relative overfitting rate repairs it.
Own the call on whether a corrected bootstrap belongs in your evaluation standard at all - weigh its bias-variance advantage on small samples against the cost of reporting a blended number that stakeholders and reviewers cannot interpret or reproduce.
## Two biased estimates pointing in opposite directions When you estimate prediction error by resampling, you have two numbers available and both are wrong in a known direction. **The apparent error** (also called resubstitution or training error) is the error the model makes on the very rows it was fitted to. It is **optimistically biased**: the model has already adapted to those rows. For a flexible enough model it can be zero while true predictive error is dreadful. **The leave-one-out bootstrap error** is the error on the rows each replicate left out, averaged over replicates (for each row, average its error over only those replicates that did not draw it). It is **pessimistically biased**, and the reason is subtle but concrete: a replicate contains only about 63.2% distinct rows, so each model is trained on roughly two-thirds of an effective sample. Learning curves slope downward - more training data means lower error - so a model fitted on 63% of the rows predicts worse than the model you will finally ship on 100% of them. The estimate therefore overstates the error of the thing you actually care about. The effect is largest exactly where it hurts most: small samples and high-variance learners, where the learning curve is still steep. ## Efron's .632 estimator If one estimate is too low and the other too high, blend them. Efron's proposal weights them by the in-bag and out-of-bag probabilities: ``` Err_632 = 0.368 * apparent_error + 0.632 * loo_bootstrap_error ``` The 0.632 is not a tuning knob picked by taste - it is `1 - 1/e`, the probability that any given row appears in a replicate. The reasoning is that the left-out error effectively evaluates a model trained on 63.2% of the data, so weighting it at 0.632 and giving the remaining mass to the (optimistic) apparent error approximately cancels the two biases. Empirically it works well for models with moderate flexibility and small samples, which is precisely where a single hold-out is unaffordable. ## Why it breaks, and the .632+ repair The blend assumes the apparent error carries real information about optimism. For a **1-nearest-neighbour classifier** it carries none: every training row is its own nearest neighbour, so apparent error is identically zero no matter how noisy the problem is. Plug that in and ``` Err_632 = 0.368 * 0 + 0.632 * loo_bootstrap_error = 0.632 * loo_bootstrap_error ``` On a hopeless problem where the left-out error is 50%, the .632 estimator reports 31.6% - a badly optimistic answer produced by a model that has learned nothing. Any strongly overfitting learner triggers the same failure to a lesser degree. The **.632+ estimator** (Efron and Tibshirani) makes the weight respond to how much overfitting is actually happening. It introduces two extra quantities: - The **no-information error rate** `gamma`: the error the fitted model would make if the inputs and the labels were unrelated. It is estimated by scoring every prediction against every label - all n-squared pairings - or equivalently by permuting the labels. For a balanced two-class problem it is around 0.5. - The **relative overfitting rate** `R = (loo_error - apparent_error) / (gamma - apparent_error)`, which lands near 0 when the model barely overfits and near 1 when it overfits maximally. The weight becomes `w = 0.632 / (1 - 0.368 * R)`, and the estimator is `(1 - w) * apparent_error + w * loo_error`. When `R = 0` this reduces to the plain .632 rule; when `R = 1` the weight goes to 1 and the estimate falls back entirely on the left-out error - which is exactly what you want for 1-nearest-neighbour, where the estimator then reports the honest 50% instead of 31.6%. ## How this compares with k-fold k-fold makes the same trade in a different place. Ten-fold trains on 90% of the rows, so its pessimistic bias is small; five-fold trains on 80% and is a little more pessimistic. The bootstrap trains on ~63% distinct rows, so its raw bias is the largest of the three - but it averages over hundreds of fits instead of five or ten, which makes the estimate noticeably more stable from run to run. The corrected .632 and .632+ estimators are attempts to keep the low variance while removing the bias, which is why they show up mostly on genuinely small samples where a stable number matters more than a simple one. ## When to reach for it, and when not to Use a corrected bootstrap estimate when the sample is small, the learning curve is steep, and you need a low-variance error number. Do not use it when you cannot compute a meaningful apparent error, when the rows are clustered or time-ordered (resampling rows independently destroys the dependence structure and produces a badly optimistic answer either way), or when the audience will not understand a blended number - a plain cross-validated error that everyone can reason about often beats a cleverer estimate that nobody trusts. And whatever you report, state which estimator produced it: ".632+ bootstrap error, 500 replicates" is interpretable, "error 0.19" is not.
- How does the bootstrap's bias compare with 10-fold cross-validation?Ten-fold trains each model on 90% of the rows, so its pessimistic bias is mild; a bootstrap replicate carries only about 63% distinct rows, so its raw bias is larger. In exchange the bootstrap averages over hundreds of fits rather than ten, giving a lower-variance estimate. The .632 family exists to keep that variance advantage while removing the bias.
- How do you estimate the no-information rate used by .632+?It is the error the fitted model would make if inputs and labels were independent. Estimate it by scoring every prediction against every label - all n-squared pairings - and averaging, or equivalently by permuting the labels and re-scoring. For a balanced two-class problem with a symmetric loss it sits near 0.5.
- Would you actually report a .632+ number to a stakeholder?Usually not as the headline. It is a blended quantity that is hard to explain and easy to distrust. I would use it internally when the sample is too small for a stable cross-validated number, and report a plainly-defined error alongside it. An estimator nobody in the room can interpret does not help a decision, however well calibrated it is.
saying these in an interview costs you the question
- Thinks the left-out bootstrap error is unbiased
- Believes .632 refers to a confidence level
- Says the 0.632 weight was chosen arbitrarily
- Applies plain .632 to a zero-training-error model
- Gets the bias direction backwards - calls it optimistic
- Confuses optimism correction with regularisation