skip to content

How does a BCa bootstrap interval differ from a percentile interval, and when is it worth it?

level: seniorimportance: nice to knowfreq 32%

answer

  1. same replicates, different quantile levels
  2. two correction constants, both often zero
  3. bias correction from the replicate fraction
  4. acceleration usually from the jackknife
  5. second-order accuracy on skewed statistics

basics

~20 s

A percentile interval reads the 2.5th and 97.5th percentiles of the bootstrap distribution as they are. BCa shifts which percentiles it reads, correcting for an off-centre distribution and a standard error that moves with the parameter.

solid answer

~50 s

The percentile interval reads the alpha/2 and 1 - alpha/2 quantiles straight off the bootstrap distribution. That is exactly right only when some monotone transformation would make the estimator unbiased, symmetric and constant in variance, which is often close enough and sometimes not. BCa keeps the same replicates but moves the quantile levels it reads, using two constants: a bias-correction `z0`, estimated from the fraction of replicates falling below the original point estimate, and an acceleration `a`, usually estimated from jackknife leave-one-out values, capturing how fast the standard error changes with the parameter. When both come out near zero, BCa collapses back to the percentile interval. The payoff is second-order accuracy and better coverage on skewed statistics such as a correlation or a variance ratio; the cost is n extra evaluations and instability when n is small or the statistic is non-smooth.

go deeper

for a junior

Know that more than one recipe turns a bootstrap distribution into an interval, and that the percentile recipe, cutting 2.5% off each tail, is the simplest and the one to describe by default.

for a middle

Be able to say what each BCa constant is for: one shifts for a bootstrap distribution that is not centred on the estimate, the other for a standard error that changes as the parameter changes.

for a senior

Demonstrate the judgment call. Look at the skew of the bootstrap distribution, decide whether the extra jackknife pass earns its keep, and know that BCa is least reliable exactly where the statistic is non-smooth or the sample is small.

for a principal

Own which interval flavour is the house default and how it is reported, so results from different teams are comparable and nobody quietly switches recipes until an interval clears a threshold.

## Turning replicates into an interval A bootstrap gives you B replicate values of a statistic. Converting that cloud into a confidence interval is a separate decision, and several recipes exist. The two worth knowing well are the percentile interval and the bias-corrected and accelerated interval, usually written BCa. ## The percentile interval Sort the replicates and take the alpha/2 and 1 - alpha/2 quantiles: for 95%, the 2.5th and 97.5th percentiles. Its great virtue is that it is **transformation-respecting**. If you compute a percentile interval for a parameter and then apply any monotone transformation to both endpoints, you get exactly the percentile interval you would have obtained by bootstrapping the transformed statistic directly. That is why a percentile interval for a correlation can never escape the range from -1 to 1: no replicate can. Its weakness is subtler. The percentile interval is exactly correct under a hidden condition: that there exists some monotone transformation under which the estimator is unbiased, normally distributed, and has constant variance. You never need to find that transformation, which is the point, but you do need it to exist. For a mean of reasonably behaved data it effectively does. For a correlation near the boundary, a variance ratio, or a small-sample skewed statistic, it does not, and the interval is misplaced: it can be shifted relative to where it should be, and its real coverage can be several points below the nominal level, typically failing asymmetrically so one tail errs much more than the other. ## What BCa adds BCa reads the same replicates but at **adjusted quantile levels**. Two constants drive the adjustment. The bias-correction constant, z0, is the inverse standard normal CDF applied to the proportion of bootstrap replicates that fall strictly below the original point estimate. If the bootstrap distribution is centred on the estimate, that proportion is 0.5 and z0 is 0. A proportion far from a half means the bootstrap distribution sits off to one side of the estimate, and z0 records how far in normal-score units. The acceleration constant, a, measures how fast the standard error of the estimator changes as the parameter changes, which is a skewness effect. It is customarily estimated from jackknife leave-one-out values: with theta_(i) the statistic computed with observation i removed and theta_(.) their mean, a is the sum over i of (theta_(.) - theta_(i)) cubed, divided by six times the sum of (theta_(.) - theta_(i)) squared raised to the power three halves. When the leave-one-out values are symmetric about their mean, the cubes cancel and a is zero. The adjusted lower level is Phi( z0 + (z0 + z_lo) / (1 - a * (z0 + z_lo)) ), where Phi is the standard normal CDF and z_lo is the inverse normal CDF at alpha/2. The upper level uses the inverse normal CDF at 1 - alpha/2 in the same expression. You then read the bootstrap distribution at those two adjusted levels rather than at 2.5% and 97.5%. Set z0 and a to zero and the expression collapses to the original levels, so BCa contains the percentile interval as a special case. ## Accuracy, in one sentence each The percentile interval is **first-order accurate**: its coverage error shrinks like 1/sqrt(n). BCa is **second-order accurate**: its coverage error shrinks like 1/n. On a well-behaved statistic with a decent sample, the two intervals nearly coincide and the distinction is academic. On a skewed statistic, or a small sample, the difference in real coverage is visible. ## The basic, or reverse-percentile, interval A third recipe reflects the bootstrap quantiles about the point estimate: the endpoints are two times the estimate minus the upper bootstrap quantile, and two times the estimate minus the lower one. The logic is that the bootstrap distribution of (replicate minus estimate) approximates the distribution of (estimate minus truth), so you subtract rather than read directly. It corrects crudely for bootstrap bias but, unlike the percentile and BCa intervals, it is not transformation-respecting and can produce endpoints outside the parameter's legal range, for instance a correlation bound past 1. ## When to use which Plot the replicates first. If the bootstrap distribution is roughly symmetric and centred on the point estimate, z0 and a will both be near zero, BCa will reproduce the percentile interval, and the extra jackknife pass buys nothing. If the distribution is visibly skewed, or the statistic is bounded and the estimate sits near a boundary, BCa is the better report. The cases where BCa should be treated warily are small n and non-smooth statistics, precisely because the acceleration constant is estimated from a jackknife, which is itself unreliable in exactly those cases. A wildly estimated acceleration can push an adjusted level near 0 or 1 and place an endpoint somewhere indefensible. When that happens, the honest report is a plain percentile interval labelled as such, together with a note that the sample is too small for the refinement to be trusted. Finally, fix the recipe up front. Computing several interval flavours and reporting the one that clears a decision threshold is a selection effect dressed up as a method choice.

  • How is the bias-correction constant z0 in a BCa interval actually computed?
    Count the bootstrap replicates falling strictly below the original point estimate, divide by B, and pass that proportion through the inverse standard normal CDF. If the bootstrap distribution is centred on the estimate the proportion is a half and z0 is zero, so no shift happens. A proportion far from a half signals a distribution sitting off to one side, and z0 records that offset in normal-score units.
  • When would you not bother with BCa?
    When the bootstrap distribution is visibly symmetric and centred on the estimate, both constants land near zero and BCa simply reproduces the percentile interval at extra cost. Skip it too when the sample is small or the statistic is non-smooth, because the jackknife-based acceleration term is unstable exactly there and can push an endpoint somewhere indefensible. A plain percentile interval, honestly labelled, is then the safer report.
  • What is the basic, or reverse-percentile, interval and how does it relate to these two?
    It reflects the bootstrap quantiles about the estimate: the endpoints are two times the point estimate minus the upper bootstrap quantile, and two times the estimate minus the lower one. It corrects crudely for bootstrap bias, but unlike the percentile and BCa intervals it is not transformation-respecting, so it can produce endpoints outside a parameter's legal range, such as a correlation bound past 1.

saying these in an interview costs you the question

  • Thinks BCa uses more replicates rather than different quantile levels
  • Applies BCa blindly to tiny samples or non-smooth statistics
  • Says a percentile interval requires the statistic to be normal
  • Cannot say what the acceleration constant measures
  • Treats percentile intervals as adequate regardless of skew
  • Computes several interval flavours and reports the most favourable

context