skip to content

What is the jackknife, and why does it break down for the sample median?

level: middleimportance: nice to knowfreq 26%

answer

  1. leave one observation out at a time
  2. n recomputations, no randomness, no seed
  3. pseudo-values give bias and variance
  4. a linear approximation to the bootstrap
  5. non-smooth statistics barely move

basics

~20 s

The jackknife recomputes a statistic n times, leaving out one observation each time, and uses the spread of those values to estimate bias and standard error. It fails for non-smooth statistics such as the median.

solid answer

~50 s

The jackknife is leave-one-out resampling. With `theta_hat` computed on all n observations, compute `theta_(i)` on the sample with observation i removed, for every i, and let `theta_(.)` be their mean. The bias estimate is `(n - 1) * (theta_(.) - theta_hat)` and the variance estimate is `((n - 1) / n) * sum over i of (theta_(i) - theta_(.)) squared`. Equivalently, form pseudo-values `n * theta_hat - (n - 1) * theta_(i)` and treat them as n observations whose mean is the bias-corrected estimate. It predates the bootstrap, costs only n evaluations, is fully deterministic, and is essentially a linear approximation to the bootstrap, which is exactly why it fails on non-smooth statistics. Removing one point from an odd-sized sample shifts the median only to a neighbouring order statistic, so the leave-one-out values take very few distinct values and the variance estimate is inconsistent.

go deeper

for a junior

Recall the mechanic: drop one observation, recompute the statistic, repeat for all n, and look at how much the answer moved. There is no randomness and no seed involved.

for a middle

Explain the pseudo-value construction and the (n - 1) scaling in the bias and variance formulas, and be able to say why deleting a single point barely moves a median.

for a senior

Know when it is still the right tool despite its age, namely a deterministic standard error for smooth statistics and the acceleration term in a BCa interval, and reach for delete-d or the bootstrap when the statistic is non-smooth.

for a principal

Own the guidance your team follows on which resampling tool fits which statistic, so nobody ships a jackknife standard error for an order statistic and nobody re-derives the same choice every quarter.

## The mechanics The jackknife is the oldest of the resampling methods and the simplest to state. Given a sample of n observations and a statistic theta_hat computed on all of them: 1. For each i from 1 to n, delete observation i and recompute the statistic on the remaining n - 1 observations. Call the result theta_(i). 2. Let theta_(.) be the average of the n leave-one-out values. 3. The **bias estimate** is (n - 1) times (theta_(.) - theta_hat). Subtracting it gives the bias-corrected estimate, which can also be written as n times theta_hat minus (n - 1) times theta_(.). 4. The **variance estimate** is (n - 1)/n times the sum over i of (theta_(i) - theta_(.)) squared. There is no randomness anywhere: run it twice and you get identical numbers, with no seed and no B to argue about. The cost is exactly n recomputations. ## Pseudo-values An equivalent presentation defines the i-th **pseudo-value** as n times theta_hat minus (n - 1) times theta_(i). The average of the pseudo-values is precisely the bias-corrected estimate, and their sample variance divided by n is precisely the jackknife variance estimate. The appeal is conceptual: the pseudo-values behave, for smooth statistics, roughly like n independent observations of the quantity of interest, so ordinary formulas for a mean apply to them. For a sample mean the pseudo-values are exactly the original observations, which is a useful sanity check on the definitions. ## Why the (n - 1) scaling appears twice Both formulas carry an inflation factor, and for the same underlying reason. Leave-one-out samples are enormously similar to each other: any two of them share n - 2 observations. Their raw spread therefore badly understates how much the statistic would move across genuinely independent samples, and the (n - 1) factors scale that shrunken spread back up. In the bias formula, (n - 1) converts the small observed shift between theta_(.) and theta_hat into an estimate of the bias at sample size n, exploiting the fact that bias of a smooth estimator typically behaves like a constant over n. ## The relationship to the bootstrap The jackknife can be read as a **linear approximation** to the bootstrap. It probes the statistic's sensitivity to each single observation, which is a first-order, one-point-at-a-time picture of how the statistic responds to perturbing the empirical distribution. The bootstrap perturbs the whole distribution at once and captures higher-order behaviour. When the statistic is smooth, that linear picture is enough, and the two agree closely. When it is not, the approximation misses the point entirely. ## Why the median defeats it Consider an odd-sized sample, say n = 11, sorted. The median is the sixth value. Delete an observation from the upper half and the median of the remaining ten values becomes the average of the fifth and sixth. Delete one from the lower half and it becomes the average of the sixth and seventh. Delete the median itself and you get the average of the fifth and seventh, or a similar neighbouring pair. Across all eleven deletions the leave-one-out median takes only about three distinct values, all sitting within one order statistic of each other. The resulting spread is a function of the local spacing between two or three central order statistics, not of the sampling variability of the median, and it does not converge to the right answer as n grows. The jackknife variance estimate for the sample median is **inconsistent**: more data does not fix it. The technical statement of the problem is that the median is not a smooth functional of the distribution; it responds to a small perturbation in a way that a one-point-at-a-time probe cannot see. The same warning applies to other statistics defined by ranks or thresholds, and to any estimate produced by a selection step. ## The delete-d jackknife The standard repair deletes a block of d observations at a time rather than a single one, averaging over many such subsets. Letting d grow with n rather than staying pinned at one restores consistency for statistics such as the median, because removing a block actually moves the statistic by an amount that reflects genuine sampling variability. The costs are that you now have to choose d, and that enumerating all subsets is infeasible, so you sample them and reintroduce the Monte Carlo noise that the plain jackknife avoided. ## Where the jackknife still earns its place It has two live roles. First, it is the customary way to estimate the acceleration constant in a BCa bootstrap interval, which is defined directly in terms of the leave-one-out values. Second, it gives a cheap, deterministic, seed-free standard error for smooth statistics such as means, ratios and regression coefficients, which is genuinely convenient when reproducibility matters and nobody wants to negotiate over the number of replicates. For anything non-smooth, reach for the bootstrap or a delete-d variant instead.

  • Where does the jackknife still earn its place now that the bootstrap exists?
    Two places. It is the customary way to estimate the acceleration constant in a BCa bootstrap interval, which is defined in terms of leave-one-out values. And it is deterministic and cheap, exactly n evaluations with no seed and no Monte Carlo noise, so it produces a reproducible standard error for smooth statistics such as a mean, a ratio or a regression coefficient without anyone arguing about how many replicates to draw.
  • Why does the jackknife variance formula carry an (n - 1) inflation factor?
    Leave-one-out estimates are far more alike than independent samples would be, since any two of them share n - 2 observations. Their raw spread therefore badly understates the true sampling variability of the statistic. The (n - 1)/n factor, applied to the sum of squared deviations rather than to their average, scales that shrunken spread back up to a consistent variance estimate for smooth statistics.
  • What is the delete-d jackknife?
    Instead of removing one observation it removes a block of d at a time, averaging over many such subsets. Letting d grow with n rather than staying at one restores consistency for statistics like the median, because deleting a block actually moves the statistic by a meaningful amount. The costs are choosing d and sampling subsets rather than enumerating all n cases, which reintroduces simulation noise.

saying these in an interview costs you the question

  • Uses jackknife standard errors for medians and other non-smooth statistics
  • Thinks the jackknife resamples with replacement
  • Describes the jackknife as a model-validation split
  • Omits the (n - 1) scaling from the bias or variance formula
  • Assumes the jackknife always agrees with the bootstrap

context