skip to content

With 0 successes in 20 trials, why does the Wald interval for a proportion collapse to zero width?

level: seniorimportance: should knowfreq 46%

answer

  1. look at what goes into the standard error
  2. the estimate sits inside its own error term
  3. at a boundary that term vanishes
  4. the score interval uses the hypothesised value
  5. three over n for zero events

basics

~20 s

The Wald formula plugs the observed proportion into its own standard error. With zero successes that term becomes 0, so the interval degenerates to [0, 0] - certainty from 20 trials. The Wilson score interval does not.

solid answer

~50 s

The Wald interval is `p_hat ± z × sqrt(p_hat (1 - p_hat) / n)`. At `p_hat = 0` the square root is exactly 0, so the interval is [0, 0] regardless of n — a claim of perfect certainty from 20 observations. The Wilson score interval is derived by inverting the score test, which uses the *hypothesised* proportion in the standard error rather than the observed one, so it never collapses and never leaves [0, 1]. For 0 successes in 20 trials at 95% it gives roughly [0, 0.161], close to the rule-of-three quick bound `3 / 20 = 0.15`. This is not an edge case: Wald coverage runs below its nominal level whenever n is small or the proportion is near a boundary, so Wilson or the Agresti-Coull adjustment should be the default for proportions.

go deeper

for a junior

Recognise that an interval of [0, 0] from a handful of trials cannot be right, and know the name of the standard alternative for proportions.

for a middle

Explain the mechanism: the observed proportion sits inside its own standard error, so at 0 or 1 that term vanishes and the interval degenerates.

for a senior

Show the operational judgment — pick Wilson or Agresti-Coull as a default, quote the rule-of-three bound for a zero-event run, and say what you would actually report to a stakeholder.

for a principal

Own the standard: decide which interval method your organisation reports for rates, and be able to defend the conservatism trade-off when a regulator or a safety review asks.

## What Wald does The familiar interval for a proportion is ``` p_hat ± z × sqrt( p_hat (1 - p_hat) / n ) ``` where `p_hat = x / n` is the observed success fraction. It comes from treating the estimator as approximately normal and substituting the estimate into the variance formula. That substitution is the flaw. The width depends on the *observed* proportion, so a proportion that lands at a boundary reports no uncertainty at all. With `x = 0` and `n = 20`: `p_hat = 0`, `p_hat (1 - p_hat) = 0`, the square root is 0, and the interval is `[0, 0]`. Twenty trials without a success has become a statement that the true rate is exactly zero. The same collapse happens at `x = n`, giving `[1, 1]`. ## The problem is broader than the boundary Even away from 0 and 1, Wald misbehaves. Its actual coverage oscillates and is often meaningfully below the nominal level for small n, and stays below for proportions near the extremes even at fairly large n. It can also produce limits outside `[0, 1]`, which is plainly impossible for a proportion. The rule of thumb sometimes taught — that Wald is fine once `n p_hat` and `n (1 - p_hat)` both exceed about 10 — is a patch on a formula with a better replacement available. ## What Wilson does instead The Wilson score interval is the set of proportion values `p` that a score test would not reject at the chosen level. The key difference: the score test computes the standard error from the *hypothesised* `p`, namely `sqrt(p (1 - p) / n)`, not from `p_hat`. Solving the resulting quadratic gives ``` centre = (p_hat + z^2 / (2n)) / (1 + z^2 / n) half-width = ( z / (1 + z^2 / n) ) × sqrt( p_hat (1 - p_hat) / n + z^2 / (4 n^2) ) ``` Two features fall straight out. The centre is pulled from `p_hat` toward 0.5 — the `z^2 / (2n)` term — by an amount that shrinks as n grows. And the half-width carries an extra `z^2 / (4 n^2)` inside the square root, which stays positive even when `p_hat (1 - p_hat)` is zero. That is why the interval never collapses. For `x = 0`, `n = 20`, `z = 1.96`: the centre is about 0.081 and the half-width about 0.081, giving roughly `[0, 0.161]`. It is asymmetric about `p_hat = 0` — necessarily so, since the parameter is bounded below by 0 — and it stays inside `[0, 1]` by construction. ## The rule of three When no events are observed in n trials, a quick 95% upper bound on the rate is `3 / n`. For 20 trials that is 0.15, sitting close to the Wilson upper limit of 0.161; for 300 trials it is 0.01. The derivation is short: if the true rate were `p`, the chance of seeing zero events in n independent trials is `(1 - p)^n`, which is about `exp(-n p)`. Setting that to 0.05 gives `n p ≈ 3`. It is worth memorising because *no events observed* comes up constantly — no failures in a test run, no adverse events in a trial — and the honest answer is never *the rate is zero*. ## The alternatives, ranked for practice - **Wilson (score).** Good coverage across the range, never degenerate, never leaves `[0, 1]`. A sound default. - **Agresti-Coull.** Add about two successes and two failures, then apply the Wald formula to the adjusted counts. It approximates Wilson closely, is easy to compute by hand, and is easy to explain to non-specialists as *pretend you saw a couple of each*. - **Clopper-Pearson (exact).** Built from binomial tail probabilities; guarantees coverage at or above the nominal level, at the cost of being noticeably conservative — the intervals are wider than they need to be. Reach for it when under-coverage is unacceptable, for example in a regulated setting. - **Wald.** Fine for a rough mental estimate with a large n and a proportion near the middle. Not a good default. ## How to answer this in the room Name the mechanism first — the standard error is built from the estimate, and the estimate is at a boundary — then say what you would report instead, then give a number. Interviewers are checking that you notice a degenerate interval at all. Reporting `[0, 0]` from 20 trials without flinching is the failure mode this question exists to catch.

  • If a test run shows zero failures in 300 attempts, what upper bound would you quote?
    About 1%, from the rule of three: a 95% upper bound on the rate is roughly 3 / n, and 3 / 300 = 0.01. It follows from the chance of seeing zero events being about exp(-n p), set equal to 0.05. The honest statement is that the rate is plausibly anywhere up to about 1 in 100, not that it is zero.
  • Why is the Wilson interval asymmetric around the observed proportion?
    Because it is the set of proportion values a score test would not reject, and the score test's standard error depends on the value being tested. That makes the boundary condition asymmetric near 0 and 1. Its centre is also shifted from the observed proportion toward 0.5 by a term of order z squared over 2n, which fades as the sample grows.
  • What is the Agresti-Coull adjustment and when would you prefer it to Wilson?
    Add roughly two successes and two failures to the observed counts, then apply the ordinary Wald formula to the adjusted numbers. It tracks Wilson closely, keeps the familiar symmetric form, and is easy to compute and explain by hand — which is its main advantage when you need a defensible number quickly without a solver.
  • When would you choose Clopper-Pearson over Wilson?
    When under-coverage is unacceptable and conservatism is affordable — regulated or safety-critical reporting, for instance. Clopper-Pearson is built from exact binomial tail probabilities and guarantees at least the nominal coverage, but it pays for that guarantee with intervals that are wider than necessary at most sample sizes.

Twenty coin flips that all come up tails does not prove the coin has no heads side. An interval of [0, 0] is claiming exactly that.

saying these in an interview costs you the question

  • Reports [0, 0] as a genuine interval
  • Reads zero observed events as a zero true rate
  • Claims Wald is safe once n exceeds 30 regardless of the proportion
  • Accepts an interval limit below 0 or above 1 for a proportion
  • Thinks the Wilson interval is centred on the observed proportion

context