How do you normal-approximate a Binomial(200, 0.1) probability with a continuity correction?
answer
- a count is a sum of coin flips
- match the first two moments
- discrete mass versus smooth density
- shift the boundary by half a unit
- np and n(1-p) both above ten
basics
~20 sMatch the normal's mean to np = 20 and its standard deviation to sqrt(18) = 4.24, then shift the boundary half a unit: for P(X <= 25) use 25.5, giving z = 1.30 and about 0.903.
solid answer
~50 sA binomial count is a sum of n independent Bernoulli draws, so the Central Limit Theorem applies to it directly. For `Binomial(200, 0.1)` the mean is `np = 20` and the variance is `np(1-p) = 18`, so the standard deviation is `sqrt(18) ~ 4.243`. To approximate `P(X <= 25)` you replace the discrete boundary 25 with 25.5 — the **continuity correction** — because the normal spreads each integer's probability across a unit-wide interval. Then `z = (25.5 - 20) / 4.243 ~ 1.296`, giving about 0.903. The exact binomial answer is 0.8995. Without the correction you would use `z = (25 - 20) / 4.243 ~ 1.179` and get 0.881, roughly two percentage points off. The rule of thumb for when the approximation is usable at all is `np >= 10` and `n(1-p) >= 10`; here np = 20 and n(1-p) = 180, comfortably satisfied.
go deeper
Know that a binomial count has mean np and variance np(1-p), and that a half-unit continuity correction exists when you swap a discrete count for a smooth curve.
Be able to run the full calculation live: moments, corrected boundary, z-score, tail probability. Explain in one sentence why the half-unit shift exists at all.
Demonstrate the judgment about applicability — check np and n(1-p), recognise the small-np regime where a different limit takes over, and note that tail accuracy is worse than central accuracy.
Own the framing that exact computation is cheap now, so the approximation earns its keep in closed-form reasoning and sample-size algebra rather than as a substitute for the exact number.
## Why the CLT applies to a binomial A `Binomial(n, p)` random variable counts successes in n independent trials, each succeeding with probability p. That count is literally a sum of n i.i.d. Bernoulli(p) variables, each with mean p and variance `p(1-p)`. Since a Bernoulli has finite variance, the Central Limit Theorem applies to the sum: for large n the count is approximately normal with - mean `np` - variance `np(1-p)` - standard deviation `sqrt(np(1-p))` For `n = 200, p = 0.1`: mean 20, variance 18, standard deviation about 4.243. ## The continuity correction The binomial is **discrete** — it puts probability mass on the integers 0, 1, 2, ... The normal is **continuous** — it spreads probability smoothly and assigns zero probability to any single point. Bridging them requires deciding how much of the continuous density belongs to each integer, and the natural convention gives integer k the interval from `k - 0.5` to `k + 0.5`. This produces the correction rules: - `P(X <= k)` becomes `P(Normal <= k + 0.5)` - `P(X < k)` becomes `P(Normal <= k - 0.5)` - `P(X >= k)` becomes `P(Normal >= k - 0.5)` - `P(X > k)` becomes `P(Normal >= k + 0.5)` - `P(X = k)` becomes `P(k - 0.5 <= Normal <= k + 0.5)` The rule that keeps you out of trouble: the corrected boundary always moves so as to **include** the integer you meant to include, and to **exclude** the one you meant to exclude. Note that `P(X <= k)` and `P(X < k)` differ for a discrete variable, which is exactly why they map to different corrected boundaries. ## The worked numbers Approximate `P(X <= 25)` for `Binomial(200, 0.1)`. 1. Mean `np = 200 * 0.1 = 20`. 2. Standard deviation `sqrt(200 * 0.1 * 0.9) = sqrt(18) ~ 4.243`. 3. Corrected boundary: 25.5. 4. `z = (25.5 - 20) / 4.243 ~ 1.296`. 5. Normal lower-tail probability at 1.296 is about **0.9026**. The exact binomial value is **0.8995**, so the corrected approximation is off by about 0.003. Skipping the correction gives `z = (25 - 20) / 4.243 ~ 1.179` and about **0.8807**, off by about 0.019 — more than six times the error. The correction matters most exactly where people are tempted to skip it: modest n, and boundaries near the centre where a half-unit is a meaningful slice of the distribution. ## When the approximation is usable The standard guideline is `np >= 10` **and** `n(1-p) >= 10`. Both are needed because a binomial is skewed whenever p is far from 0.5, and the skew shows up in whichever tail is squeezed against a boundary. The count cannot go below 0 or above n, so when np is small the distribution piles up against zero and cannot look symmetric no matter what the normal curve says. For `Binomial(200, 0.1)`: `np = 20` and `n(1-p) = 180`, both comfortably above 10, so the approximation is sound — as the 0.003 error confirms. Contrast `Binomial(20, 0.05)`, where `np = 1`: the normal approximation would put meaningful probability on negative counts, which is nonsense. For very small np the appropriate limit is a different one entirely — the count converges to a Poisson rather than to a normal. ## Practical notes **Tail accuracy is worse than central accuracy.** The approximation is best near the mean and degrades in the far tails, precisely where extreme probabilities are being estimated. Treat an approximated probability of, say, one in a million with scepticism. **Proportions follow the same algebra.** If you care about the observed proportion `X / n` rather than the count, divide everything by n: the mean becomes p and the standard deviation becomes `sqrt(p(1-p)/n)`. The continuity correction becomes a shift of `0.5 / n` on the proportion scale, which is easy to forget. **Say why you are approximating.** With modern computation the exact binomial probability is directly computable, so the honest framing is that the normal approximation is valuable for reasoning, for closed-form algebra like sample-size formulas, and for intuition about how far a count can drift — not because the exact number is unreachable.
- Which boundary do you use for P(X >= 26) rather than P(X <= 25)?Use 25.5 again, but as a lower bound: `P(X >= 26)` becomes `P(Normal >= 25.5)`. The boundary moves down half a unit so that 26 stays inside the region and 25 stays outside. The two answers are complements, which is a useful arithmetic check.
- Why does the np >= 10 rule involve n(1-p) as well?A binomial is symmetric only at p = 0.5; away from that it is skewed toward whichever boundary is closer. Requiring both `np >= 10` and `n(1-p) >= 10` keeps the count far enough from both 0 and n that the normal's unbounded tails do not put meaningful mass on impossible values.
- When p is very small, what limit does the count approach instead of a normal?When n is large and p is small with `np` held at a moderate constant, the binomial converges to a Poisson distribution with that same mean rather than to a normal. That regime is where the normal approximation fails worst, because the true distribution stays visibly right-skewed and bounded below at zero.
The binomial stacks probability in unit-wide bricks on the integers; the normal is a smooth ramp. The half-unit shift lines the ramp up with the edge of the brick rather than its centre.
saying these in an interview costs you the question
- Omits the continuity correction on a discrete count
- Shifts the boundary the wrong way, flipping include and exclude
- Uses np as the standard deviation instead of the variance
- Applies the approximation when np is far below ten
- Treats P(X <= k) and P(X < k) as identical for a discrete count