skip to content

When does the finite population correction meaningfully shrink a survey's standard error?

level: seniorimportance: nice to knowfreq 24%

answer

  1. sampling without replacement from a fixed list
  2. driven by the fraction sampled, not the count
  3. negligible under a five percent fraction
  4. goes to zero at a complete census
  5. corrects variance, never nonresponse bias

basics

~20 s

The finite population correction, sqrt((N-n)/(N-1)), matters only when the sample is a large fraction of the population. Below about a 5 percent sampling fraction it changes the standard error negligibly; at a census it drives it to zero.

solid answer

~50 s

The usual `s / sqrt(n)` assumes draws are independent, which holds when sampling with replacement or from an effectively unlimited population. When you sample **without replacement** from a finite list, each person you measure removes uncertainty that can never come back, so the true standard error is smaller by the factor `sqrt((N - n) / (N - 1))`. What drives it is the **sampling fraction** `n/N`, not the raw sample size. Surveying 400 of a company's 2,000 employees is a 20% fraction, giving `sqrt(1600/1999) = 0.89` - about 11% tighter than the naive figure. Below a 5% fraction the factor exceeds 0.97 and is safely ignored, which is why national polls never mention it; at `n = N` it is zero, correctly saying a census has no sampling error. Crucially it corrects **variance, not bias**: it does nothing about nonresponse.

go deeper

for a junior

Know that the standard error formula assumes the population is effectively unlimited, and that a special correction exists when your sample covers a large share of a small, fixed list of people.

for a middle

Be able to state the factor as the square root of N minus n over N minus one, evaluate it for a concrete sampling fraction, and explain that it approaches zero as the sample approaches a census.

for a senior

Show when to reach for it and when not to. Recognise internal surveys and audits as the real use cases, reject it for open-ended user populations, and refuse to let a variance correction paper over a low response rate.

for a principal

Own the survey design conversation: whether chasing a higher response rate beats enlarging the sample, how uncertainty is reported to leadership, and why a correction that tightens error bars must never be applied to a self-selected respondent pool.

## Why the usual formula needs a correction at all The standard error `sigma / sqrt(n)` is derived assuming the `n` observations are independent. That is exactly true when you sample **with replacement**, and effectively true when the population is enormous relative to the sample - drawing 1,000 people from 50 million barely changes what is left in the pool. Real surveys sample **without replacement** from a finite frame: you do not interview the same employee twice. Once you have measured someone, their contribution to the population total is known with certainty rather than estimated. In the extreme, if you interview every employee, there is no sampling uncertainty at all - yet `sigma / sqrt(n)` would still report a positive standard error. The finite population correction repairs that. ## The factor For a simple random sample of size `n` drawn without replacement from a population of size `N`: ``` SE = (sigma / sqrt(n)) * sqrt((N - n) / (N - 1)) ``` The multiplier `sqrt((N - n) / (N - 1))` is the **finite population correction**, or FPC. A near-identical form, `sqrt(1 - n/N)`, is often used interchangeably and differs only trivially for any `N` worth correcting. What matters is the **sampling fraction** `f = n / N`, not `n` on its own: | sampling fraction n/N | FPC factor | reduction in SE | |---|---|---| | 1% | 0.995 | 0.5% | | 5% | 0.975 | 2.5% | | 20% | 0.894 | about 11% | | 50% | 0.707 | about 29% | | 100% | 0 | complete - a census | ## The worked case Survey 400 of a company's 2,000 employees: ``` FPC = sqrt((2000 - 400) / (2000 - 1)) = sqrt(1600 / 1999) = sqrt(0.8004) = 0.894 ``` If engagement scores have `s = 20`, the naive standard error is `20 / sqrt(400) = 1.0`. Corrected, it is `1.0 * 0.894 = 0.89` - roughly 11% tighter. Not dramatic, but free precision, and the direction matters: ignoring the FPC is **conservative**, reporting more uncertainty than you have. ## The rule of thumb Most practitioners ignore the correction when the sampling fraction is below about **5%**, where it buys under 3% and is swamped by every other approximation in the analysis. Above roughly 10% it is worth applying, and above 30% ignoring it is a real distortion. The rule cuts both ways as a diagnostic. Someone applying an FPC because their sample of 5,000 'feels large' has misunderstood the driver: 5,000 out of 40 million is a 0.01% fraction and the factor is 0.99995. ## Where it genuinely matters - **Internal surveys.** An HR engagement survey of a 2,000-person company, or a customer survey of a 300-account enterprise book, routinely reaches double-digit sampling fractions. - **Auditing and quality inspection.** Sampling 200 invoices from a ledger of 800, or 50 units from a production lot of 200. - **School, clinic or store-level studies**, where the frame is a fixed list of a few hundred sites. ## Where invoking it is a mistake - **Consumer or web populations.** Users, sessions and requests are effectively unbounded, and the population of interest is usually a *process* generating future observations rather than a fixed list. There is no finite `N` to correct against, and inventing one to shrink error bars is an abuse. - **Anything where the target population extends into the future.** Even a complete census of this month's transactions is a sample if you want to say something about next month's. The FPC would claim zero uncertainty for a question that plainly has some. ## The correction does not fix bias This is the point interviewers actually probe. The FPC adjusts **variance** and assumes you drew a genuine simple random sample of the frame. If you invited all 2,000 employees and 400 chose to answer, you do not have a 20% random sample - you have 400 self-selected respondents. Their differences from the silent 1,600 are a **bias**, and no variance correction touches it. Shrinking the standard error by 11% in that situation makes an already-misleading estimate look more authoritative. The same reasoning explains why a very high response rate is worth more than the FPC it unlocks: it shrinks the room in which nonresponse bias can hide, which is the dominant error in most internal surveys. ## Summary for an interview Say: the correction is `sqrt((N - n) / (N - 1))`, it is driven by the sampling fraction rather than the sample size, it is negligible under about 5% and total at a census, it applies to sampling without replacement from a fixed frame, and it corrects variance only - never nonresponse bias.

  • What is the correction factor when you survey 400 of a company's 2,000 employees, and what does it buy?
    The factor is `sqrt((2000 - 400) / 1999) = sqrt(0.80) = 0.894`, so the standard error is about 11% smaller than `s / sqrt(n)` suggests. With a score spread of 20, the naive standard error of 1.0 becomes 0.89. Ignoring it is conservative rather than wrong, since it only ever overstates uncertainty.
  • Why is the correction irrelevant for a national poll of 1,000 people?
    Because the sampling fraction is microscopic. With a population in the tens of millions, `n/N` is far below 0.01%, so the factor is essentially 1 and changes the standard error in the fourth decimal place. Precision is set by the sample size, which is why a thousand-person poll works equally well for a city or a country.
  • Only 400 of 2,000 invited employees responded - can you apply the correction to that sample?
    Not legitimately as a fix. The correction assumes a genuine simple random sample of the frame; 400 self-selected respondents are not that. The dominant error is nonresponse bias - how the silent 1,600 differ - and a variance correction does nothing about bias. Shrinking the error bars here just makes a skewed estimate look more authoritative.

Tasting a spoonful tells you about the pot however big the pot is - but once you have eaten half the pot, what remains holds far fewer surprises.

saying these in an interview costs you the question

  • Applies the correction based on sample size rather than sampling fraction
  • Uses it on web or user populations with no fixed finite frame
  • Claims the correction inflates rather than shrinks the standard error
  • Treats a low response rate as a random sample of the frame
  • Thinks a census of this month proves anything about next month with zero error

context