skip to content

Why does an election poll of 1,000 respondents report a margin of error near 3 points?

level: middleimportance: should knowfreq 58%

answer

  1. precision of a percentage, not of an average
  2. a yes/no answer is one coin flip
  3. one response has variance p(1-p)
  4. worst case sits at fifty-fifty
  5. headline margin is about two standard errors

basics

~20 s

The standard error of a sample proportion is sqrt(p(1-p)/n), which at p = 0.5 and n = 1,000 is about 1.6 percentage points. A reported margin of error is conventionally about two standard errors, so roughly 3 points.

solid answer

~40 s

A yes/no answer is a coin flip with success probability `p`, so a single response has variance `p(1-p)`. Averaging `n` of them gives a sample proportion with standard error `sqrt(p(1-p)/n)`. At `n = 1,000` and `p = 0.5` that is `sqrt(0.25/1000) = 0.0158`, about 1.6 percentage points, and the headline margin of error is conventionally about two standard errors - roughly 3 points. Two features are worth naming. First, `p(1-p)` peaks at `p = 0.5` and is flat nearby, so pollsters quote the `p = 0.5` worst case and it barely changes for any realistic result; at `p = 0.2` the standard error drops only to 1.3 points. Second, the number covers **sampling** error alone. Nonresponse, a skewed sampling frame and question wording are not in it, and are often the larger problem.

go deeper

for a junior

Recall that the standard error of a sample proportion is the square root of p times one minus p over n, and be able to evaluate it at one half and a thousand to get about 1.6 points.

for a middle

Explain the mechanics: a yes/no answer is a single Bernoulli draw with variance p(1-p), a proportion is the mean of those draws, and the worst case sits at one half, which is why pollsters quote it.

for a senior

Show operational judgment - recompute the standard error for a crosstab subgroup before reacting to it, and state clearly that sampling error is only one component of a survey's total error alongside nonresponse and frame problems.

for a principal

Own how uncertainty is communicated. Decide what your organisation publishes alongside survey numbers, resist headline claims resting on subgroup noise, and set the expectation that a quoted margin is a floor on uncertainty, not a ceiling.

## A proportion is an average of zeros and ones Code each respondent as 1 for 'will vote for A' and 0 otherwise. The sample proportion `phat` is then just the sample mean of those indicators, so everything known about the standard error of a mean applies - only the spread has a special form. A single 0/1 response with success probability `p` has ``` variance = p(1 - p) standard deviation = sqrt(p(1 - p)) ``` So the standard error of the sample proportion is ``` SE = sqrt(p(1 - p) / n) ``` In practice `p` is unknown and `phat` is substituted. ## The 1,000-respondent arithmetic With `n = 1,000` and `p = 0.5`: ``` SE = sqrt(0.25 / 1000) = sqrt(0.00025) = 0.0158 ``` That is **1.58 percentage points**. A reported margin of error is conventionally about two standard errors, giving roughly **3.1 points** - which is why survey after survey quotes 'plus or minus 3' for a thousand-person sample. The number is a property of the arithmetic, not of the electorate. ## Why pollsters use p = 0.5 The function `p(1 - p)` is a downward parabola peaking at `p = 0.5`, where it equals 0.25. Quoting the margin at `p = 0.5` is therefore the **worst case**, valid for any result. It is also remarkably flat near the middle: | p | p(1 - p) | SE at n = 1,000 | |---|---|---| | 0.50 | 0.250 | 1.58 pts | | 0.40 | 0.240 | 1.55 pts | | 0.20 | 0.160 | 1.26 pts | | 0.05 | 0.0475 | 0.69 pts | Between 40% and 60% the standard error moves by less than 2% of itself, so a single headline margin genuinely covers every plausible outcome. Only for rare events - a candidate at 5% - does the true precision improve much, which matters when reporting small categories. ## Sample size, not population size `sqrt(p(1-p)/n)` contains no term for how many people live in the country. A 1,000-person sample is about equally precise for a city of 200,000 and a nation of 200 million, which strikes most people as wrong the first time they hear it. The reason is that a random sample carries information about the population's composition, not its headcount. (A correction does exist for cases where the sample is a sizeable fraction of a small population, but a national poll is nowhere near that regime.) What *does* move the number is `n`, and only as `1 / sqrt(n)`: | n | SE at p = 0.5 | approximate margin | |---|---|---| | 400 | 2.5 pts | 5 pts | | 1,000 | 1.58 pts | 3 pts | | 1,500 | 1.29 pts | 2.5 pts | | 4,000 | 0.79 pts | 1.5 pts | Getting from 3 points to 1.5 points requires four times the fieldwork - which is precisely why 1,000 is such a common sample size. It is the point where cost and precision cross for most published polling. ## Subgroups are far noisier than the headline The most abused figure in poll reporting is the crosstab. If 250 of the 1,000 respondents are under 30, their subgroup proportion has ``` SE = sqrt(0.25 / 250) = 0.0316 ``` about 3.2 points, so a margin near 6.3 points - **double** the headline. A ten-point swing in a subgroup between two consecutive polls is often entirely noise. Whenever a story turns on a slice, recompute the standard error from that slice's own `n`. ## What the margin of error leaves out This is where a strong candidate separates from a formula-reciter. `sqrt(p(1-p)/n)` quantifies the randomness of drawing `n` people **at random from the target population**. It says nothing about: - **Nonresponse.** If 4% of those contacted answer and they differ systematically from those who do not, the resulting error is not in the formula and does not shrink with `n`. - **Frame coverage.** People unreachable by the contact method are absent from every sample. - **Question wording and order**, which can shift results by more than the quoted margin. - **Turnout modelling**, since the population of interest - people who will actually vote - is not directly observable. Historically, polling misses have far more often come from these sources than from sampling noise. Treating a plus-or-minus 3 as the total uncertainty of a poll is the single most common misreading of the number. ## What to say in an interview Write `SE = sqrt(p(1-p)/n)`, evaluate it at 0.5 and 1,000 to get 1.6 points, note that the reported margin is about twice that, mention the worst-case reasoning behind `p = 0.5`, and then volunteer that the figure is sampling error only. That last sentence is usually what the question was testing.

  • Which true proportion makes the standard error of a poll largest, and why?
    A true proportion of 0.5. The variance of a single yes/no response is `p(1-p)`, a downward parabola peaking at 0.25 when `p = 0.5`. Quoting the margin there is the worst case and so is valid whatever the result turns out to be. It is also flat nearby, so the headline barely changes across any competitive race.
  • A story reports a 10-point shift among the 250 under-30 respondents in a 1,000-person poll - how do you react?
    Recompute the standard error on that subgroup's own sample size: `sqrt(0.25/250) = 0.032`, about 3.2 points, so a margin near 6 points - double the headline. Two independent subgroup readings differing by 10 points is comfortably inside noise. Crosstab swings are the most over-interpreted numbers in survey reporting.
  • Does the reported margin of error account for people refusing to take part?
    No. The formula quantifies only the randomness of drawing `n` people at random from the target population. Nonresponse, unreachable groups, question wording and turnout modelling sit entirely outside it, and unlike sampling noise they do not shrink as the sample grows. Historically they explain most large polling misses.

saying these in an interview costs you the question

  • Thinks the margin of error covers nonresponse and question wording
  • Believes a larger country needs a proportionally larger poll
  • Reports the standard error itself as the margin of error
  • Applies the headline margin to a small crosstab subgroup
  • Treats the margin of error as a guaranteed bound on the truth

context