Why did the 1936 Literary Digest poll get the election wrong despite millions of returned ballots?
answer
- count who the list could never reach
- phone and car owners, 1936
- only about a quarter mailed it back
- two selection steps, not one
- n shrinks the interval, not the offset
basics
~20 sTwo selection failures compounded. The mailing list was built from telephone directories, car registrations and subscriber rolls, which excluded poorer households; and only about a quarter of recipients mailed a ballot back, and returners differed from non-returners.
solid answer
~40 sThe Literary Digest mailed ballots to roughly ten million names drawn from telephone directories, automobile registration lists and its own subscriber list. In 1936 those lists over-represented affluent households, so the **sampling frame** did not cover the electorate — that is undercoverage, or frame error. On top of that, only about a quarter of recipients bothered to return a ballot, and people motivated to respond were not a random subset of those mailed — that is non-response bias. The poll predicted a comfortable win for Alf Landon; Franklin Roosevelt took 46 of the 48 states. The number of responses is irrelevant to either problem: the standard error shrinks like `1 / sqrt(n)` while the bias is a fixed offset, so an enormous biased sample gives you a very tight interval around the wrong answer.
go deeper
Be ready to say that the ballots went to phone and car owners, who were not a cross-section of voters in 1936, and that a huge number of responses does not make a sample representative.
Separate the two mechanisms out loud: a frame that gave part of the electorate zero chance of selection, and a roughly one-in-four response rate that filtered again. Explain why bias is a fixed offset while the standard error falls like one over the square root of n.
Show how you would have caught it live: benchmark respondent composition against known population figures, monitor response rate by subgroup rather than raw counts, and refuse to publish a margin of error that ignores coverage.
Frame the organisational lesson: set a standard that every survey result ships with its frame definition and response rate, because a precise interval around an uncovered population is the most expensive kind of wrong number a company can publish.
## What happened In 1936 the *Literary Digest* ran what was, by volume, the largest election poll ever attempted. It mailed roughly ten million straw-vote ballots and received back on the order of 2.4 million. It forecast a comfortable victory for the Republican challenger, Alf Landon. Franklin Roosevelt won 46 of the 48 states. Meanwhile a far smaller poll run by George Gallup, with a sample in the tens of thousands, called the result correctly. The episode is the canonical demonstration that **sample quality dominates sample size**. ## Failure one: the sampling frame A **sampling frame** is the list from which you actually draw. It is not the population; it is your operational stand-in for the population. The Digest built its frame from telephone directories, automobile registration records and its own subscriber list. In the middle of the Depression, owning a telephone or a car was a marker of relative affluence, and affluence was correlated with the very thing being measured — vote intention. Large parts of the electorate had **zero probability** of ever being sampled. That is **undercoverage**, also called frame error or coverage error. The formal point matters: any sampling theory that promises unbiasedness assumes every member of the target population has a known, non-zero inclusion probability. When a subgroup has probability zero, no amount of sampling from the frame recovers it, and no confidence interval computed from the sample accounts for the missing group — the interval describes uncertainty about the frame's mean, not the population's. ## Failure two: who mailed the ballot back About one in four recipients returned a ballot. Response is a second, self-administered selection step layered on top of the frame. Whether someone bothers to fill in and post a straw ballot depends on how strongly they feel, and strength of feeling was itself related to which candidate they preferred. The people who responded therefore differed from the people who did not, on the exact variable being estimated. Here non-response is best understood as **a frame-and-weighting problem**: the effective sample is not "people mailed a ballot" but "people mailed a ballot who chose to answer", and unless you can characterise how those two groups differ, you cannot correct the estimate. A high absolute response *count* tells you nothing; the response *rate*, and how it varies across subgroups, is the informative quantity. ## Why the enormous n did not save it Decompose the error of an estimator into bias and variability: `error = (estimate - frame mean) + (frame mean - population mean)` The first term is sampling variability and shrinks roughly as `1 / sqrt(n)`. The second term is the bias induced by frame and response selection, and it does **not** depend on `n` at all. Going from ten thousand to two million responses divides the first term by about fourteen and leaves the second untouched. The practical result is worse than a small poll: the reported margin of error becomes tiny, which manufactures confidence in a wrong number. A biased big sample is more dangerous than a biased small one because it looks authoritative. ## What a better design would have looked like - **Fix the frame first.** Use a frame that reaches the whole target population — in that era, area or household-based sampling rather than lists of consumer goods owners. - **Control who responds, not just who is contacted.** Chase non-respondents, offer alternative response modes, and measure the response rate by subgroup so the shortfall is visible. - **Compare the achieved sample to known population composition.** If your respondents are richer, older or more urban than the population, you can see the coverage problem before the result embarrasses you. - **Report the response rate next to the estimate.** A poll quoting a raw count without a rate is hiding the second selection step. ## The modern version Swap telephone directories for an opt-in web panel, a banner ad recruiting respondents, or an intercept survey on a single page of a site, and the identical two-step failure reappears: a frame that reaches only part of the population, and a response step that filters again by interest. Very large modern samples reproduce the Digest's mistake routinely, and for the same reason — the analyst reports `n` and never reports who could not be reached and who declined. ## What interviewers listen for A strong answer separates the two failures rather than blurring them into "the sample was bad", states that bias does not shrink with `n`, and names the diagnostic that would have caught it: compare the composition of respondents against a trusted external benchmark, and look at the response rate rather than the response count.
- With the same budget, what design change would have fixed it?Spend far less on volume and far more on frame coverage and follow-up. Draw from a frame that reaches the whole electorate rather than lists of consumer-goods owners, use a much smaller sample, and pursue non-respondents so the response rate is high and measurable. A well-covered sample of tens of thousands beats a badly covered one of millions, which is exactly what happened in 1936.
- How would you detect this failure mode in a survey you are running today?Compare the composition of your respondents against trusted external benchmarks on variables you can observe for everyone contacted, and track the response rate separately by subgroup rather than quoting a total count. A large gap between contacted and responding composition is direct evidence of a selection problem, and early-versus-late responders give a crude proxy for what the never-responders look like.
- Is undercoverage the same thing as non-response?No. Undercoverage means a person could never have been selected because the frame does not contain them, so their inclusion probability is zero. Non-response means they were selected and declined, so they exist in the frame and could in principle be characterised and followed up. The remedies differ: undercoverage needs a better frame, non-response needs follow-up or weighting.
saying these in an interview costs you the question
- Blames bad luck or ordinary sampling error
- Says a larger sample would have fixed it
- Confuses the frame problem with biased question wording
- Assumes a big response count implies a good response rate
- Treats undercoverage and non-response as the same failure