skip to content

Sampling Design & Bias

How the rows were chosen: simple random, stratified and cluster designs, and the convenience samples, non-response and survivorship that skew every later number. Interviewers probe who is missing.

on this pageshow

explore

questions

15

What is the difference between simple random sampling and stratified sampling?

level: juniorimportance: must knowfreq 76%

answer

  1. who gets a chance of being picked
  2. one pool versus several pools
  3. guarantees small groups appear
  4. between-group variation leaves the estimator

basics

~20 s

Simple random sampling gives every unit in the frame an equal chance of selection. Stratified sampling first splits the population into non-overlapping groups, then samples inside each one, guaranteeing every group appears and usually giving a more precise estimate.

solid answer

~50 s

Simple random sampling draws n units so that every subset of size n from the sampling frame is equally likely; each unit has inclusion probability `n/N`. Stratified sampling first partitions the frame into non-overlapping strata — say country crossed with plan tier in a customer survey — and then draws an independent random sample inside each stratum. Stratifying buys two things: small but important groups are guaranteed to appear instead of being left to chance, and variation *between* strata no longer enters the estimator, so a stratified estimate is usually at least as precise as a simple random one of the same size. The costs are that the stratum labels must already be on the frame before you draw, and that unless you allocate proportionally the sample stops being self-weighting, so weights have to be carried into every later calculation.

go deeper

for a junior

Be ready to define both designs in one sentence each and name a concrete stratifying variable, such as country or plan tier in a customer survey.

for a middle

Explain why stratification helps: the estimator carries only within-stratum variation, so strata that are internally similar and different from each other buy precision.

for a senior

Show that you check whether the frame actually carries the stratum labels, whether each stratum gets enough units to report on, and what weights the allocation forces downstream.

for a principal

Own the tradeoff between a simple self-weighting design every analyst can use correctly and a finely stratified one that is more precise but demands weights be honoured in every later query.

## What "random" actually guarantees A **simple random sample** (SRS) of size `n` from a frame of `N` units is drawn so that every possible subset of size `n` is equally likely; equivalently, every unit has the same inclusion probability `n/N`. That is a statement about the *selection procedure*, not about how the resulting data looks. A sample that "looks random" — the first 500 rows, whoever answered the email — is not an SRS, and no amount of scatter in the data makes it one. An SRS needs a **sampling frame**: an enumerable list of the units you could have drawn. The frame is where a design either succeeds or quietly fails, because a unit that is not on the list has inclusion probability zero no matter how carefully you randomise afterwards. The payoff of an SRS is simplicity. Every unit carries the same weight, the plain sample mean is an unbiased estimator of the population mean, and every downstream analyst can compute an average without knowing anything about how the rows were chosen. ## Stratified sampling A **stratified** design partitions the frame into `H` non-overlapping strata that together cover the whole population, and draws an independent sample inside each. Every unit belongs to exactly one stratum, and every stratum is sampled — that last part is what separates strata from clusters, where only some groups are selected at all. A concrete case: a customer survey stratified by country (US, DE, JP) crossed with plan tier (free, pro, enterprise) gives nine strata. Inside each of the nine you take a simple random sample of accounts. Two motives justify the extra machinery. **Guaranteed representation.** If the enterprise tier is 1% of accounts, a simple random sample of 500 accounts gives about 5 enterprise customers on average, and a real chance of getting one or none. You cannot report on a group that the draw happened to skip. Stratification fixes the count in advance: you decide that 60 of the 500 come from enterprise, and they do. **Precision.** Total variability in the population can be decomposed into variation *within* strata and variation *between* stratum means. Because a stratified design fixes how many units come from each stratum instead of letting chance decide, the between-stratum component drops out of the estimator's sampling variability — the estimator only "feels" the within-stratum scatter. So the gain is large when strata are internally homogeneous and clearly different from each other, and essentially zero when every stratum has the same mean. Stratify on variables that separate the outcome, not on whatever happens to be convenient. ## Putting the strata back together The stratified estimate of a population mean is a weighted combination of stratum means: `ybar_str = sum over h of W_h * ybar_h`, where `W_h = N_h / N` is the stratum's share of the population and `ybar_h` is its sample mean. Two consequences follow. If the allocation is **proportional** — each stratum's share of the sample equals its share of the population — then every unit had the same inclusion probability, the design is *self-weighting*, and `ybar_str` is numerically identical to the plain sample mean. Nothing downstream has to change. If the allocation is **disproportionate** — you deliberately took more enterprise accounts than their population share — the plain average of the responses is no longer an estimate of anything you wanted. It estimates the mean of a fictional population that is mostly enterprise. You must apply weights. ## With or without replacement A second design choice sits underneath both schemes: whether a unit can be drawn twice. Drawing 100 rows from a 1,000-row table **without replacement** yields 100 distinct rows; the draws are slightly negatively correlated, and the sample is a little more precise than an equivalent independent draw — the finite-population correction factor `(N - n) / (N - 1)` is what quantifies that gain. Drawing **with replacement** makes the draws independent but lets rows repeat: at `n = 100`, `N = 1000` you expect roughly 95 distinct rows among the 100 draws. When the sampling fraction `n/N` is small the two are nearly indistinguishable; when you are sampling half the table they are not. ## What stratification does not do It does not repair a frame that omits part of the population, and it does not turn a self-selected group of respondents into a probability sample — those are frame and non-response problems, and they live outside the design itself. Stratification also multiplies bookkeeping: more strata means fewer units per stratum, and a stratum with three respondents is not reportable no matter how carefully it was drawn. ## A short checklist Before choosing between the two designs, ask: does the frame already carry the candidate stratifying variable? Do the strata differ on the outcome? Will you need to report on any group separately? Will the sample stay self-weighting, and if not, does everyone consuming the data know that weights exist? An SRS that everyone uses correctly often beats a clever stratification that one analyst quietly averages unweighted.

  • If you sample 100 rows from a 1,000-row table with replacement instead of without, what changes?
    With replacement a row can be drawn more than once, so among 100 draws you expect only about 95 distinct rows, and the draws are independent. Without replacement all 100 rows differ, the draws are slightly negatively correlated, and the sample is a bit more precise for the same n — the finite-population correction (N - n) / (N - 1) captures that. At a 10% sampling fraction the difference is minor; at 50% it is not.
  • When does stratifying give you essentially no gain over a simple random sample?
    When the strata do not differ on the quantity you are estimating. Stratification removes between-stratum variation from the estimator, so if every stratum has the same mean there is nothing to remove and precision is unchanged. You still pay for building the strata and, under a disproportionate allocation, for carrying weights. Stratify on variables that actually separate the outcome.
  • How do you choose a stratifying variable in practice?
    Pick something that is known for every unit on the frame, that splits the population into groups differing on the outcome, and that you will want to report on separately. Country and plan tier in a customer survey qualify on all three counts. Keep the number of strata small enough that each still receives enough units to estimate within.

Simple random sampling is scooping one cup from a stirred soup and trusting that every ingredient made it in. Stratified sampling is deliberately taking a spoonful of broth, a spoonful of beans and a spoonful of meat, in known proportions.

saying these in an interview costs you the question

  • Says stratified sampling means surveying only the interesting group
  • Treats any convenient split of the data as a random sample
  • Claims stratification repairs a frame that omits people
  • Forgets that a disproportionate allocation requires weights
  • Confuses strata, which are all sampled, with clusters, which are not

context

open as a page

What is the difference between MCAR, MAR and MNAR missing data?

level: juniorimportance: must knowfreq 74%

basics

~20 s

MCAR means missingness is unrelated to any variable. MAR means it depends only on variables you observed. MNAR means it depends on the missing value itself, as when the highest earners are the ones who leave salary blank.

open as a page

How does survivorship bias inflate the average return in a table of funds that still exist today?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Badly performing funds get closed or merged away, so a table of funds still open today lists mostly winners. Averaging it measures the survivors, not the return an investor could have expected when picking a fund years ago.

open as a page

Why does filling missing numeric values with the column mean understate the standard deviation?

level: middleimportance: must knowfreq 58%

basics

~20 s

Every filled value sits exactly at the mean, so it adds nothing to the sum of squared deviations while still adding to the row count. The numerator is unchanged and the denominator grows, so variance and standard deviation shrink.

open as a page

Why did the 1936 Literary Digest poll get the election wrong despite millions of returned ballots?

level: middleimportance: must knowfreq 58%

basics

~20 s

Two selection failures compounded. The mailing list was built from telephone directories, car registrations and subscriber rolls, which excluded poorer households; and only about a quarter of recipients mailed a ballot back, and returners differed from non-returners.

open as a page

In a stratified survey, when does Neyman allocation beat proportional allocation?

level: middleimportance: should knowfreq 44%

basics

~20 s

Proportional allocation sizes each stratum by its share of the population. Neyman allocation sizes it by population share times within-stratum standard deviation, so it wins when strata differ sharply in variability: it puts interviews where the answers scatter most.

open as a page

Why can systematic sampling of every 7th day of logs give a biased estimate?

level: middleimportance: should knowfreq 33%

basics

~20 s

Taking every 7th day always lands on the same weekday. Systematic sampling picks every k-th unit after a random start, so when the ordering carries a cycle of period k, the sample sees one phase of that cycle and misses the rest.

open as a page

Why are app-store star ratings a biased estimate of average user satisfaction?

level: middleimportance: should knowfreq 42%

basics

~20 s

Reviewers volunteer. Posting a rating takes effort, and the people willing to spend it are disproportionately delighted or furious, while the indifferent majority stays silent. The published average estimates the mean rating among reviewers, not among users.

open as a page

Why does sampling whole classrooms instead of individual students cost you effective sample size?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Students in a classroom resemble one another, so each extra student adds less new information than an independently drawn one. That similarity inflates the variance by the design effect, roughly 1 + (m - 1) * rho for clusters of size m.

open as a page

Complete-case analysis drops 40% of your rows — when does that bias the estimate?

level: seniorimportance: should knowfreq 50%

basics

~20 s

The share dropped controls precision; the mechanism controls bias. If values are missing completely at random, the surviving 60% are a random subsample and the estimate stays unbiased, only noisier. Under MAR or MNAR it can be badly off.

open as a page

How does multiple imputation produce standard errors that reflect missing-data uncertainty?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Several completed datasets are built by drawing missing values from a model, each analysed separately, then pooled. The variability of the estimates across datasets is added to the average within-dataset variance, so the standard error grows with how much was missing.

open as a page

How would you post-stratify an age-skewed online panel to census shares, and when does bias remain?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Split respondents into age cells, weight each cell by population share divided by sample share, and average the cell means with population weights. Bias remains whenever respondents inside a cell still differ from non-respondents on the outcome.

open as a page

How do you estimate an overall mean when power users were deliberately oversampled?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Weight each respondent by the inverse of their selection probability and report the weighted mean, sum(w_i * y_i) / sum(w_i). Oversampled power users get small weights and ordinary users large ones, so the estimate reflects the real population mix.

open as a page

A key field may be MNAR and the data cannot prove otherwise — how do you proceed?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Treat MNAR as an assumption to declare rather than a hypothesis to test. Report the missingness rate, run a sensitivity analysis that shifts the imputed values until the conclusion flips, and where the stakes justify it, go and observe the non-responders.

open as a page

How do you decide whether a KPI computed only over still-active users is trustworthy enough to publish?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Ask what the number claims to describe. A metric whose denominator is this month's actives moves when weak users leave, so it reports composition, not behaviour. Publish it only alongside a cohort-anchored version, and name the conditioning in the metric itself.

open as a page