What is the difference between simple random sampling and stratified sampling?
answer
- who gets a chance of being picked
- one pool versus several pools
- guarantees small groups appear
- between-group variation leaves the estimator
basics
~20 sSimple random sampling gives every unit in the frame an equal chance of selection. Stratified sampling first splits the population into non-overlapping groups, then samples inside each one, guaranteeing every group appears and usually giving a more precise estimate.
solid answer
~50 sSimple random sampling draws n units so that every subset of size n from the sampling frame is equally likely; each unit has inclusion probability `n/N`. Stratified sampling first partitions the frame into non-overlapping strata — say country crossed with plan tier in a customer survey — and then draws an independent random sample inside each stratum. Stratifying buys two things: small but important groups are guaranteed to appear instead of being left to chance, and variation *between* strata no longer enters the estimator, so a stratified estimate is usually at least as precise as a simple random one of the same size. The costs are that the stratum labels must already be on the frame before you draw, and that unless you allocate proportionally the sample stops being self-weighting, so weights have to be carried into every later calculation.
go deeper
Be ready to define both designs in one sentence each and name a concrete stratifying variable, such as country or plan tier in a customer survey.
Explain why stratification helps: the estimator carries only within-stratum variation, so strata that are internally similar and different from each other buy precision.
Show that you check whether the frame actually carries the stratum labels, whether each stratum gets enough units to report on, and what weights the allocation forces downstream.
Own the tradeoff between a simple self-weighting design every analyst can use correctly and a finely stratified one that is more precise but demands weights be honoured in every later query.
## What "random" actually guarantees A **simple random sample** (SRS) of size `n` from a frame of `N` units is drawn so that every possible subset of size `n` is equally likely; equivalently, every unit has the same inclusion probability `n/N`. That is a statement about the *selection procedure*, not about how the resulting data looks. A sample that "looks random" — the first 500 rows, whoever answered the email — is not an SRS, and no amount of scatter in the data makes it one. An SRS needs a **sampling frame**: an enumerable list of the units you could have drawn. The frame is where a design either succeeds or quietly fails, because a unit that is not on the list has inclusion probability zero no matter how carefully you randomise afterwards. The payoff of an SRS is simplicity. Every unit carries the same weight, the plain sample mean is an unbiased estimator of the population mean, and every downstream analyst can compute an average without knowing anything about how the rows were chosen. ## Stratified sampling A **stratified** design partitions the frame into `H` non-overlapping strata that together cover the whole population, and draws an independent sample inside each. Every unit belongs to exactly one stratum, and every stratum is sampled — that last part is what separates strata from clusters, where only some groups are selected at all. A concrete case: a customer survey stratified by country (US, DE, JP) crossed with plan tier (free, pro, enterprise) gives nine strata. Inside each of the nine you take a simple random sample of accounts. Two motives justify the extra machinery. **Guaranteed representation.** If the enterprise tier is 1% of accounts, a simple random sample of 500 accounts gives about 5 enterprise customers on average, and a real chance of getting one or none. You cannot report on a group that the draw happened to skip. Stratification fixes the count in advance: you decide that 60 of the 500 come from enterprise, and they do. **Precision.** Total variability in the population can be decomposed into variation *within* strata and variation *between* stratum means. Because a stratified design fixes how many units come from each stratum instead of letting chance decide, the between-stratum component drops out of the estimator's sampling variability — the estimator only "feels" the within-stratum scatter. So the gain is large when strata are internally homogeneous and clearly different from each other, and essentially zero when every stratum has the same mean. Stratify on variables that separate the outcome, not on whatever happens to be convenient. ## Putting the strata back together The stratified estimate of a population mean is a weighted combination of stratum means: `ybar_str = sum over h of W_h * ybar_h`, where `W_h = N_h / N` is the stratum's share of the population and `ybar_h` is its sample mean. Two consequences follow. If the allocation is **proportional** — each stratum's share of the sample equals its share of the population — then every unit had the same inclusion probability, the design is *self-weighting*, and `ybar_str` is numerically identical to the plain sample mean. Nothing downstream has to change. If the allocation is **disproportionate** — you deliberately took more enterprise accounts than their population share — the plain average of the responses is no longer an estimate of anything you wanted. It estimates the mean of a fictional population that is mostly enterprise. You must apply weights. ## With or without replacement A second design choice sits underneath both schemes: whether a unit can be drawn twice. Drawing 100 rows from a 1,000-row table **without replacement** yields 100 distinct rows; the draws are slightly negatively correlated, and the sample is a little more precise than an equivalent independent draw — the finite-population correction factor `(N - n) / (N - 1)` is what quantifies that gain. Drawing **with replacement** makes the draws independent but lets rows repeat: at `n = 100`, `N = 1000` you expect roughly 95 distinct rows among the 100 draws. When the sampling fraction `n/N` is small the two are nearly indistinguishable; when you are sampling half the table they are not. ## What stratification does not do It does not repair a frame that omits part of the population, and it does not turn a self-selected group of respondents into a probability sample — those are frame and non-response problems, and they live outside the design itself. Stratification also multiplies bookkeeping: more strata means fewer units per stratum, and a stratum with three respondents is not reportable no matter how carefully it was drawn. ## A short checklist Before choosing between the two designs, ask: does the frame already carry the candidate stratifying variable? Do the strata differ on the outcome? Will you need to report on any group separately? Will the sample stay self-weighting, and if not, does everyone consuming the data know that weights exist? An SRS that everyone uses correctly often beats a clever stratification that one analyst quietly averages unweighted.
- If you sample 100 rows from a 1,000-row table with replacement instead of without, what changes?With replacement a row can be drawn more than once, so among 100 draws you expect only about 95 distinct rows, and the draws are independent. Without replacement all 100 rows differ, the draws are slightly negatively correlated, and the sample is a bit more precise for the same n — the finite-population correction (N - n) / (N - 1) captures that. At a 10% sampling fraction the difference is minor; at 50% it is not.
- When does stratifying give you essentially no gain over a simple random sample?When the strata do not differ on the quantity you are estimating. Stratification removes between-stratum variation from the estimator, so if every stratum has the same mean there is nothing to remove and precision is unchanged. You still pay for building the strata and, under a disproportionate allocation, for carrying weights. Stratify on variables that actually separate the outcome.
- How do you choose a stratifying variable in practice?Pick something that is known for every unit on the frame, that splits the population into groups differing on the outcome, and that you will want to report on separately. Country and plan tier in a customer survey qualify on all three counts. Keep the number of strata small enough that each still receives enough units to estimate within.
Simple random sampling is scooping one cup from a stirred soup and trusting that every ingredient made it in. Stratified sampling is deliberately taking a spoonful of broth, a spoonful of beans and a spoonful of meat, in known proportions.
saying these in an interview costs you the question
- Says stratified sampling means surveying only the interesting group
- Treats any convenient split of the data as a random sample
- Claims stratification repairs a frame that omits people
- Forgets that a disproportionate allocation requires weights
- Confuses strata, which are all sampled, with clusters, which are not