skip to content

In a stratified survey, when does Neyman allocation beat proportional allocation?

level: middleimportance: should knowfreq 44%

answer

  1. how many units per stratum
  2. population share versus internal spread
  3. N_h times S_h, not N_h alone
  4. the sample stops being self-weighting

basics

~20 s

Proportional allocation sizes each stratum by its share of the population. Neyman allocation sizes it by population share times within-stratum standard deviation, so it wins when strata differ sharply in variability: it puts interviews where the answers scatter most.

solid answer

~50 s

Proportional allocation sets `n_h` proportional to `N_h`, the stratum's population count. The result is self-weighting — every unit has the same inclusion probability and the plain sample mean is already the right estimate. Neyman allocation sets `n_h` proportional to `N_h * S_h`, where `S_h` is the within-stratum standard deviation; that choice minimises the variance of the overall estimate for a fixed total sample size. It pays off when strata are very unequal in spread: a small enterprise tier whose spend ranges over orders of magnitude deserves far more interviews than its head count suggests, while a homogeneous free tier needs few. The costs are that you need a decent prior estimate of each `S_h`, the sample is no longer self-weighting so weights must be carried, and an allocation tuned for one outcome can be poor for another. With unequal per-unit costs `c_h`, the optimum adds a `1 / sqrt(c_h)` factor.

go deeper

for a junior

Know that allocation is the decision of how many units each stratum contributes, and that proportional allocation simply copies the population's shares into the sample.

for a middle

Be able to state both rules — proportional to N_h, versus proportional to N_h times S_h — and say which quantity you have to estimate before the survey runs.

for a senior

Show judgment about reportable segments, floors on small strata, and the operational cost of shipping a sample that is no longer self-weighting.

for a principal

Own the call between an allocation optimal for one metric and a near-proportional design that serves many metrics and many analysts without weighting mistakes.

## The allocation problem Stratification asks two separate questions. First, *what are the strata* — how do you partition the frame. Second, *how big is each stratum's sample* — allocation. The first is about which variables separate the outcome; the second is a pure optimisation, and it has clean answers. Write `N_h` for the population count in stratum `h`, `W_h = N_h / N` for its share, `S_h` for the standard deviation of the outcome *within* that stratum, and `n_h` for the number of units you sample there, with `sum of n_h = n` fixed. ## Proportional allocation Set `n_h = n * W_h`. Each stratum gets its population share of the sample. Two properties follow, and both matter operationally more than they look on paper. The design is **self-weighting**: every unit's inclusion probability is `n_h / N_h = n / N`, the same everywhere. So the plain unweighted sample mean is already the correct estimate of the population mean, and every analyst who touches the table afterwards gets the right answer without knowing the design existed. And it is **robust across outcomes**. A single survey usually measures many things — spend, satisfaction, feature usage. Proportional allocation is not optimal for any of them, but it is never badly wrong for any of them either. ## Neyman allocation Set `n_h` proportional to `N_h * S_h`. This is the allocation that minimises the variance of the stratified mean estimator subject to a fixed total `n`. The intuition is direct: sampling effort should go where two things coincide — a lot of population, and a lot of disagreement inside it. A stratum where everyone answers nearly the same number is cheap to pin down with a handful of interviews; a stratum whose answers span orders of magnitude needs many. A worked case. A customer base has 90,000 free accounts with a within-stratum spend standard deviation of about 5, and 10,000 enterprise accounts with a standard deviation of about 50. Total budget: 1,000 interviews. - Proportional: 900 free, 100 enterprise. - Neyman: the products `N_h * S_h` are `90,000 * 5 = 450,000` and `10,000 * 50 = 500,000`, summing to 950,000. So free gets `1000 * 450/950` which is about 474, and enterprise about 526. The enterprise tier is 10% of the population and receives more than half the sample. That is not a mistake — it is where the uncertainty about the population total lives. ## Optimal allocation with costs If a response costs `c_h` in stratum `h` and the constraint is a budget rather than a headcount, the variance-minimising rule becomes `n_h` proportional to `N_h * S_h / sqrt(c_h)`. Expensive-to-reach strata are sampled less than their variability alone would justify. Neyman allocation is exactly the special case where every stratum costs the same. ## What Neyman allocation costs you **You must guess `S_h` in advance.** The optimum depends on a quantity you are running the survey to learn. In practice you use a pilot, last year's survey, or an operational proxy such as the spread of billing records. A wrong guess costs precision, not correctness: as long as you use the true inclusion probabilities as weights, the estimator remains unbiased — you have simply spent interviews in the wrong places and landed somewhere between the optimum and proportional. **The sample is no longer self-weighting.** Inclusion probabilities now differ by stratum, so the plain average of the responses estimates the wrong thing, and every downstream calculation must apply weights. This is a real organisational cost, not a formality. **It is optimal for one variable at a time.** The `S_h` that minimises variance for spend is not the one that minimises it for satisfaction. Surveys measuring many outcomes usually settle for a compromise — an allocation near proportional, nudged toward the more variable strata. ## Domain estimates change the calculus Allocation rules above optimise the *overall* estimate. If you also need to report each stratum separately, precision inside a stratum depends on that stratum's own sample size, and a rule that hands a 200-account segment four interviews is useless. Practical designs therefore impose a floor: a minimum sample per reportable stratum, then allocate what is left by proportional or Neyman logic, then weight everything back together. ## What an interviewer is listening for That you can state both rules without hesitation; that you know Neyman needs `S_h`, not the stratum *mean*; that you catch the variance-versus-standard-deviation confusion (an allocation proportional to the *variance* is wrong — it is the standard deviation that enters); and that you volunteer the weighting obligation that a disproportionate allocation creates rather than being reminded of it.

  • What happens if your prior estimate of a stratum's variability is badly wrong?
    You lose precision, not correctness. As long as you weight by the actual inclusion probabilities, the estimator stays unbiased; you have merely spent interviews in the wrong places and end up somewhere between the Neyman optimum and proportional. That risk is why practical designs put a floor on each stratum's sample size and cap how far the allocation drifts from proportional.
  • How does an unequal cost per response change the optimal allocation?
    Under a fixed budget the variance-minimising rule becomes n_h proportional to N_h * S_h / sqrt(c_h), where c_h is the cost of one response in stratum h. Strata that are expensive to reach are sampled less than their variability alone would justify. Neyman allocation is the special case where every stratum costs the same.
  • Why might you over-sample a small stratum beyond what any allocation rule says?
    Because you want to report on it separately, not merely have it contribute to the total. Precision for a segment estimate depends on that segment's own sample size, so a 200-account tier may need a few hundred responses to be reportable even when the optimal rule for the grand mean would give it a handful. You then weight so the overall estimate stays unbiased.

Proportional allocation staffs each department by head count. Neyman allocation staffs by how unpredictable each department is: the one whose numbers swing wildly gets the extra auditors.

saying these in an interview costs you the question

  • Says Neyman allocation means equal sample sizes per stratum
  • Allocates on stratum size alone and ignores internal variability
  • Uses the stratum mean instead of its standard deviation
  • Makes the allocation proportional to the variance rather than the standard deviation
  • Forgets that a disproportionate allocation obliges weighting

context