skip to content

Why does sampling whole classrooms instead of individual students cost you effective sample size?

level: seniorimportance: should knowfreq 40%

answer

  1. cheap to collect, expensive in information
  2. units inside a group are alike
  3. intraclass correlation between classmates
  4. 1 + (m - 1) * rho

basics

~20 s

Students in a classroom resemble one another, so each extra student adds less new information than an independently drawn one. That similarity inflates the variance by the design effect, roughly 1 + (m - 1) * rho for clusters of size m.

solid answer

~50 s

In a multistage cluster design you first sample schools, then classrooms inside them, then measure every student in the chosen classrooms. It is cheap — you visit 40 sites instead of chasing 1,000 scattered students — and often it is the only option, because you may have a list of schools but no list of students. The price is that observations inside a cluster are correlated: same teacher, same neighbourhood, same curriculum. The intraclass correlation `rho` measures that similarity, and with equal clusters of size `m` the design effect is roughly `Deff = 1 + (m - 1) * rho`, with effective sample size `n_eff = n / Deff`. At `m = 25` students per classroom and `rho = 0.1`, `Deff` is 3.4, so 1,000 measured students carry about the information of 294 independently drawn ones. The design lesson: more clusters, fewer units per cluster.

go deeper

for a junior

Know that cluster sampling selects whole groups, such as classrooms, and measures the units inside them, rather than picking individuals from the population directly.

for a middle

Be able to state the design effect for equal clusters, 1 + (m - 1) * rho, and convert a nominal sample size into an effective one.

for a senior

Demonstrate that you plan around the design effect up front and would refuse to report a precision claim computed as if clustered observations were independent.

for a principal

Own the budget argument: what the fieldwork saving per site is worth against the effective sample size lost, and what precision the decision at hand actually requires.

## Why anyone clusters in the first place Cluster sampling selects naturally occurring groups — schools, classrooms, city blocks, households, stores — and measures units inside the selected groups. **Multistage** sampling nests the idea: sample schools (primary sampling units), then classrooms inside the selected schools (secondary units), then measure every student in the selected classrooms, or a subsample of them. Two reasons drive it, and neither is statistical elegance. **Cost.** Measuring 1,000 students scattered across a country means 1,000 travel legs. Measuring 40 classrooms of 25 means 40 site visits. The fieldwork saving is enormous. **Frames.** You often cannot draw a simple random sample of students because no national list of students exists — but a list of schools does. Multistage designs let you build the frame one stage at a time: you only need to enumerate students inside the schools you actually selected. ## The cost: correlated observations Units inside a cluster are alike. Students in one classroom share a teacher, a curriculum, a catchment area and a peer group; households on one block share income level and infrastructure. The **intraclass correlation** `rho` quantifies this: it is the correlation between two units drawn from the same cluster, equivalently the share of total variance that sits *between* clusters rather than within them. `rho = 0` means cluster membership tells you nothing; `rho = 1` means every unit inside a cluster is identical, so the second student measured in a classroom adds literally nothing. This is exactly the opposite of what stratification exploits. Strata are built to be internally *homogeneous* and are *all* sampled, which helps. Clusters are internally homogeneous and only *some* are sampled, which hurts. The same variable — school — can play either role: sample every school and it is a stratum, sample 40 of 2,000 and it is a cluster. ## The design effect The **design effect** is the ratio of the variance your design actually produces to the variance a simple random sample of the same size would have produced: `Deff = Var(design) / Var(SRS, same n)`. For a cluster design with equal cluster size `m` and intraclass correlation `rho`, `Deff is approximately 1 + (m - 1) * rho` Read the formula. When `rho = 0`, `Deff = 1` and clustering is free. When `m = 1` — one unit per cluster, so you sampled units individually after all — `Deff = 1` again. Everywhere else `Deff` grows with *both* cluster size and correlation, and `m` enters multiplied by `rho`, which is why even a tiny `rho` becomes serious once clusters are large. The practical translation is **effective sample size**: `n_eff = n / Deff`, the number of independently drawn observations that would carry the same information. Work the classroom case. `m = 25`, `rho = 0.1`: `Deff = 1 + 24 * 0.1 = 3.4`. Measure 1,000 students and you hold roughly `1000 / 3.4`, about 294 independent students' worth of information. Two thirds of the fieldwork has bought you nothing in precision — and it may still have been the right call, because 1,000 clustered measurements can be cheaper than 294 scattered ones. ## Fewer per cluster, more clusters The formula immediately implies the design rule. Halve `m` from 25 to 10 at the same `rho = 0.1` and `Deff` falls from 3.4 to 1.9. With a fixed total of 400 students, 40 classrooms of 10 give a far more precise mean than 10 classrooms of 40. Precision is driven mainly by the **number of clusters**, because clusters are the units that vary independently. Against this pushes cost. Each additional site has a fixed cost — travel, permission, setup — while each additional student inside an already-visited classroom is nearly free. The optimal cluster size balances that cost ratio against `rho`: the higher the intraclass correlation and the cheaper the extra site, the smaller the clusters should be. ## Beyond the simple formula The `1 + (m - 1) * rho` expression assumes equal cluster sizes. Unequal cluster sizes inflate the design effect further, and so do unequal sampling weights — those contributions are usually folded into a single reported `Deff` for the survey. A design effect below 1 is possible and means the design beat simple random sampling: stratification typically produces `Deff` slightly under 1, and a well-ordered systematic draw can too. Cluster designs essentially always produce `Deff` above 1, because `rho` in real social, educational and geographic data is small but reliably positive. ## Where candidates lose the point The fatal move is to plan or report as if the 1,000 clustered students were 1,000 independent observations. Every precision statement, every sample-size calculation and every test built on that assumption is overconfident by a factor of `sqrt(Deff)` in the standard-error scale. The correct planning move is the reverse: decide the effective sample size the decision requires, guess `rho` from prior studies of similar populations, choose `m` from cost, and inflate the nominal sample size by `Deff`. Interviewers ask this question because it separates people who have designed field data collection from people who have only analysed clean rectangles of independent rows.

  • With a fixed budget, do you sample more clusters or more units per cluster?
    More clusters, nearly always. The design effect grows with cluster size m, and precision is driven by the number of independently varying clusters, so 40 classrooms of 10 beat 10 classrooms of 40 at the same total. The counterweight is cost: another site has a fixed price while another student in a visited classroom is nearly free, so the optimum balances that ratio against rho.
  • How do you tell a stratum from a cluster?
    You sample from every stratum; you sample only some clusters. Strata are constructed to be internally similar and different from one another, which raises precision. Clusters are naturally occurring groups taken whole for convenience, and their internal similarity lowers it. The same variable, such as school, can serve as either depending on whether you take all of them or a sample.
  • What does an intraclass correlation of zero imply for a cluster sample?
    That units inside a cluster are no more alike than units drawn at random, so the design effect is 1 and the clustered sample carries the same information as a simple random sample of equal size. Real education and household data rarely behave that way — rho there is small but reliably positive, which is enough to matter once clusters get large.

Interviewing 25 students from one classroom is closer to interviewing one classroom 25 times than to interviewing 25 students.

saying these in an interview costs you the question

  • Treats 1,000 clustered students as 1,000 independent observations
  • Confuses clusters with strata
  • Assumes the design effect is always greater than one
  • Buys precision by adding students inside already-visited classrooms
  • Ignores the design effect when planning sample size

context