skip to content

In a store-level A/B test, why is your sample size the number of stores, not shoppers?

level: seniorimportance: should knowfreq 42%

answer

  1. count the coin flips, not the people
  2. everyone in a store shares its conditions
  3. extra footfall adds little information
  4. power lives in between-store variance
  5. match on pre-period sales before randomizing

basics

~20 s

Every shopper in a store receives the same assignment and the same local conditions, so their outcomes move together. The independent replicates are the stores, and adding shoppers inside a store adds far less information than adding another store.

solid answer

~50 s

When the assignment unit is a whole store — because the change is a shelf layout or a store-wide price that cannot be delivered to one shopper at a time, or because shoppers in the same store influence each other — the coin was flipped once per store. All the shoppers under that flip share not only the treatment but also the store's staff, catchment area, weather and local promotions, so their outcomes are correlated. Statistically that means 60 stores with 5,000 shoppers each is much closer to 60 observations than to 300,000. Power comes from the number of stores and from how much stores differ from one another, not from footfall. The practical consequences: analyse at the store level, expect a small n and therefore a large minimum detectable effect, and reduce between-store variance by pairing or blocking stores on pre-period sales before randomizing. Reporting a shopper-level test on a store-level design produces spectacular and entirely fake significance.

go deeper

for a junior

Remember the basic rule: the sample size is the number of things that were randomly assigned, so a design that assigns whole stores has as many observations as it has stores.

for a middle

Explain why shoppers inside a store are correlated — shared treatment, staff, catchment and local conditions — and why that makes extra footfall nearly worthless for precision.

for a senior

Show how you would run it: metric aggregated per store, power sized on store count and the spread of store means, matched-pair or blocked randomization on pre-period sales.

for a principal

Own the expectation-setting: decide when a small, low-powered cluster experiment is worth running at all versus buying more units, and what evidence you accept when it is not.

## Why a whole cluster becomes the unit Some treatments cannot be delivered to an individual. A rearranged shelf, a store-wide price, a new curriculum in a classroom, a city-wide delivery fee — the intervention exists at the level of the group. Sometimes the reason is behavioural instead: members of the group talk to each other or share a resource, so treating one member effectively touches the others. In both cases the honest design flips one coin per store, per classroom, per city, and every member inherits that arm. ## What the sample size becomes The rule is simple and often surprising to people used to web experiments: **the number of independent replicates equals the number of units that were randomized.** Sixty stores means sixty draws, however many shoppers walk through them. The reason is correlation within the cluster. Two shoppers in the same store share the arm, but also everything else about that store: the manager's execution of the change, staffing levels, the neighbourhood's income, local weather, a competitor opening down the road. Their outcomes therefore rise and fall together. Averaging 5,000 correlated shoppers gives a very precise estimate of *that store's* mean, but the quantity that limits the experiment is how much store means vary from store to store — and that variation is not reduced at all by measuring more shoppers inside each store. Mechanically: the store mean has an irreducible store-specific component that no amount of within-store sampling removes. Once you have enough shoppers to pin down each store's mean, additional shoppers buy essentially nothing. Additional *stores* buy a full new observation. ## The practical consequences **Small n, large minimum detectable effect.** With 40 or 60 stores, only fairly large effects are detectable. Teams routinely arrive with a plan sized on shopper counts and have to be told the experiment is powered for a 6% lift, not a 0.5% one. **Analyse at the store level.** Compute one number per store — sales per visit, conversion rate, revenue per opening hour — and compare those numbers across arms. That keeps the analysis unit equal to the randomization unit. **Variance reduction is where the power is.** Since n is fixed by the world, the lever is the variance between store means. Two standard moves: use each store's pre-period value as a covariate or as a baseline the outcome is measured against, and randomize within matched pairs or blocks of similar stores (by size, format, region) so the arms are comparable before the change starts. Blocking on a strong predictor of the outcome can cut the required number of stores substantially. **Beware a few dominant clusters.** Store sizes are usually very unequal. One flagship store can swing a store-weighted average; a shopper-weighted average lets it swing even harder. Decide before the run whether the estimand is per store or per shopper, and if it is per shopper, remember the uncertainty still has to respect that only the stores were randomized. **Randomization checks matter more.** With 60 units, chance imbalance on a variable like store size is common, not rare. Checking pre-period balance and, better, randomizing within matched blocks avoids arms that differ before anything happened. ## The failure mode to name The classic error is running a shopper-level significance test on a store-level design. It combines two mistakes — the analysis unit is finer than the randomization unit, and shoppers within a store are correlated — and it produces standard errors that are wrong by a large factor. The result is a confident report that a shelf change lifted sales, from an experiment that could not have detected anything smaller than several percent. ## What a strong answer sounds like "We randomized stores, so n is stores. Shoppers inside a store share the treatment and the store's own conditions, so their outcomes are correlated and extra footfall adds little. I'd analyse one metric per store, size the experiment on the number of stores and the spread of store means, and use pre-period sales for matched-pair randomization to get the detectable effect down."

  • With only 50 stores available, how do you get the detectable effect down?
    Attack the variance rather than the count. Use a pre-period window to measure each store's baseline, randomize within matched pairs or blocks of similar stores, and analyse the change relative to baseline instead of the raw level. Running longer also helps, because a longer window averages away week-to-week noise inside each store, though it never adds independent stores.
  • Your stores differ enormously in footfall. Does that change the analysis?
    It forces a choice of estimand. An unweighted average of store-level metrics answers "what happens to a typical store"; a footfall-weighted version answers "what happens to a typical shopper". Decide before the run, and either way the uncertainty must be driven by variation across stores. Very unequal sizes also make chance imbalance likelier, which is another argument for blocking on size.

Weighing sixty sacks of grain tells you about sixty sacks; scooping a thousand handfuls out of each one does not give you a thousand sacks.

saying these in an interview costs you the question

  • Quotes footfall as the experiment's sample size
  • Runs a shopper-level test on a store-level split
  • Assumes more weeks of traffic fixes too few stores
  • Ignores that stores differ before the change starts
  • Treats one flagship store as just another observation

context