skip to content

If 20,000 reviews come from just 200 sellers, how many independent observations do you have?

level: seniorimportance: nice to knowfreq 28%

answer

  1. count clusters, not rows
  2. repetition adds less than novelty
  3. the design effect from survey sampling
  4. variance, not bias, is what goes wrong
  5. the interval widens by roughly sqrt(deff)

basics

~20 s

Far fewer than 20,000. Reviews from one seller are correlated, so they repeat information rather than adding it. With a design effect near 20 the sample behaves like about a thousand independent rows, so every error bar widens.

solid answer

~50 s

Rows clustered inside 200 sellers are not 20,000 independent draws. Reviews by one seller share product category, writing style and rating habits, so another review from a seller you already have adds far less than a review from a fresh seller. Survey sampling quantifies this with the design effect: `deff = 1 + (m - 1) * ICC`, where `m` is average cluster size and `ICC` is the share of outcome variance lying between clusters. With `m = 100` and an ICC of 0.2, `deff` is about 21, so the effective sample size is close to a thousand. The consequence is **uncertainty, not bias**: the interval around the held-out score is several times wider than a `1/sqrt(n)` calculation suggests, so a 0.3-point gap between two models is noise. And if the model will serve *new* sellers, the unit being sampled is sellers — you have 200.

go deeper

for a junior

Recall that repeated rows from the same source count for less than the row total suggests, and that a big row count is not automatically a big sample. Naming the intuition is enough here.

for a middle

Explain the mechanism: correlated rows within a cluster repeat information, so the standard error falls more slowly than one over the square root of the row count. Be able to sketch the design-effect correction.

for a senior

Demonstrate that you check what a single draw from the deployment population is before trusting any score, and that your intervals and model comparisons are computed over clusters rather than rows.

for a principal

Own the decision about the evaluation unit and what evidence a launch requires when the real sample is a few hundred entities. Be ready to argue for collecting breadth over volume when the two compete for the same budget.

## The independence half of i.i.d., in a concrete case The i.i.d. assumption has two halves, and discussions usually concentrate on the "identically distributed" one. This question is about the other half. Take 20,000 marketplace reviews written by 200 sellers, averaging 100 reviews per seller. Nothing here has drifted — the population is stable and the labelling rule is fixed. The rows are simply not independent draws. Why not? Reviews under one seller share the seller's catalogue, price band, shipping reliability, customer demographic and the phrasing conventions of that seller's buyers. If you know how twenty of a seller's reviews went, you can predict the twenty-first far better than chance. That predictability is exactly what independence rules out. ## Effective sample size and the design effect The standard error of a mean over `n` independent observations shrinks like `1 / sqrt(n)`. Under clustering it shrinks more slowly, and the survey-sampling literature gives the correction directly: ``` deff = 1 + (m - 1) * ICC n_eff = n / deff ``` - `m` is the average number of rows per cluster (here 100). - `ICC`, the intraclass correlation, is the share of total outcome variance that sits **between** clusters rather than within them. `ICC = 0` means seller identity tells you nothing and the rows really are independent; `ICC = 1` means every review from a seller is a copy and you effectively have 200 rows. With `m = 100` and `ICC = 0.2`: `deff = 1 + 99 * 0.2 = 20.8`, so `n_eff` is about 960. Twenty thousand rows carry the statistical weight of roughly a thousand. Two things follow immediately. The width of a confidence interval scales like `1 / sqrt(n_eff)`, so intervals are about `sqrt(20.8)`, roughly four and a half times wider than the naive computation. And any comparison between two candidate models has to clear that widened bar before it means anything. Note what is **not** claimed: clustering by itself does not bias the point estimate. If the held-out sellers are a fair sample of the sellers you will serve, the average loss over their reviews still centres on the right value. What is wrong is the confidence you attach to it. ## What the population of interest actually is The subtler question is what unit you are sampling. If the model will score new reviews for sellers already in the file, reviews are a reasonable unit and you have plenty of them. If it will score sellers it has never seen, then the thing being sampled is **sellers**, and 200 is your sample size no matter how many reviews each one contributed. A model can look excellent by memorising seller-level idiosyncrasy — this seller's buyers write short, generous reviews — and that skill transfers to exactly zero new sellers. This is the question to ask before the arithmetic: *what does one draw from my deployment population look like?* Answer that and the effective sample size usually becomes obvious. ## Putting an honest error bar on the score The practical habit is to make every uncertainty calculation operate on clusters rather than rows. When you resample to get an interval around the held-out score, resample **whole sellers** with replacement and recompute the score on all their reviews, rather than resampling individual reviews — resampling rows silently re-imposes the independence you do not have and reproduces the too-tight interval. The same logic applies to any significance claim about a model comparison. A rough sanity check needs no formal machinery: compute the score separately for each held-out seller and look at the spread across sellers. If per-seller scores range from 0.6 to 0.9, the pooled number is not accurate to three decimal places, whatever the row count says. ## What this leaf does not cover Keeping a given seller's rows entirely on one side of a split is a **splitting-protocol** question and is handled separately. It is worth being clear that the two problems are different: a clean by-seller split removes the contamination of scoring on sellers the model has memorised, but it does **not** create more independent information. After the cleanest possible split you still have 200 clusters, and the error bar is still the wide one. Candidates frequently answer the sample-size question with the split answer and stop there; the interviewer is usually listening for the recognition that the split fixes optimism while the effective sample size governs precision.

  • Does clustering bias the held-out score, or only its uncertainty?
    Mainly the uncertainty. If the held-out sellers are a fair sample of the sellers you will serve, the average loss over their reviews still centres on the right value; what is wrong is the too-narrow interval around it. Bias enters separately, if the seller mix on the held-out side is unrepresentative or if the model was allowed to memorise those sellers.
  • How would you put an honest confidence interval on a score computed over clustered rows?
    Resample whole clusters rather than individual rows: draw sellers with replacement, recompute the score over all of each drawn seller's reviews, and read the interval off the resulting spread. Resampling rows re-imposes the independence you do not have. A cheap sanity check is to score each held-out seller separately and inspect the spread across sellers.
  • When are 20,000 clustered reviews genuinely worth close to 20,000 rows?
    When the intraclass correlation is near zero — seller identity explains almost none of the outcome variance — and when deployment means scoring more reviews for sellers already represented. Then each review is close to a fresh draw from the relevant population and the design effect is close to one.

Interviewing one person a hundred times is not a hundred-person survey. You learn a great deal about that person and almost nothing new about the population.

saying these in an interview costs you the question

  • Quotes 20,000 as the sample size without qualification
  • Says grouping the split fixes the sample-size problem
  • Thinks clustering biases the point estimate rather than the interval
  • Bootstraps individual rows to get an interval on clustered data
  • Treats a 0.3-point model difference on clustered data as real

context