skip to content

Twenty patients are each measured ten times and analysed as 200 independent observations - what breaks?

level: seniorimportance: must knowfreq 46%

answer

  1. ask what a single row really is
  2. effective sample size, not row count
  3. correlation inside each cluster
  4. standard error too small by a factor
  5. the CLT cannot repair dependence

basics

~20 s

The independence assumption breaks. Measurements from one patient are correlated, so 200 rows carry far less information than 200 independent ones: standard errors come out too small, confidence intervals too narrow and p-values far too optimistic.

solid answer

~50 s

Rows from the same patient are correlated, so the effective sample size is much closer to 20 than to 200. Any standard error computed as if there were 200 independent draws is understated by roughly the square root of the design effect: with ten measurements per patient and an intraclass correlation of 0.5 the design effect is `1 + 9 x 0.5 = 5.5`, the effective n is about 36, and the reported standard error is more than twice too small. The consequence is inflated false positives - a null effect rejects far more often than the nominal 5%. Normality is not the problem here, and the central limit theorem does not help, because dependence does not average away. The simplest fix is to collapse each patient to one summary value and analyse those 20 numbers; the alternative is a method that accounts for the within-patient correlation explicitly.

code

python · 14 lines
python
import random, statistics

def trial(n_patients=20, reps=10):
    xs = []
    for _ in range(n_patients):
        patient = random.gauss(0, 1)      # patient-level offset, mean zero
        xs += [patient + random.gauss(0, 1) for _ in range(reps)]
    n = len(xs)
    se = statistics.stdev(xs) / (n ** 0.5)   # pretends all 200 rows are independent
    return abs(statistics.mean(xs) / se) > 1.96

random.seed(0)
rejections = sum(trial() for _ in range(2000))
print(rejections / 2000)

go deeper

for a junior

Be able to spot the structure: if one identifier appears many times, those rows are not independent and the sample is smaller than the row count suggests.

for a middle

Explain the mechanism numerically - the design effect of 1 plus (m minus 1) times the intraclass correlation, the effective sample size it implies, and why the standard error is understated by its square root.

for a senior

Diagnose and repair it in a real analysis: identify the clustering from the data dictionary, state the direction of the error, choose between collapsing to patient-level summaries and modelling the correlation, and defend the choice.

for a principal

Set the guardrails. Decide how experiment platforms and analysis templates declare the unit of randomisation, so nobody ships a result whose confidence interval silently assumed independent rows.

## What independence means and where it fails Independence says that knowing one observation tells you nothing about another. It is a property of how the data were generated, not of the numbers on the page, and it is the assumption behind almost every standard error you have ever computed - the `s / sqrt(n)` formula counts every row as a fresh, independent piece of information. Repeated measurement is the most common way it fails. Ten readings from the same patient share everything about that patient: their baseline, their physiology, their measurement conditions. If a patient runs high, all ten of their rows run high. The same structure appears whenever a unit contributes multiple rows - several sessions per user, pupils inside a classroom, orders inside a household, repeated observations of the same store. Treating those rows as independent has a name in the experimental-design literature: pseudoreplication. ## Quantifying the damage The within-cluster similarity is summarised by the intraclass correlation, the proportion of total variance that sits between clusters rather than within them. Write it as `rho`. For equally sized clusters of `m` observations, the variance of the estimated mean is inflated by the design effect `DEFF = 1 + (m - 1) * rho` and the effective sample size is the row count divided by that factor. With `m = 10` and `rho = 0.5`, `DEFF = 5.5`, so 200 rows carry about as much information as 36 independent ones, and the standard error you should report is `sqrt(5.5)` - roughly 2.3 times - larger than the naive one. With a weaker `rho = 0.3`, `DEFF = 3.7` and the effective n is about 54: still nowhere near 200. Because the test statistic is an estimate divided by its standard error, dividing by something 2.3 times too small inflates the statistic by the same factor. A procedure advertising a 5% false-positive rate can easily produce false positives dozens of percent of the time. Notice the direction: positive within-cluster correlation makes the analysis anti-conservative, which is the dangerous direction - you find effects that are not there. ## Why more data makes it worse This is the counter-intuitive part worth saying out loud in an interview. Increasing the number of measurements per patient raises `m`, which raises the design effect, while the effective sample size creeps toward a ceiling set by the number of patients. So the row count grows, the reported standard error shrinks, and the honest standard error barely moves - the gap between them widens. Precision about the population comes from more patients, not more repeats. The repeats buy precision about each individual patient, which is a different and usually smaller-value quantity. Equally important: the central limit theorem does not rescue you. It concerns the shape of the sampling distribution of a mean built from independent draws. Dependence violates its premise, so no sample size fixes it. This is why independence outranks normality in any honest ranking of assumptions - normality problems shrink as data accumulate, dependence problems grow. ## Detecting it You detect this by reading the data, not by testing it. Ask what a single row represents and whether any identifier repeats across rows. A patient id, a user id, a device id, a session id appearing 200 times across 20 distinct values is the entire diagnosis. If you want a numeric feel, compute the variance of the per-patient means and compare it with the pooled within-patient variance: a large between-patient component is a large intraclass correlation. Sorting the residual-free raw values by patient and eyeballing the groups usually makes it obvious. ## Fixing it The simplest and most defensible fix in this scenario is to collapse: compute one summary value per patient - the mean of their ten readings - and analyse those 20 numbers as 20 independent observations. It looks like throwing away data, but the ten repeats mainly reduce measurement noise inside each patient, and averaging retains that benefit while leaving the between-patient variation that the standard error should be reflecting. The analysis is then transparently valid, at the honest cost of admitting n = 20. When the design is unbalanced, when repeats differ systematically, or when the within-patient structure is itself of interest, the alternative is a method that models the correlation between measurements on the same patient rather than assuming it away. Either route is defensible. What is not defensible is reporting an interval built on n = 200. ## How to answer Name the assumption, quantify the consequence with the design effect, state the direction of the error (too-narrow intervals, too-small p-values), point out that more repeats worsen it, and note explicitly that normality was never the problem here. That sequence is what an interviewer is listening for.

  • How would you estimate how badly the standard error is understated?
    Through the intraclass correlation. With equally sized clusters, the variance is inflated by a design effect of `1 + (m - 1) * rho`, where `m` is the measurements per patient. Divide the row count by that factor for the effective sample size, and take its square root to see how far off the standard error is. At `m = 10` and `rho = 0.3` the design effect is 3.7, so 200 rows are worth about 54 independent ones.
  • If you collapse to one value per patient, haven't you thrown away data?
    You have thrown away precision within a patient, not information about the comparison you care about. The ten repeats mostly reduce measurement noise inside each patient, and averaging keeps that benefit while leaving the between-patient variation, which is exactly what the standard error should reflect. What you lose is the ability to model within-patient structure, which is the reason to prefer a method that keeps both levels when the design is unbalanced.
  • Does the problem get better if you take more measurements per patient?
    No, it gets worse. Adding repeats to the same 20 patients raises the row count and the design effect together, while the effective sample size creeps toward a ceiling set by the number of patients. So the naive standard error keeps shrinking while the honest one barely moves, and the gap between the reported answer and the truth widens. Real precision comes from recruiting more patients.

saying these in an interview costs you the question

  • Says 200 rows means n = 200 regardless of who produced them
  • Blames non-normality when the real failure is dependence
  • Thinks the central limit theorem covers correlated observations
  • Proposes more repeats per subject as a way to gain power
  • Reports a narrow confidence interval without asking what a row is

context