Why is treating three repeated blood-pressure readings per patient as three independent observations wrong?
answer
- same patient, not new information
- n counts units, not rows
- errors correlated within patient
- standard errors too small, p-values too small
- pseudo-replication
basics
~20 sReadings from one patient are correlated, so three readings carry far less information than three different patients. Counting them as independent inflates the sample size, shrinks the standard errors and p-values, and manufactures significance that is not there.
solid answer
~50 sRepeated measurements on one patient are positively correlated: knowing the first reading tells you a lot about the second, so three readings sit closer to one independent data point than to three. Ordinary regression assumes independent errors, so 300 readings from 100 patients make it behave as if it had 300 independent units and divide the variance by far too large an `n`. The coefficients themselves usually stay unbiased when the mean model is right, but the standard errors are too small, the confidence intervals too narrow and the p-values too small, so the real Type I error rate runs well above the nominal 5%. That is pseudo-replication: replicating the measurement is not replicating the unit. The honest fixes are to make the patient the unit — aggregate per patient, cluster the standard errors by patient, or fit a random intercept per patient.
go deeper
Be ready to say plainly that rows are not units: several readings from one patient do not count as several patients. Naming the within-patient correlation and its effect on the p-value is enough at this level.
Expect to explain the mechanics: the errors share a patient-specific offset, the variance formula divides by far too large an n, so the standard error is too small while the coefficient itself stays unbiased.
Show that you would catch this in the data, not in the write-up. Spot the repeating id column before modelling, decide the unit deliberately, and justify your choice among aggregation, clustered standard errors and a random intercept.
Own the standard: the analysis unit is a design decision, not a post-hoc repair, and it belongs in every review checklist your team uses. Be ready to explain to a stakeholder why a result with a tiny p-value had to be withdrawn once repeated readings were counted honestly.
## Rows are not units Every standard-error formula in ordinary least squares rests on one assumption that has nothing to do with normality or straight lines: the errors are independent of each other. Independence is what licenses the arithmetic `SE = s / sqrt(n)` and, more generally, the whole `(X'X)^-1 * sigma^2` variance formula. When a dataset has several rows per patient, per user, per store or per classroom, that assumption is usually false, and the sample size the software believes in is not the sample size you actually have. Call the readings on patient `j` at visits 1, 2 and 3 `y_j1`, `y_j2`, `y_j3`. Patients differ from each other in baseline blood pressure for reasons the model does not contain — age, weight, medication history, the anxiety of being in a clinic. Write that persistent patient-specific offset as `u_j` and the visit-to-visit noise as `e_jt`, so the error on any row is `u_j + e_jt`. Two rows from the same patient share `u_j`, so their errors are positively correlated; two rows from different patients share nothing and stay independent. The strength of the sharing is the intraclass correlation, `ICC = var(u) / (var(u) + var(e))`. ## What actually breaks Three things are worth separating, because interviewers probe exactly this distinction. **The coefficients are usually fine.** Under within-patient correlation, ordinary least squares remains unbiased for the coefficients as long as the mean model is correctly specified and the clustering is not itself confounded with the predictor. It is no longer the minimum-variance estimator — it is inefficient — but it is not systematically wrong. **The standard errors are not fine.** The reported variance assumes `n` independent errors. With positive within-patient correlation there is less independent information than `n`, so the reported standard error is too small. Confidence intervals are too narrow, t-statistics too large, p-values too small. A test you believe rejects 5% of the time under the null can easily reject 20% or 40% of the time. This is a false-positive machine, and it is silent: nothing in the output looks wrong. **The degrees of freedom are not fine either.** Software counts residual degrees of freedom from rows. With 300 rows from 100 patients, the reference distribution is far too generous. The severity is not uniform across predictors. A predictor that is constant within a patient — a patient-level treatment assignment, sex, a chronic diagnosis — suffers the worst inflation, because each patient really contributes just one observation of that predictor while the software counts three. A predictor that varies freely from visit to visit and is uncorrelated within the patient suffers much less. The damage grows with both the correlation of the errors within a cluster and the correlation of the predictor within a cluster. ## Recognising it before you model The diagnostic is structural, not graphical: look at the identifier columns. If a patient id, user id, device id or session id repeats down the rows, you have clustered data, and the burden of proof is on the analyst to show it does not matter. Counting distinct ids against total rows takes one line and tells you the worst case immediately: 300 rows and 100 patients means your true sample size is somewhere between 100 and 300, never above 300, and closer to 100 the more the patient dominates the variation. ## The three honest analyses **Aggregate to the unit.** Average each patient's readings into one number and analyse 100 rows. Simple, transparent, hard to argue with when cluster sizes are equal and you only care about patient-level predictors. It costs you any within-patient effect and some precision when cluster sizes are unequal. **Keep the rows, fix the inference.** Fit on all 300 rows but compute standard errors clustered by patient. This allows arbitrary correlation inside a patient while assuming independence across patients, and it changes only the uncertainty, never the estimates. **Model the structure.** Fit a random intercept per patient, which estimates `var(u)` explicitly, weights patients sensibly and recovers efficiency. It asks more of you: the random effect must be uncorrelated with the predictors, or the coefficients themselves become biased. ## The interview point The sentence that lands is that a p-value is only as trustworthy as the independence claim underneath it, and adding visits to the same patients buys measurement precision, not statistical power over patients. If you want more power over patients, recruit more patients.
- Does pseudo-replication bias the coefficients themselves, or only the standard errors?Usually only the inference. With a correctly specified mean model, ordinary least squares stays unbiased under within-cluster correlation; it is merely inefficient. What breaks is the variance estimate, so the standard errors, intervals and p-values are all too small. Bias in the coefficients appears only when the clustering is confounded with the predictor, for example when patients with more visits also differ systematically in treatment.
- For which kind of predictor is the understatement of the standard error worst?For a predictor that is constant within a cluster, such as a patient-level treatment or a demographic. Each patient then contributes essentially one observation of that predictor while the software counts every row, so the naive standard error can be several times too small. A predictor that moves from visit to visit and is uncorrelated within the patient suffers far less inflation.
- Would averaging each patient's readings into a single value be an acceptable analysis?Often yes, and it is the simplest defensible answer when every patient has a similar number of readings and the question is about patient-level predictors. You lose the ability to estimate anything that varies within a patient, and you throw away precision when cluster sizes are very unequal, since a mean of ten readings deserves more weight than a mean of two.
Asking one person the same question three times is not a three-person survey. You get a more precise reading of that one opinion, not three opinions.
saying these in an interview costs you the question
- More rows always mean more statistical power
- It is fine because every reading is a genuine measurement
- Independence only matters for the normality assumption
- Add more visits per patient to gain power over patients
- The p-value came out tiny, so the effect must be real