You randomize by user but run the t-test over page-views — why is that p-value wrong?
answer
- two different units in play
- coin flips versus rows
- SE = s / sqrt(n), and n is a lie
- heavy users dominate a row-weighted mean
- collapse to one value per assigned unit
basics
~20 sThe test counts page-views as independent, but page-views from one person share that person's single assignment. The effective sample size is the number of users, so the standard error is understated and false positives run far above the nominal rate.
solid answer
~50 sThis is a unit-of-analysis mismatch: assignment happened once per user, but the comparison is computed over page-views. A two-sample t-test assumes the observations are independent draws; page-views from the same person are not, because they share one arm, one taste and one session pattern. Two things go wrong. First, n is wildly inflated — 40,000 users become 3 million page-views — and the standard error, which shrinks like 1 over the square root of n, is far too small. Second, the mean is now page-view-weighted, so a handful of very heavy users dominate the estimate. The result is confidence intervals that are too narrow and p-values that are too small; an A/A test analysed this way will flag "significant" differences far more often than 5% of the time. The simple fix is to make the analysis unit match the randomization unit: collapse each user to one number, then compare those user-level values across arms.
go deeper
Recall that the rows you compare should correspond to the things you randomized, and that a million events from forty thousand people is not a million independent observations.
Explain the mechanism: dependent rows break the sqrt(n) standard error, and a row-weighted mean silently re-weights the population toward the heaviest users.
Show you would catch this in a real pipeline — compare reported n against assigned users, run a null split to measure the actual false-positive rate, and state the aggregation you would use instead.
Own the platform answer: make the analysis unit derive automatically from the recorded assignment unit so this class of error cannot be authored by hand.
## The mismatch An experiment has two units that people confuse. The **unit of randomization** is what received a coin flip — here, the user. The **unit of analysis** is what each row of the comparison represents — here, a page-view. When the analysis unit is finer than the randomization unit, the analysis is claiming more independent evidence than the design produced. ## Why independence matters to the t-test A two-sample t-test compares two means using an estimate of how much a mean would bounce around under repeated sampling. For independent observations with standard deviation s, the standard error of the mean is SE = s / sqrt(n) Everything downstream — the t statistic, the confidence interval, the p-value — is built on that quantity, and that formula is only correct when the n observations are independent. Page-views from a single person are strongly dependent: they share the same assigned arm, the same interest in the product, the same device and session habits. Ten page-views from one person carry nowhere near ten observations' worth of information about how a user responds to the treatment. ## Two distinct errors **Inflated n.** Suppose 40,000 users produce 3 million page-views. Plugging n = 3,000,000 instead of n = 40,000 into the formula shrinks the standard error by a factor of about sqrt(75), roughly 8.7. Intervals that should be plus or minus one percentage point come out plus or minus a tenth of a point. Differences that are pure noise clear any significance threshold. **Weighting change.** The page-view-level mean is not the average user; it is the average page-view, which weights each user by how many page-views they generated. A small group of extremely heavy users can dominate. That is not automatically wrong — sometimes the page-view-weighted quantity is the business question — but it is a different estimand, and it is usually not what "the effect on a user" means. ## How to detect it The cheapest detector is a null run: split users into two arms with no treatment difference at all and analyse exactly as you plan to. If the analysis is honest, roughly 5% of such runs produce p < 0.05. Under a unit-of-analysis mismatch you will see 30%, 50%, or nearly always significant, depending on how many page-views each user contributes and how strongly users differ from one another. A second smell test: compare the reported n in the analysis against the number of assigned users. If they differ by an order of magnitude, ask why. ## The fix Make the analysis unit match the randomization unit. Reduce each user to a single value, then run the comparison over users: - **Rate per user.** Each user's own conversions divided by their own page-views, then compare the average of those per-user rates across arms. This weights every user equally. - **Binary per user.** Did this user convert at least once during the experiment? Yes or no, one value each. - **Count per user.** Total page-views, purchases or minutes per user. Often heavy-tailed, so trimming or capping extreme values, decided before looking at the results, keeps the comparison stable. With one row per user, the observations correspond to the coin flips, independence holds by design, and the usual standard error formula applies again. n is the number of users — which is also the number that should have driven the power calculation before the experiment started. ## The uncomfortable consequence Doing it correctly costs power: the honest n is far smaller than the row count, so the experiment needs more users or a longer run than the mismatched analysis suggested. That is not a loss, it is the removal of a false gain. The mismatched version was never detecting real effects; it was manufacturing significance out of within-user correlation. ## What interviewers listen for They want three things: the phrase that names the problem (the analysis unit is finer than the randomization unit), the mechanism (dependent observations break the standard error, and the mean is re-weighted toward heavy users), and a remedy stated concretely (aggregate to one value per user, or otherwise account for the dependence). Candidates who only say "the sample size is wrong" have half the answer; the weighting shift and the null-run diagnostic are what separate a solid answer from a strong one.
- How would you demonstrate to a sceptical stakeholder that the page-view-level analysis is broken?Run a null comparison: split users into two arms that receive identical experiences and analyse over page-views exactly as before. A correct procedure flags significance about 5% of the time; the mismatched one will flag it far more often. Repeating that split many times and showing the false-positive rate is far more persuasive than a formula, because it uses your own data and your own pipeline.
- If you aggregate to one rate per user, does every user count equally? Should they?Yes — averaging per-user rates gives each user weight one, regardless of activity. Whether that is right depends on the decision. Equal weighting answers "what happens to a typical user" and resists domination by a few heavy accounts. If the business cares about total volume, a volume-weighted quantity is the target instead, but then you must say so explicitly and handle the dependence rather than pretending the rows are independent.
- Does the mismatch bias the point estimate or only its uncertainty?Mainly the uncertainty: the standard error is far too small, so intervals and p-values are wrong. The point estimate is not biased for the page-view-weighted effect, but that is a different quantity from the per-user effect. If treatment changes how many page-views people generate, the weighting itself shifts between arms and the point estimate becomes hard to interpret too.
It is like polling ten households, asking every family member the same question, and then reporting a survey of forty independent voters.
saying these in an interview costs you the question
- Says more rows always means more power
- Confuses the assignment unit with the metric grain
- Reports n as the row count of the events table
- Believes a huge n makes any test trustworthy
- Thinks the only issue is heavy-tailed data