skip to content

Average order value is measured per order but users were randomized: how do you get a valid standard error?

level: seniorimportance: should knowfreq 46%

answer

  1. one row per user, not per order
  2. repeat buyers break independence
  3. collapse to per-user revenue and count
  4. delta method or resample the users
  5. point estimate unchanged, interval widens

basics

~20 s

Aggregate each user into revenue and order-count totals, then compute uncertainty across users: either the delta method on those two per-user means, or a bootstrap that resamples users and carries all their orders. Never treat orders as independent rows.

solid answer

~50 s

The metric stays `total revenue / total orders`, but the variance has to be computed at the unit that was randomized. Collapse the order table to one row per user with `x_i` = that user's revenue and `y_i` = that user's order count, including users with zero orders if they are in the eligible population. Then either apply the delta method to the two per-user means, which gives a closed-form standard error, or bootstrap by resampling users with replacement and recomputing the ratio of totals from whichever users were drawn. Both give roughly the same answer; the delta method is cheap and scales, the resampling approach is easier to defend when per-user totals are wildly skewed. What you must not do is run a two-sample test over the pooled order table — a repeat buyer's orders are correlated, and that test will report a standard error far smaller than the truth.

go deeper

for a junior

Recall that orders from the same customer are not independent observations, so a test run over the raw order table is not valid. Knowing to aggregate to one row per user first is the expected answer here.

for a middle

Explain the aggregation concretely — per-user revenue and per-user order count — and name both routes to the standard error, the delta method and resampling users. Be clear that the metric definition itself does not change.

for a senior

Show operational judgment: how you handle zero-order users, what you do when the treatment moves order counts, and how you sanity-check the result with A/A coverage. Expect to be pushed on why the interval got wider.

for a principal

Take a position on how the platform should treat revenue ratio metrics by default, including whether per-user revenue is the better primary metric for launch decisions and what to tell teams whose past wins do not survive an honest interval.

## Why the obvious analysis is wrong The order table is tempting: one row per order, a value column, an arm column, run a two-sample comparison of means. That analysis answers a question about *orders drawn independently*, and orders were not drawn independently. A customer who places eleven orders contributes eleven rows that share their price sensitivity, basket habits and the arm they were assigned to. The nominal sample size is the order count, the real one is the customer count, and if repeat buyers dominate the two can differ by an order of magnitude. The result is a standard error that is too small, intervals that are too narrow, and a false-positive rate above nominal. ## Step one: fix the unit Aggregate to the randomization unit before doing anything statistical. For each user i in the eligible population: - `x_i` = total revenue from that user in the experiment window - `y_i` = number of orders from that user in the same window The metric is unchanged: ``` AOV = sum(x_i) / sum(y_i) ``` One decision to make explicitly: users with `y_i = 0` contribute (0, 0) and correctly add sampling noise to the denominator without changing the point estimate. Dropping them is defensible only if the eligible population is genuinely "users who ordered", and that has to be a stated definition rather than an accident of a join. ## Step two: two ways to get the standard error **Delta method.** Treat AOV as the ratio of the two per-user means and use the first-order variance approximation, which combines the per-user variance of revenue, the per-user variance of order count and their covariance, divided by the number of users. It is a closed-form expression over per-user sufficient statistics, so it costs one pass over the aggregated table and can be computed for every metric in a catalogue on a schedule. It assumes a denominator mean comfortably above zero and enough users for the linearisation to hold. **Resampling users.** Draw users with replacement to form a sample the same size as the original, take *all* orders belonging to each drawn user, recompute the ratio of totals, and repeat. The spread of the recomputed ratios estimates the standard error. The critical detail is the resampling unit: resample users, never orders. Resampling orders reconstructs precisely the independence assumption you set out to avoid, so it reproduces the too-small standard error with extra compute. Both approaches converge to similar numbers when the assumptions hold, which is itself a useful cross-check: if the closed-form and resampled standard errors disagree materially, something about the aggregation or the skew deserves attention. ## Step three: compare the arms Compute the standard error separately in treatment and control, then combine. Because the arms hold disjoint sets of independently randomized users, the variances add, and the standard error of the difference is the square root of the sum. Report the absolute difference or the relative lift with an interval built from that standard error. ## What to watch for in a real analysis - **Sensitivity to a handful of users.** Per-user revenue is skewed; a small number of customers can move the point estimate. Reporting how much the estimate shifts when the largest contributors are excluded is an honest robustness note, distinct from the variance calculation. - **Order counts differing across arms.** If the treatment changes how many orders users place, the denominator itself is affected by treatment, and the ratio blends two effects. That is a real interpretive hazard: a shift in AOV may reflect a change in ordering behaviour, not in basket size. Say so rather than declaring a basket-size win. - **Standard errors that shrank after the fix.** They should grow. A user-level standard error smaller than the per-order one is a sign of a bug in the aggregation. - **Under-powered results.** Once the standard error is honest, the experiment may simply not be conclusive. That is information, not failure, and it feeds the sample-size planning for the next run. ## How to say it in an interview Name the mismatch in one sentence, name the aggregation step, then name both variance routes and their tradeoff. The strongest signal is stating that the point estimate is unchanged while only the uncertainty was wrong — it shows you know this is a variance problem, not a bias problem.

  • Should users who placed no orders be included in the aggregation?
    It depends on the declared eligible population. If the metric is defined over all exposed users, keep them as a (0, 0) pair: they leave the point estimate alone and correctly contribute uncertainty to the denominator. If the metric is explicitly about users who ordered, exclude them — but as a stated definition, not because a join silently dropped them.
  • What if the treatment changes how many orders users place?
    Then the denominator is itself affected by the treatment, and a movement in average order value mixes a basket-size effect with a change in order frequency. The standard error is still computed the same way, but the interpretation needs care: report the order-count change alongside the ratio, and avoid calling the result a pure basket-size win.
  • How would you check that your new standard error is actually correct?
    Run A/A analyses on historical data: split control users at random, compute the metric and its interval, and repeat many times. About 95% of the intervals should cover zero difference. If the coverage is materially below that, the standard error is still too small; if far above, it is too conservative.

saying these in an interview costs you the question

  • Runs a two-sample test over the pooled order table
  • Resamples orders instead of users when bootstrapping
  • Believes the point estimate needs correcting too
  • Silently drops users with zero orders
  • Expects the corrected standard error to get smaller

context