Which users belong in the denominator of a checkout conversion metric?
answer
- who could plausibly be affected
- snapshot as of assignment
- nothing downstream of the treatment
- same filter, different populations
- narrower denominator, different estimand
basics
~20 sOnly users whose eligibility is settled by facts that the treatment cannot change, snapshotted at assignment. Narrowing to a subgroup such as users with a saved payment method is safe if that status was already true at assignment, and unsafe if the treatment can create it.
solid answer
~50 sThe denominator defines the population the effect is estimated over, so it must be determined by information that the treatment cannot influence — typically a snapshot taken at assignment time. Restricting a checkout metric to users who already had a payment method on file is legitimate when card-on-file status is read as of assignment: both arms then contain the same kind of people and randomisation still holds inside the subgroup. It becomes invalid the moment the treatment itself can change who qualifies — for example if the variant adds a save-this-card prompt, the treatment arm's denominator now contains users the control arm's never had, and any difference mixes the effect with the selection. The second consideration is interpretation: a narrower denominator is a different estimand. A 4% lift among card-on-file users is not comparable to a 4% lift among all users, and it does not translate into the same business impact. Pick the population deliberately, freeze the rule before launch, and report which population the number refers to.
go deeper
Be able to say that the denominator is the population the result is about, and that it should be decided from what was true before the user entered the experiment, not from what they did afterwards.
Explain mechanically how a post-treatment eligibility filter breaks comparability: the treatment changes who passes the filter, so the two arms end up holding different kinds of users even though the rule was identical.
Demonstrate the operating habits — snapshot eligibility at assignment, check per-arm denominator counts before reading any result, and refuse to convert a narrowed-population lift into a company-wide forecast.
Own the tradeoff between sensitivity and representativeness. Be ready to argue when a narrow eligible population is the honest estimand for a segment-targeted feature and when it quietly excludes the very users the feature was built for.
## What a denominator commits you to When you write `checkout conversions / eligible users`, the denominator is doing two jobs at once. It is the normaliser that makes arms comparable, and it is the definition of the population your result is about. Getting it wrong breaks either the causal claim or its interpretation, and the two failures look very different. ## The pre-assignment rule Randomisation buys you one guarantee: the two arms are exchangeable *as assigned*. Any filter applied to the denominator that uses information generated after assignment can destroy that guarantee, because the treatment may change who passes the filter. Concretely, suppose the metric is conversions divided by users who have a payment method on file. Two versions: - **Snapshot at assignment.** You record, for each user, whether a card was on file at the moment they entered the experiment, and you never update it. The filter uses a pre-treatment fact. Inside the filtered set, assignment is still random, so the comparison is still causal. It answers a narrower question — the effect among users who already had a card saved — but it answers it correctly. - **Evaluated at analysis time.** You filter on whether the user has a card on file when you build the table. If the variant nudges people to save a card, the treatment arm's denominator now contains a group of users the control arm's denominator never gets. You are comparing card-savers-under-treatment with card-savers-under-control, which are different populations. The difference you measure is part effect and part selection, and no amount of extra sample fixes it. The general principle is that a legitimate denominator filter is a function of pre-randomisation state only. Applying the same rule symmetrically to both arms is not enough: symmetric rules on post-treatment variables still select asymmetric populations, because the treatment changed the variable. ## Sensitivity versus estimand The reason teams narrow denominators is sensitivity. Users who could never plausibly convert — someone who bounced on the landing page, someone in a market where the feature is not shipped — add zeros to both arms. They contribute variance in the denominator and nothing to the signal, so removing them raises the measured effect size and can shorten the test. That gain is real, but it is not free. The narrowed metric estimates a different quantity. A lift of 4% among users with a card on file, who may be 20% of the base, is not a 4% lift on the business. When someone converts the experiment result into a forecast, the population the denominator described has to be carried along with the number, or the forecast is wrong by the ratio of the two populations. This is why a metric definition should always name its population in the metric name itself — `conversion per card-on-file user` rather than `conversion rate`. There is also a subtler asymmetry: narrowing to an eligible population can make an effect look larger while making it *less* representative of what launching would do, because the excluded users are exactly the ones the feature may have been meant to reach. If a feature's whole thesis is bringing in users who do not yet have a payment method, measuring only card-on-file users guarantees you will miss it. ## Practical rules 1. **Snapshot eligibility at assignment** and store it as an attribute of the assignment record, not as a join computed later. This makes the filter reproducible and auditable. 2. **Never filter on anything downstream of the treatment** — not on later behaviour, not on attributes the treatment can edit, not on whether the user was still around at the end. 3. **Name the population in the metric.** Two metrics with the same numerator and different denominators need two names, or someone will compare them. 4. **Freeze the rule before launch.** Changing eligibility mid-flight means the early days and later days are measuring different populations; if you must change it, recompute both arms from raw logs under one rule, or restart. 5. **Report the denominator counts per arm.** If the two arms' denominators differ by more than sampling noise, something is filtering asymmetrically and the result is suspect until explained. 6. **Decide whether the narrow or the broad population is the decision-relevant one.** For a feature aimed at a specific segment, the narrow one is often right; for a launch decision with a company-wide forecast attached, the broad one usually is. Reporting both, clearly labelled, is a reasonable default. ## What a good answer sounds like A strong candidate does not simply say "everyone who was randomised". They say: start from everyone randomised; narrow only on pre-assignment attributes; state the population in the metric's name; check the denominator counts per arm before believing anything; and be explicit that a narrower denominator answers a narrower question that cannot be reused as a company-level lift.
- Restricting the denominator to users with a saved payment method doubles the measured lift. Is that legitimate?It is legitimate only if card-on-file status was snapshotted at assignment and the treatment cannot create it. Even then, the bigger number is not a bigger effect on the business — it is an effect over a smaller population, and multiplying it by total users overstates impact. Report it as the effect among that segment, with the segment's size next to it.
- How do you keep the denominator stable if the eligibility rule has to change mid-experiment?Version the definition and recompute both arms from raw logs under a single rule, so no read mixes two populations. If the change alters who could have been assigned rather than just who is counted, the clean move is to restart the experiment. Never patch the rule forward only, which makes early and late data incomparable.
- Two arms show denominator counts differing by 3% on a 50/50 split. What do you check first?Treat it as a sample-ratio problem, not a result. Check assignment logging, bot and internal-traffic filters, whether one arm errors before the exposure event fires, and whether the eligibility attribute is being read at a different time in each arm. Until the imbalance is explained, the metric difference is not trustworthy in either direction.
saying these in an interview costs you the question
- Filters the denominator on behaviour that happened after assignment
- Argues a symmetric filter is safe even on post-treatment attributes
- Compares a narrowed-population lift against an all-user lift
- Assumes narrowing the denominator only removes noise
- Changes eligibility rules mid-flight and re-reads the same test