skip to content

Metrics & Guardrails

Choosing the one metric a launch decision hangs on, the guardrails that can veto it, and the driver metrics that explain movement. A bad metric invalidates an otherwise perfect experiment.

on this pageshow

explore

questions

30

In an A/B test, what is the difference between a count, a rate, and a share metric?

level: juniorimportance: must knowfreq 72%

answer

  1. sum, normalise, or compose
  2. what sits under the line
  3. denominator fixed by assignment or by behaviour
  4. shares within a partition sum to 100%
  5. a rate needs a per-what

basics

~20 s

A count metric sums events, such as total orders. A rate metric divides a count by the units that were exposed, such as orders per visitor. A share metric divides one count by a related total, such as the percentage of orders placed on mobile.

solid answer

~50 s

A count is a raw sum of events in an arm: total orders, total add-to-carts. It scales with how many units landed in that arm, so it is never directly comparable between arms unless the exposure is identical. A rate normalises that count by the exposed population: `orders / exposed visitors`. The denominator is fixed by the assignment, so the two arms are comparable by construction. A share is compositional: `mobile orders / all orders`. Its denominator is itself an outcome, and the shares across categories are forced to sum to 100%, so a share can move because a different category moved. The practical rule is that the denominator is part of the metric name. "Conversion rate = 12%" is not a metric until you say 12% of what — visitors, sessions, or eligible users are three different numbers with three different interpretations.

go deeper

for a junior

Be ready to name the three shapes and give one concrete example of each, and to say out loud that a rate is meaningless until you state its denominator. Practise turning a vague brief like measure conversion into a full definition.

for a middle

Explain why a rate normalised by the randomised unit is comparable across arms while a count and a share are not, and describe how a de-duplication rule changes what a count metric actually measures.

for a senior

Show that you write metric definitions others can implement: numerator event, eligible denominator, window, and counting rule, all fixed before launch so two analysts produce the same number from the same logs.

for a principal

Own the question of which shapes are allowed to drive decisions at all. Argue when a share belongs on a review deck as context rather than as a result, and what the cost is when teams publish shares without their underlying counts.

## The three shapes Almost every experiment metric is one of three shapes, and confusing them is one of the most common analysis errors. **Count metrics** are raw sums over an arm: total orders, total searches, total support tickets. They are the easiest to instrument and the least useful to compare. A count answers "how much happened here", and how much happened depends on how many units were sent here. If the split is 50/50 and both arms logged the same number of exposures, a count difference is interpretable; if the split is 90/10, if one arm had a logging outage, or if the ramp changed mid-experiment, the count difference is mostly an artefact of exposure. Counts are still worth reporting as a sanity check and as a business-size figure, but the decision should not hang on them. **Rate metrics** (also called per-unit or normalised metrics) divide a count by the number of units that were exposed: `orders / exposed visitors`, `searches / exposed user`, `revenue / exposed user`. The essential property is that the denominator is fixed by the randomisation: every assigned unit contributes exactly one to the denominator, whatever it then does. That makes the two arms comparable by construction and makes the metric an average over a well-defined population. Two sub-shapes live here. A **proportion** has a binary numerator per unit (did this visitor convert at all: 0 or 1) and is bounded in [0, 1]. A **mean** has an unbounded numerator per unit (how many orders, how much revenue) and can be skewed. Both share the same virtue: one row per randomised unit. **Share metrics** (composition or mix metrics) divide one count by a related total rather than by the exposed population: percentage of orders placed on mobile, percentage of sessions that used search, percentage of revenue from subscriptions. The denominator here is another behavioural outcome, not a fixed population. Two consequences follow. First, the metric moves when the numerator moves, when the denominator moves, or when both move — a rise in mobile share is fully consistent with mobile orders being flat and desktop orders falling. Second, shares within a partition are constrained to sum to 100%, so they cannot all improve; one category's gain is arithmetically another's loss. Shares are excellent descriptive and diagnostic metrics and poor decision metrics on their own. ## Why the denominator is part of the definition The word "rate" carries no information about the denominator. "Conversion rate" can mean conversions per visitor, per session, per eligible user, per search, or per product-page view, and these are genuinely different quantities that can move in different directions in the same experiment. A metric definition is only complete when it states four things: 1. **The numerator event** — which logged event counts, and under what conditions. 2. **The denominator population** — which units are in scope, and as of when. 3. **The observation window** — over what period both are accumulated. 4. **The counting rule** — whether repeats count, and at what granularity they are de-duplicated. The fourth point is where a lot of silent damage happens. If a user clicks "add to cart" three times because the button did not respond, a raw event count records three; a per-user-per-session de-duplicated count records one. A treatment that only changes button responsiveness will move the first metric a long way and the second not at all. Neither counting rule is wrong; what is wrong is not deciding, or deciding differently between arms or between reads. De-duplication is a definitional choice with a rationale — count occurrences when intensity is the thing you care about, count distinct units when adoption is the thing you care about — and it belongs written down next to the metric, not buried in a query. ## Choosing between them Start from the question. If the question is "does this change help a user do the thing", you want a rate whose denominator is the randomised unit, because randomisation guarantees the denominators are comparable. If the question is "how big is this in absolute business terms", report a count alongside, and expect to scale it by exposure. If the question is genuinely about composition — which channel, which device, which plan — a share is the right shape, but report the underlying counts next to it so a reader can see whether the share moved because the part grew or the whole shrank. A useful discipline: for every headline share you publish, publish the numerator and the denominator as per-unit rates too. It costs one extra column and removes an entire class of misreadings.

  • Why is a raw count of orders a poor metric to compare arms when the traffic split is uneven?
    A count scales with the number of units assigned, so a 90/10 split produces roughly a nine-fold count difference with no behavioural change at all. Logging gaps, ramp changes and bot filtering shift exposure too. Normalising by exposed units removes that dependency, which is why the decision metric is a rate and the count is reported only as a business-size figure.
  • When you count add-to-cart events, should repeats inside one session be de-duplicated?
    It depends on the question, and the choice must be written into the definition. De-duplicate to one per user per session when you are measuring adoption — did they do it at all. Keep raw occurrences when intensity matters, for example basket-building behaviour. The hazard of raw counts is that retries and double-taps inflate them, so a treatment that only changes button responsiveness moves the metric without changing real behaviour.
  • Is revenue per exposed user a rate or a share?
    A rate — specifically a mean, because the denominator is the exposed population fixed by assignment and the numerator is unbounded per user. It is not a share, because nothing constrains it to a total. It differs from a proportion in that it is heavily skewed by a small number of large spenders, so its uncertainty behaves differently from a conversion proportion.

A count is the total distance a car travelled, a rate is its fuel economy per litre, and a share is the fraction of that distance driven on motorways. Only the middle one is comparable between two cars driven for different lengths of time.

saying these in an interview costs you the question

  • Says conversion rate without naming the denominator
  • Compares raw event counts between arms of different sizes
  • Assumes a rising share means the numerator grew
  • Treats de-duplication as a query detail, not part of the definition
  • Calls any x-per-y number a rate, including shares

context

open as a page

What is a guardrail metric in an A/B test, and how does it differ from a driver metric?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A guardrail is a metric the experiment must not damage — latency, crash-free sessions, unsubscribes — and a big enough degradation can veto the launch. A driver metric only explains why the primary metric moved; it never vetoes.

open as a page

Why does revenue per user make an A/B test hard to call when a few users spend far more than the rest?

level: juniorimportance: must knowfreq 74%

basics

~10 s

Almost all the variance in revenue per user comes from a handful of very large spenders, so the interval around the treatment-control difference stays wide. Ordinary experiment sizes then cannot resolve realistic revenue effects.

open as a page

What is an Overall Evaluation Criterion (OEC) in an A/B test?

level: juniorimportance: must knowfreq 74%

basics

~20 s

The Overall Evaluation Criterion is the single metric, agreed before launch, that an experiment's ship-or-not decision is defined against. It expresses what the team means by the product getting better, so the result cannot be reinterpreted afterwards.

open as a page

What does a DAU/MAU stickiness ratio of 0.2 tell you about how a product is used?

level: juniorimportance: must knowfreq 68%

basics

~10 s

A DAU/MAU of 0.2 means the typical monthly active user opens the product on about 6 days out of 30. It measures return frequency, not audience size, and averages over very different user types.

open as a page

What does capping revenue per user at the 99th percentile do to the effect you are estimating in an A/B test?

level: middleimportance: must knowfreq 64%

basics

~20 s

Capping replaces values above the 99th percentile with that value, so you estimate mean capped spend, not mean revenue. Variance falls sharply and power rises, but a real effect in the tail is shrunk or erased.

open as a page

When choosing an A/B test OEC, how do you trade sensitivity against long-term alignment?

level: middleimportance: must knowfreq 63%

basics

~20 s

Pick the most aligned metric that can still detect a realistic effect with the traffic you have. Sensitive short-term metrics decide fast but drift from what the business wants; truer long-run metrics are often too noisy or too slow to see.

open as a page

Why does a per-impression standard error for click-through rate understate uncertainty when users are randomized?

level: middleimportance: must knowfreq 72%

basics

~20 s

A per-impression standard error treats every impression as an independent draw, but impressions from one user are correlated and a few heavy users supply most of them. The effective sample size is users, not impressions.

open as a page

Why do N-day, rolling and bracket retention give three different numbers for one cohort?

level: middleimportance: must knowfreq 74%

basics

~20 s

N-day retention counts users active on exactly day N, rolling counts users active on day N or any later day, and bracket counts users active at least once inside a day range. Each rule admits a different set, so the numbers differ.

open as a page

Conversion per visitor rose but conversion per session fell in the same A/B test — why?

level: seniorimportance: must knowfreq 54%

basics

~20 s

The treatment almost certainly increased sessions per visitor. Sessions are an outcome, not a fixed denominator, so extra sessions dilute the per-session rate even while more visitors convert. Report the per-visitor number as primary and sessions per visitor beside it.

open as a page

Your treatment lifts sign-ups but degrades p95 page latency by 20ms — how do you decide whether to ship?

level: seniorimportance: must knowfreq 56%

basics

~10 s

Confirm the 20ms is real and not an instrumentation artefact, compare it against the pre-agreed latency budget, segment it by device, then ship, fix, or escalate. Never move the threshold to fit the result.

open as a page

What changes when an active-user metric uses a 28-day activity window instead of 7?

level: middleimportance: should knowfreq 41%

basics

~20 s

A 28-day window counts anyone with a single qualifying action in the past four weeks, so it reads higher, moves slowly and hides recent declines. A 7-day window reacts faster, is noisier, and is far more exposed to day-of-week effects.

open as a page

Which users belong in the denominator of a checkout conversion metric?

level: middleimportance: should knowfreq 56%

basics

~20 s

Only users whose eligibility is settled by facts that the treatment cannot change, snapshotted at assignment. Narrowing to a subgroup such as users with a saved payment method is safe if that status was already true at assignment, and unsafe if the treatment can create it.

open as a page

What counter-metrics would you track for a test that increases ad density on a page?

level: middleimportance: should knowfreq 50%

basics

~10 s

Track what a revenue win costs users: unsubscribe and churn rate, ad-blocker installs, return-visit frequency, complaints, and page latency. Short-run ad revenue almost always rises, so revenue alone settles nothing.

open as a page

Why isn't a non-significant guardrail p-value enough to conclude an A/B test did no harm?

level: middleimportance: should knowfreq 46%

basics

~20 s

A non-significant result means the data does not rule out zero difference; it never proves the difference is zero. Claiming no harm needs a stated degradation margin and a guardrail interval sitting entirely inside it.

open as a page

Why winsorise rather than trim extreme spenders when analysing an A/B test on revenue per user?

level: middleimportance: should knowfreq 46%

basics

~10 s

Winsorising clips extreme values to a threshold but keeps every randomised user, so the arms stay comparable. Trimming deletes users for their own outcome, conditioning on something the treatment could have caused.

open as a page

How do you choose the cap for a winsorised revenue metric so the A/B comparison stays fair?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Fix one numeric threshold before the experiment reads out, derive it from pre-period data the treatment cannot have touched, and apply that same number to both arms. Report results at neighbouring thresholds as a sensitivity check.

open as a page

Why can clicks-per-user as an experiment's OEC reward clickbait content?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Clicks measure attention captured, not value delivered. A sensational headline earns the click before the user can tell whether the content is worth it, so a variant that misleads scores as a win on the criterion while making the product worse.

open as a page

How does the delta method turn per-user means into a standard error for a ratio metric?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The delta method takes a first-order Taylor expansion of the ratio of two per-user means around their expectations, then reads off the variance. The result combines the numerator variance, the denominator variance and, subtracted, their covariance.

open as a page

Average order value is measured per order but users were randomized: how do you get a valid standard error?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Aggregate each user into revenue and order-count totals, then compute uncertainty across users: either the delta method on those two per-user means, or a bootstrap that resamples users and carries all their orders. Never treat orders as independent rows.

open as a page

A cohort's retention curve flattens near 20% while another decays toward zero — what does each imply?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A curve flattening at a positive level means part of every cohort becomes a durable core, so the active base compounds as you keep acquiring. A curve decaying to zero means the base is capped by acquisition rate times average lifetime.

open as a page

In a signup funnel where email verification has the worst step conversion, where do you invest?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Rank funnel steps by expected additional conversions - users lost times a plausible lift times downstream conversion - not by the worst step rate. And confirm the drop is real rather than a tracking gap before investing.

open as a page

Your capped revenue metric is flat but uncapped revenue is up 6% — how do you decide whether to launch?

level: principalimportance: should knowfreq 38%

basics

~20 s

Treat the uncapped 6% as unresolved until you see its confidence interval and how much survives dropping the largest few spenders. The pre-registered capped metric governs the launch; a whale-driven swing calls for a tail-targeted follow-up, not an override.

open as a page

How would you set the weights in a composite OEC combining purchases, subscriptions and content consumption?

level: principalimportance: should knowfreq 40%

basics

~20 s

Put the components on a common scale first, then set weights from an exchange rate the business will defend — how much consumption one subscription is worth. Weights are a stated value judgment, fixed before launch, not a statistical output.

open as a page

Why is percentage of orders placed on mobile a weak metric to judge an experiment on?

level: middleimportance: nice to knowfreq 27%

basics

~20 s

Its denominator is another outcome. Mobile share rises when mobile orders grow, when desktop orders shrink, or when total orders fall — including outcomes nobody wants. Report it beside the underlying per-visitor counts so a reader can see which side moved.

open as a page

Why is queries-per-user a dangerous OEC for a search product?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Queries per user has no fixed good direction. More queries can mean users are engaged, or that results were bad and they had to keep reformulating. A criterion whose direction is ambiguous cannot decide a launch on its own.

open as a page

In a weekly acquisition cohort table, why is averaging down a column misleading?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

A cohort table is triangular: recent cohorts have not lived long enough to appear in later columns. Averaging a column therefore describes only the older cohorts that reached that age, blending different acquisition weeks and channels.

open as a page

When would you replace revenue per user with a bounded proxy such as purchased-or-not in an A/B test?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Use a bounded proxy when the hypothesis is about whether users buy at all, not how much they spend, and revenue per user cannot be powered at available traffic. A yes/no purchase indicator has variance at most 0.25.

open as a page

How do you decide which guardrail metrics get authority to veto a launch across an organisation?

level: principalimportance: nice to knowfreq 34%

basics

~10 s

Give hard veto power to a few non-tradeable trust metrics such as crash-free session rate. Everything else gets a written threshold and a named owner who can accept a recorded breach.

open as a page

Your experiment platform reports per-event standard errors for every ratio metric: how do you prioritise fixing it?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Quantify the damage first by re-analysing past experiments with user-level variance and counting how many decisions flip. Then make correct variance the platform default for ratio metrics, and prepare teams for wider intervals and larger sample-size requirements.

open as a page