What is a guardrail metric in an A/B test, and how does it differ from a driver metric?
answer
- one decides, one vetoes, one explains
- the thing you must not break
- asymmetric: only degradation matters
- latency, crashes, unsubscribes
- drivers decompose the primary metric
basics
~20 sA guardrail is a metric the experiment must not damage — latency, crash-free sessions, unsubscribes — and a big enough degradation can veto the launch. A driver metric only explains why the primary metric moved; it never vetoes.
solid answer
~40 sA guardrail metric is one the experiment is not allowed to damage: page latency, crash-free session rate, error rate, unsubscribe rate. It carries veto power — a large enough degradation blocks the launch even when the primary metric wins — but a guardrail moving in the good direction is never by itself a reason to ship. A driver metric sits underneath the primary metric and explains it: if purchases move, add-to-cart rate and checkout completion rate tell you which step changed. Drivers are diagnostic, not decisive. The clean way to say it in an interview is that each metric type carries a different decision rule: the primary decides, guardrails can veto, drivers explain.
go deeper
Be ready to name the three roles and give one concrete example of each: a primary metric the decision hangs on, a guardrail such as latency or crash rate, and a driver such as add-to-cart rate.
Explain the asymmetry — guardrails are checked for degradation only, against a stated threshold — and why driver metrics are more sensitive than the primary but are never allowed to decide a launch.
Show you declare metrics and thresholds before launch, keep the standing guardrail set small, and attach a consequence to every guardrail so a breach triggers a defined action rather than a debate.
Own the framing that guardrails encode which tradeoffs the organisation refuses to make. Argue for which are non-negotiable trust metrics and which are business floors with a negotiable threshold and a named owner.
### Three roles, three decision rules Every metric on an experiment scorecard should have a written decision rule attached to it before the test starts. Three roles cover almost everything: - **Primary metric** — the one the ship decision hangs on. There is normally exactly one, fixed before launch. - **Guardrail metrics** — things the change must not break. They cannot win an experiment; they can lose one. A guardrail is checked against a *degradation threshold*, not against "did it improve". - **Driver metrics** (also called supporting, secondary, or diagnostic metrics) — the intermediate steps that decompose the primary metric. They are read to *explain* a result and to build confidence that the mechanism is the one you intended. They never override the primary or the guardrails. The asymmetry in the guardrail rule is the part candidates most often miss. A primary metric is two-sided in spirit: you are trying to move it and you accept the answer either way. A guardrail is one-sided in spirit: nobody launches a feature *because* crash rate fell; you are only asking whether it got meaningfully worse. ### What actually belongs on the guardrail list Guardrails split roughly into three families. **Trust metrics** are about product integrity rather than product strategy: crash-free session rate, client and server error rate, p95 page or request latency, availability. They apply to essentially every experiment because almost any code change can break them, and they are rarely tradeable — an organisation does not usually accept "we crash 2% more often but sign-ups are up". **Business-floor metrics** protect a value the current business already depends on: revenue per user on a test aimed at engagement, or engagement on a test aimed at monetisation. These *are* tradeable, and that is exactly why they need a pre-agreed threshold rather than an argument after the fact. **Counter-metrics** are the subclass aimed at gaming: the metric you would expect to move badly if the treatment wins by extracting value from users instead of creating it. Unsubscribes and support-ticket volume against an aggressive upsell flow are typical. Every counter-metric is a guardrail, but not every guardrail is a counter-metric — latency is a guardrail and nobody is gaming it on purpose. **Sanity or invariant checks** — traffic split, sample counts per arm, instrumentation coverage — are sometimes lumped in with guardrails. They are worth separating in your answer: a broken invariant means the experiment is invalid and must be discarded, whereas a breached guardrail means the experiment is valid and the result is bad. ### Driver metrics in practice Suppose the primary metric is purchases per user and it rises 2%. Driver metrics tell you *how*: search result click-through rose, add-to-cart rate rose, checkout completion was flat. That story is coherent, so you believe the result. If instead purchases rose while every intermediate step was flat, the result is suspicious — most likely a measurement or assignment problem rather than a real effect. Drivers are also usually more sensitive than the primary: they sit closer to the change, have less noise between the change and the observation, and so move detectably in tests where the primary metric cannot reach a conclusion. That sensitivity is why they are useful for *reading* an experiment and dangerous for *deciding* one — a driver can move without the outcome you care about moving at all. ### The discipline around the list Four habits distinguish a mature setup: 1. **Declare all three groups before launch.** Metrics discovered after the result is known are inseparable from the temptation to pick the ones that support the conclusion you already prefer. 2. **Attach a number to each guardrail.** "Latency must not get worse" is unenforceable, because a large enough experiment can detect a degradation far too small to matter, and a small one cannot detect a degradation large enough to hurt. "p95 latency must not degrade by more than 10 ms" is enforceable. 3. **Keep the standing guardrail set small and stable** so it is not renegotiated every launch, and add test-specific guardrails only where the change has a plausible specific risk. 4. **Say in advance what a breach triggers** — automatic block, escalation to a named owner, or fix-and-rerun. A guardrail with no consequence attached is a chart, not a guardrail. The same metric can play different roles in different experiments. Latency is a guardrail for a ranking-algorithm test and the primary metric for an infrastructure test. The role lives in the experiment's plan, not in the metric itself.
- Can the same metric be a guardrail in one experiment and the primary metric in another?Yes, and it usually is. Latency is a guardrail for a ranking or UI test — you are not trying to move it, you just must not wreck it — and the primary metric for an infrastructure or caching test aimed squarely at speed. The role is a property of the experiment's decision plan, not of the metric definition.
- Where do trust metrics such as crash-free session rate fit in that scheme?They are standing guardrails applied to every experiment, not test-specific ones. Almost any change can introduce a crash or an error, and organisations rarely accept trading product stability for a business win, so these tend to be the guardrails with the hardest thresholds and the least room for negotiation after the fact.
- What is the difference between a breached guardrail and a failed sanity check?A failed sanity check — skewed traffic split, missing events in one arm — means the experiment itself is untrustworthy, so you discard the result and fix the setup. A breached guardrail means the experiment was run correctly and the honest answer is that the change causes harm. One invalidates the measurement; the other reports a real cost.
A guardrail is the medical trial's safety monitoring: the trial is run to test efficacy, but a bad enough safety signal stops it regardless of how well the efficacy endpoint is doing.
saying these in an interview costs you the question
- Calls every secondary metric on the scorecard a guardrail
- Tries to optimise a guardrail instead of protecting it
- Says a big primary win lets you ignore guardrails
- Adds guardrails only after seeing the results
- States a guardrail with no threshold or consequence