skip to content

How do you decide which guardrail metrics get authority to veto a launch across an organisation?

level: principalimportance: nice to knowfreq 34%

answer

  1. two tiers: refusal versus price
  2. short non-negotiable trust list
  3. named owner for tradeable breaches
  4. alarm fatigue on one side, drift on the other
  5. budgets tracked across launches

basics

~10 s

Give hard veto power to a few non-tradeable trust metrics such as crash-free session rate. Everything else gets a written threshold and a named owner who can accept a recorded breach.

solid answer

~50 s

I split guardrails into two tiers. A short list of trust metrics — crash-free session rate, error rate, availability, a hard latency ceiling — gets automatic veto power, because those are not things the organisation will trade for a business win at any price. Everything else is a business floor: it gets a written degradation threshold and a named owner who can accept the breach as an explicit, recorded tradeoff. The failure modes sit on both sides. Too many hard vetoes and teams learn to route around them, or alarm fatigue sets in and breaches get waved through as noise. Too few, and each launch degrades things slightly within its own rights until the product is measurably worse with no decision anyone can point to. The thing I care most about is that thresholds are set before results are seen and consumption is tracked across launches, not per test.

go deeper

for a junior

Be ready to say that guardrails need thresholds and consequences agreed before a test runs, and that some metrics such as crash rate are treated as non-negotiable.

for a middle

Explain the two tiers — non-tradeable trust metrics versus business floors with a price — and why a threshold too fine to detect at realistic sample sizes ends up ignored.

for a senior

Show you have operated this: named owners, escalation records that capture size and affected segments, and the judgement about which slow metrics belong in production monitoring rather than in a launch gate.

for a principal

Own the balance directly. Argue for the smallest blocking set that holds, defend it against both alarm fatigue and cumulative erosion, and make the tradeoffs the regime permits visible and attributable afterwards.

### The question behind the question Deciding which metrics can veto a launch is not a statistics question — it is a question about which tradeoffs the organisation refuses to make, and about who is allowed to make the ones that remain. The statistics only tell you whether a degradation is real and how large it might be. ### Two tiers **Tier 1 — non-tradeable trust metrics.** A short, stable list applied to every experiment: crash-free session rate, unhandled error rate, availability, a hard latency ceiling, and data-integrity or security invariants. These earn automatic veto power for three reasons. They are about the product working at all rather than about strategy; the harm they measure is hard to reverse, since users who hit a broken experience often do not return to see the fix; and no plausible business win is worth them, which means arguing the exchange rate case by case wastes time and produces inconsistent outcomes. Making them non-negotiable is a *simplification*, not a bureaucratic tax. **Tier 2 — business floors.** Revenue on an engagement-focused test, engagement on a monetisation test, support volume, unsubscribes. These are genuinely tradeable — the whole point of experimentation is finding favourable trades — so a hard veto is the wrong instrument. Each gets a written degradation threshold and a named owner empowered to accept a breach as a recorded tradeoff. The distinction to defend in an interview is *why* the tiers differ: tier 1 encodes a refusal, tier 2 encodes a price. ### Setting the thresholds Three properties make a threshold real: - **Set before results are visible.** A threshold negotiated after the numbers land is not a constraint, it is a rationalisation. If a threshold turns out to be wrong, change it prospectively as a standing policy, not retroactively for the launch that tripped it. - **Detectable at realistic sample sizes.** A threshold tighter than the experiment can resolve is theatre: the guardrail can never be shown to be met, so in practice it gets ignored. Check at design time that a typical experiment produces an interval narrower than the margin. - **Attached to a consequence.** Block, escalate to a named person, or fix-and-rerun. A guardrail with no defined consequence is a chart. ### The failure modes on both sides **Too many hard vetoes.** Every metric someone cares about becomes blocking. Two things follow. Alarm fatigue: breaches become routine, so overrides become routine, and the whole regime loses force — which is worse than having fewer, respected guardrails. And avoidance: teams shrink experiments, shorten runs, or ship outside the experimentation system entirely, moving risk somewhere with no measurement at all. **Too few.** Each launch degrades something slightly, every degradation is individually defensible, and the aggregate is a product that is slower, noisier and more annoying than it was a year ago, with no single decision responsible. This ratchet is the strongest argument for guardrails as *cumulative budgets* rather than per-test checks: track how much each shipped experiment consumed and hold a standing total, so the drift becomes visible and someone owns replenishing it. ### Governance details that separate a strong answer - **Ownership.** Each tier-2 guardrail needs a named owner — usually the team whose outcome it protects, not the team running the test. Self-approval of one's own breaches recreates the problem the guardrail was meant to solve. - **A standing set versus per-test additions.** Keep a small, universal list so it is not re-litigated each launch, and allow test-specific guardrails where a change carries a specific plausible risk. - **Alerting versus deciding.** Not every guardrail needs to block a decision; some are better as production monitors after rollout, especially slow-moving metrics that cannot resolve within a test window. - **Review cadence.** Revisit the list periodically. Guardrails accumulate — every incident adds one, and almost nothing removes one — until the scorecard is unreadable and nobody looks at it. - **Escalation quality.** A good escalation records the size of the degradation, its confidence interval, the affected segments, the budget consumed and the decision rationale. That record is what lets a future team see how much room is left. ### The summary position A small non-negotiable core protects the things the organisation will not sell at any price; everything else has a price, a written threshold, an owner, and an audit trail. The measure of a good regime is not how many launches it blocks but whether the tradeoffs it permits are visible, deliberate and attributable afterwards.

  • How do you keep the guardrail list from growing after every incident?
    Review it on a cadence with an explicit removal bar: a guardrail stays only if it has plausibly changed a decision, or protects something the organisation still refuses to trade. Incident-driven additions should default to production monitoring rather than launch-blocking status, since most incidents are better caught by alerting than by an experiment scorecard nobody can read.
  • Should a team be allowed to accept a breach of a guardrail it owns itself?
    No. The owner should be the team whose outcome the guardrail protects, not the team whose launch it constrains, because self-approval reproduces exactly the conflict the guardrail exists to manage. Where those coincide, require a second signature outside the team and a written record of the size of the degradation and the reasoning.
  • What signals tell you the guardrail regime has too many hard vetoes?
    Overrides becoming routine, breach reviews turning into formalities, and teams shortening or shrinking experiments to slip under the checks. The strongest signal is work shipping outside the experimentation system entirely — the risk has not gone away, it has just moved somewhere with no measurement attached.

saying these in an interview costs you the question

  • Makes every guardrail blocking and calls that rigour
  • Lets the launching team approve its own breaches
  • Negotiates thresholds after the results are known
  • Sets thresholds finer than experiments can detect
  • Checks guardrails per test with no cumulative accounting

context