How do you set the ship threshold on probability of superiority for a Bayesian experiment program?
answer
- 95% is convention, not law
- price the cost of a wrong launch
- tier the bar by reversibility
- certainty is not magnitude
- commit the rule before the data lands
basics
~20 sThe threshold should encode the cost of being wrong, not a borrowed convention. Set it tighter for expensive, hard-to-reverse changes and looser for cheap reversible ones, pair it with a magnitude requirement, and fix it before the experiment runs.
solid answer
~50 sThere is no law making 95% the bar; the threshold is a business parameter expressing how badly a wrong launch hurts relative to a missed win. Three principles hold. First, tier it by reversibility and cost: a copy change you can roll back in an afternoon does not need the bar a data migration needs. Second, never let probability stand alone — a 96% probability of superiority on a posterior lift centred on 0.05% is near-certain and still not worth the migration cost, so the rule must also demand a difference large enough to matter, via an expected-loss tolerance in metric units or a probability keyed on a meaningful lift rather than on zero. Third, fix the rule before the data lands, because a threshold that moves after a near-miss is not a threshold. Guardrails invert the logic: they demand a low probability of harm, not a high probability of gain.
go deeper
Know that the ship bar is a choice someone made, not a constant of statistics, and that a high probability of the variant winning does not by itself justify shipping.
Be able to explain why probability alone is insufficient: it rises with precision as well as with effect size, so a near-certain but tiny lift can clear any probability bar.
Show how you would operationalise a rule for a specific decision — bar, magnitude requirement, guardrail conditions — and defend why that decision warrants those numbers.
Own the program-level scheme: tiers by reversibility and cost, pre-registration and the governance around exceptions, and an explicit account of what an over-tight bar costs in learning speed.
## The bar is a business parameter The 95% figure that shows up as a default in experimentation tools is inherited convention, not a mathematical requirement. In a Bayesian decision framework the threshold on probability of superiority is a statement about how you trade a wrong launch against a missed opportunity. Two organisations with identical data and different cost structures should rationally set different bars, and a leader who cannot articulate why theirs is where it is has not made a decision, only accepted a default. ## Tier the bar by the decision, not by the metric The most useful structure is a small set of tiers keyed to what the change costs to make and to undo: - **Cheap and reversible** — copy, layout, a flagged UI variant. The cost of a wrong ship is a rollback. A looser bar is rational; the organisation's real risk here is moving too slowly, not shipping a dud. - **Expensive or slow to reverse** — a platform migration, a pricing change, anything with contractual or data-model consequences. Here a wrong call is paid for over quarters, and the bar should be materially tighter, often supported by a follow-up holdback. - **Guardrails** — latency, errors, refunds, support contacts. The logic inverts: you are not looking for a high probability of gain, you are requiring a low probability of meaningful harm. A change can ship on a strong primary result while still being blocked by a guardrail whose harm probability is uncomfortably high. Tiering also solves an organisational problem: it gives teams a defensible answer to "why did you ship on 90% when the other team needed 97%?" that is not negotiation. ## Probability alone is never the rule The sharpest illustration is a result with 96% probability of superiority attached to a posterior lift centred on 0.05%. The direction is close to settled; the magnitude is nothing. If the change costs a migration, the correct call is not to ship, and any rule keyed purely on probability crossing a bar will get it wrong. At high traffic this case is not exotic — it is the default outcome of a large program testing many small changes, because probability of superiority rises with precision as well as with effect size. Two repairs, and mature programs use both: - **Threshold a meaningful lift instead of zero.** Require high posterior probability that the lift exceeds an amount worth having, rather than merely exceeds zero. - **Add a cost-denominated tolerance.** Require the expected loss of the decision, in metric units, to be under a tolerance the business signed off on. The second is more legible to non-analysts, because it is quoted in conversion or revenue rather than in probability. ## Fix the rule before the data lands A threshold that can be revisited after a near-miss is decoration. The practical machinery: the decision rule — bar, magnitude requirement, guardrail conditions and how long the experiment runs — is recorded in the experiment plan before launch, and changing it afterwards requires the same review as launching without one. This is less about statistical purity than about incentives. Everyone will have a plausible story for why the specific result in front of them warrants an exception, and the aggregate of those exceptions is a program that ships noise. ## The costs of getting the level wrong A bar set too high is not the safe choice, which is the point leaders most often miss. It buys certainty with time and traffic, and its costs are real: slower learning cycles, fewer experiments completed per quarter, and a bias toward the incumbent that compounds into stagnation. A bar set too low fills the product with changes that do not replicate and erodes trust in the experimentation system itself, which is the more expensive failure because it is hard to reverse. The right level depends on how many experiments you run, how large your true effects tend to be, and how much traffic a decision costs. ## What a strong answer sounds like Name the threshold as a cost tradeoff rather than a statistical constant. Give a tiering scheme tied to reversibility. Insist that magnitude enters the rule alongside probability, with the near-certain-but-tiny case as the illustration. Commit the rule in advance and treat exceptions as a governed event. And acknowledge, without prompting, that an over-tight bar has its own cost, so the goal is calibration rather than caution.
- Would you use the same threshold for a guardrail metric as for the primary metric?No — the direction of the question flips. On the primary you demand a high probability of gain; on a guardrail you demand a low probability of meaningful harm, usually against a one-sided harm boundary. A change can clear the primary bar and still be blocked because the probability of a real latency or refund regression is uncomfortably high.
- What keeps teams from moving the bar once they see a near-miss?Make the rule part of the pre-launch plan and make changing it a reviewed event rather than an analyst's judgement call. Every borderline result comes with a persuasive story for why it is the exception; the defence is process, not willpower. Publishing the pre-registered rule alongside the readout makes deviations visible.
- What is the cost of setting the bar too high?Slower learning and a structural bias toward the incumbent. A tighter bar consumes more traffic and calendar time per decision, so you complete fewer experiments and reject real but modest wins. Over-caution is a choice with a price, not a free safety margin, and a program that never ships anything marginal is usually under-calibrated rather than rigorous.
saying these in an interview costs you the question
- Copies 95% from convention without asking what it buys
- Ships on probability alone regardless of the size of the lift
- Applies one threshold to reversible and irreversible changes alike
- Loosens the bar after seeing a near-miss result
- Treats a tighter threshold as free caution with no cost