How does a region of practical equivalence change a Bayesian ship decision?
answer
- a band you would call no difference
- declared before you look at the posterior
- compare the interval to the band
- entirely inside means stop, not continue
- ship neither is a real verdict
basics
~20 sA region of practical equivalence is a band of differences you would call 'no real difference', fixed before analysis. If the posterior for the lift sits entirely inside it, the arms are equivalent and you choose on cost.
solid answer
~50 sYou declare in advance a band around zero — say plus or minus 0.5% relative lift on checkout conversion — inside which any difference is too small to be worth acting on. Then compare the posterior for the difference to that band. Three verdicts follow. If the credible interval for the lift lies entirely inside the band, you have positive evidence of practical equivalence: ship neither on statistical grounds and pick the arm that is cheaper to run or simpler to maintain. If the interval lies entirely outside the band on one side, you have a difference that matters and you act on it. If it straddles a boundary, you are undecided and either collect more data or accept the ambiguity. The real gain is the first verdict: without a band, you can only ever fail to find a difference, never conclude that there is not a meaningful one.
go deeper
Know the idea: a band around zero, agreed in advance, inside which a difference is too small to care about. If the result falls inside it, the arms count as equivalent.
Explain the three verdicts — inside, outside, straddling — and state which interval you compare against the band, since a whole-posterior rule is far stricter than a 95% interval rule.
Demonstrate that 'inside the band' is an affirmative finding you act on, not a failed test, and that the band is derived from the cost of the change and locked before analysis.
Own how bands are set across the org: who signs them off, how they differ for guardrails and headline metrics, and what stops a team widening one after an inconvenient result.
## The problem it solves A difference of exactly zero is a measure-zero event; no continuous posterior ever concentrates on a point. So "are the arms the same?" has no useful answer as literally posed. What decision-makers actually mean is "is the difference small enough that I do not care?" — and that question is answerable, but only once someone states how small is small enough. A region of practical equivalence, often abbreviated ROPE, is that statement: an interval of differences declared indistinguishable from no effect for the purpose of this decision. On a checkout funnel it might be plus or minus 0.5% relative lift. Inside that band, the two arms are treated as the same product. ## The decision rule With a posterior for the difference in hand, compare it to the band: - **Posterior entirely inside the band.** Accept practical equivalence. This is the verdict that does not exist without a band, and it is a genuine finding: you now have evidence that whatever difference exists is too small to matter. The decision moves to non-statistical grounds — which arm is cheaper to operate, simpler to maintain, better aligned with where the product is heading. On a migration that costs engineering time, this verdict says do not migrate. - **Posterior entirely outside the band, on one side.** The difference is real and large enough to act on. Act. - **Posterior straddling a boundary.** Undecided. Either gather more evidence or decide on other grounds while acknowledging the ambiguity. Teams differ on whether to compare the *whole* posterior or a 95% credible interval — typically a highest-density interval, the narrowest interval containing 95% of posterior belief. The stricter whole-posterior version almost never fires at realistic sample sizes; the interval version is what gets used. Either way, state which one you mean, because the strictness differs. ## Where the band comes from It must be set before you look at the result, and it must come from the cost side of the business, not from the data. The question to ask the product owner is: below what change would you not bother making this change at all? For a migration that costs two engineer-months and adds a service to operate, the honest answer may be a full percent; for a one-line copy change with no maintenance cost, it may be far tighter. Setting the band after seeing the posterior is the failure mode that voids the whole exercise. A band drawn around whatever the data produced can be made to yield any verdict you want, and it makes the equivalence claim unfalsifiable. The band is also directional in some decisions. For a guardrail metric — latency, error rate, refund rate — you often care only that the variant is not meaningfully worse, so the meaningful boundary is one-sided. ## An example verdict A checkout test runs to a posterior for relative lift that spans roughly minus 0.2% to plus 0.3%, entirely within a pre-declared band of plus or minus 0.5%. Read carefully, this is not a failed experiment. It is a successful one: the variant, whatever else it does, does not move checkout conversion by an amount worth the change. If the variant carries an ongoing cost — a new dependency, more code, a vendor bill — the correct call is to ship neither and retire the variant. Teams that lack the band read the same posterior as "inconclusive, run it longer", and burn traffic pursuing a difference they have already established is too small to want. ## The relationship to loss-based rules A practical-equivalence band and an expected-loss tolerance are two ways of encoding the same underlying fact: that the business has a threshold below which differences are not worth acting on. The band expresses it as an interval on the effect scale and answers "are these the same?"; a loss tolerance expresses it as a cost and answers "how much do I risk by choosing?" They are complementary, and mature experiment readouts show both. ## What interviewers listen for That you set the band in advance and from cost. That "inside the band" is an affirmative conclusion rather than a failure to find something. That you say which interval you are comparing. And that you do not let a band be widened after the fact until a decision appears.
- Where does a number like plus or minus 0.5% come from?From the cost of making the change, not from the data. Ask the owner what improvement would be too small to justify building, operating and maintaining the variant. A migration costing engineering months and a new service to run deserves a wide band; a copy tweak with no ongoing cost deserves a tight one. It is a product and finance call, recorded before analysis.
- What do you do when the posterior straddles a band boundary?Call it undecided. You have neither established a difference worth acting on nor established equivalence, so the honest options are to collect more evidence or to decide on other grounds while stating that the statistical question is open. What you must not do is widen the band until one of the clean verdicts appears.
- Does the band ever need to be one-sided?Often, on guardrail metrics. For latency, error rate or refunds you usually care only that the variant is not meaningfully worse, so the decision is whether the posterior stays above a single harm boundary. Improvement on a guardrail is welcome but is not what the check is for, which makes a symmetric band the wrong shape.
It is the tolerance stamped on a machined part. Nothing is ever exactly the specified width; the tolerance is what turns 'close enough' into a decision you can sign off.
saying these in an interview costs you the question
- Sets the equivalence band after seeing the posterior
- Reads a posterior inside the band as 'needs more data'
- Insists on an exactly zero difference to claim equivalence
- Widens the band until a clean verdict appears
- Applies a symmetric band to a one-sided guardrail metric