skip to content

What policy should govern segment reporting in A/B test readouts across an experimentation team?

level: principalimportance: nice to knowfreq 32%

answer

  1. separate permission from authority
  2. three tiers, one decision
  3. print the cut count automatically
  4. promotion needs a powered rerun
  5. too strict also fails

basics

~20 s

Make the all-up metric the decision, allow a short pre-declared segment list per experiment, keep exploratory cuts in a labelled section with the cut count disclosed, and require a powered confirmatory run before any segment result changes a launch.

solid answer

~50 s

I would separate the readout into three tiers with different authority. The all-up primary metric and its interval is the launch decision. A short list of segments declared at design review, each with a mechanism and a sample plan, can qualify that decision. Everything else is exploratory: shown, labelled, reported with the number of cuts examined, and structurally unable to move a decision on its own — promotion requires a confirmatory experiment powered for that segment. The tooling should make this cheap: record which cuts were pulled, print the count in the readout, and default templates to the pre-declared list. The real design tension is that a strict policy suppresses genuine heterogeneity while a loose one fills the roadmap with phantom wins, so I would tune the confirmatory bar by stakes — cheap, reversible targeting changes can ship on weaker evidence than expensive, hard-to-unwind ones.

go deeper

for a junior

Be ready to say that a readout should separate planned segment reads from exploratory ones and that the number of cuts examined belongs in the report.

for a middle

Explain why the tiers exist: exploratory cuts carry no valid error rate, so they can inform the next experiment but not the current launch decision.

for a senior

Show how you would operate this on a real team, including powering the declared segments, running the confirmatory experiment, and expecting shrinkage between discovery and replication.

for a principal

Own both failure modes and the tooling that resolves them, and be ready to defend where you set the confirmatory bar as a function of how reversible the resulting decision is.

## Why this needs a policy at all Segment reading is the most common route by which an experimentation programme starts producing wins that do not show up in the business. Individual analysts rarely intend to mislead; the pathology is structural. A flat result is unwelcome, slicing is free, tooling makes thirty cells available in one click, and nobody records how many cells were looked at. A policy is how you make the honest path the default path rather than an act of personal discipline repeated under deadline pressure. ## A three-tier readout **Tier 1 — the decision.** The primary metric, all-up, with its interval. This is what launches or does not launch. Nothing in the tiers below overrides it. **Tier 2 — pre-declared segments.** A short list, agreed at design review, each entry carrying: the cut, the mechanism you expect, and a sample plan showing the segment has enough traffic to detect an effect worth acting on. Short is load-bearing — three declared cuts you can power beats fifteen you cannot. Tier 2 results may qualify the decision: gate a launch to a segment, hold back a rollout, trigger a follow-up. **Tier 3 — exploration.** Everything else, shown openly, in a section labelled exploratory, with **the number of cuts examined printed at the top**. Tier 3 generates hypotheses. By policy it cannot carry a launch on its own. ## The promotion rule The policy lives or dies on what it takes to move a Tier 3 observation into a decision. A workable rule: a confirmatory experiment, the segment declared in advance, powered for that segment specifically, with the decision rule fixed before the run. Expect a real effect to come back smaller than its discovery estimate, because the discovery was selected for being extreme; the confirmatory number is the one you forecast with. Tune the bar by stakes rather than applying one threshold everywhere: - **Cheap and reversible** (a copy variant shown to one segment) — a weaker bar is rational, because being wrong costs a rollback. - **Expensive or sticky** (a permanently divergent experience, a pricing rule, a headcount commitment) — insist on the full confirmatory run, because unwinding costs quarters. ## Making it cheap to comply Policies that depend on virtue decay. Build the rule into the tooling: - The experiment template asks for pre-declared segments at design time, before a launch button appears. - The analysis surface logs which cuts were pulled, and the readout renders the count automatically. Nobody has to remember to disclose. - The readout template ships with the three tiers as sections, so an exploratory finding physically cannot appear in the decision block. - Segment reads carry intervals by default, so a cell too small to say anything looks visibly uninformative rather than exciting. ## The tension to own A principal-level answer names the cost on both sides. **Too strict** and you suppress real heterogeneity: treatment effects genuinely do vary across users, and a programme that never looks will never find the segmentation that matters, while analysts route around a rule they find unreasonable and do the slicing in private spreadsheets — which is worse than doing it in the open. **Too loose** and every flat experiment yields a segment win, the roadmap fills with launches whose promised lift never materialises, and within a year leadership stops believing experiment readouts at all. That loss of credibility is the expensive failure, because it removes the programme's ability to stop bad ideas. The resolution is not a stricter threshold; it is a clear *status* distinction plus a cheap promotion path. Exploration stays fully permitted and fully visible; only its authority is constrained. ## Cultural levers - **Make flat results publishable.** If a team is judged on wins, they will find one in the grid. Review rituals that treat a well-run flat experiment as a good outcome remove most of the pressure that causes segment fishing. - **Review readouts, not just designs.** Design review catches unpowered segments; readout review catches the ones added later. - **Track the promotion record.** Count how many Tier 3 findings later replicated. A team that sees its own hit rate — often startlingly low — self-corrects faster than any policy document achieves. ## What an interviewer is listening for They want to hear that you separate *permission* from *authority*, that you have an operational promotion rule rather than a slogan, that you would build it into tooling rather than training, and that you can articulate the cost of over-correcting. A candidate who simply says "we should not p-hack" has not answered the question; the job is designing the system in which the right behaviour is also the easy one.

  • How do you keep a strict segment policy from suppressing genuine heterogeneity findings?
    By constraining authority rather than permission. Exploration stays fully allowed and fully visible in a labelled tier, and the promotion path is cheap enough to use: one confirmatory run powered for the segment. Analysts only route around a rule when compliance is expensive, so the design goal is a policy that costs a follow-up experiment rather than an argument.
  • What single metric would tell you whether the policy is working?
    The replication rate of promoted segment findings. Track how many Tier 3 observations that went to a confirmatory run reproduced, and at what fraction of their discovered size. A collapsing rate says the exploratory bar is too low; a rate near one hundred percent with few attempts suggests the team has stopped exploring at all.
  • Should the confirmatory bar be the same for every segment finding?
    No. Tie it to reversibility and cost. A copy variant shown to one segment can ship on weaker evidence because a rollback is cheap, while a permanently divergent experience, a pricing rule or a staffing commitment should require the full powered rerun. A uniform threshold either blocks cheap learning or waves through expensive mistakes.

saying these in an interview costs you the question

  • Bans exploratory segment analysis outright
  • Relies on analyst discipline instead of tooling defaults
  • Never discloses how many cuts were examined
  • Applies one evidence bar regardless of reversibility
  • Treats a flat experiment as a failed experiment

context