How would you set sample ratio mismatch policy for a platform running hundreds of experiments?
answer
- the check runs on every experiment
- false-alarm rate times portfolio size
- a warning nobody reads is not a control
- rank segments, do not page on them
- triage cost decides whether the gate survives
basics
~20 sMake the check automatic and blocking: every experiment is tested on arm counts, a flagged test hides its metric readout, and clearing it requires a named, documented cause. Set the threshold strict enough that alerts stay believable at portfolio scale.
solid answer
~40 sThree decisions matter. First, the **threshold**: because the check runs on every experiment, a 0.05 cutoff flags one healthy test in twenty, so most mature platforms sit at 0.001 or stricter and accept missing tiny mismatches. Second, **gate or advisory**: an advisory warning decays into wallpaper, so the default should be a hard gate that suppresses the metric readout, with an override requiring a written cause rather than a checkbox. Third, **triage cost**: the expensive part is the days of investigation, not the alert, so ship automated diagnostics reporting per-stage counts and the worst-offending segments alongside it. Run segment-level checks as a ranked diagnostic rather than a pager — with dozens of segments some will breach any threshold by chance. Track platform mismatch rate as a health metric.
go deeper
Know that platforms run this check automatically on every experiment and that a flagged test normally has its results withheld rather than shown with a warning.
Explain why the alerting threshold is far stricter than 0.05: the check runs on every experiment, so the conventional cutoff would flag one healthy test in twenty and destroy trust in the alert.
Argue the gate-versus-advisory choice concretely and show you have thought about triage cost — per-stage counts, ranked segments and deploy overlays delivered with the alert rather than left to the investigator.
Own the whole system: threshold as a documented tradeoff, override governance with real evidence requirements, prevention rules on filters and instrumentation, and the handful of metrics that reveal whether the policy is respected or quietly bypassed.
## What the policy is actually trading off On one side: shipping decisions based on invalid experiments. On the other: blocked teams, wasted investigation hours, and — the failure mode that quietly kills these programs — an alert nobody believes. The policy has to be tight enough that a flag means something and loose enough that flags are rare. ## Threshold The check runs on every experiment, so its false-alarm rate is multiplied by the size of the portfolio. At the conventional 0.05, roughly one healthy experiment in twenty is flagged; with three hundred concurrent tests that is fifteen spurious investigations at any moment, and teams rationally start ignoring the alert. Cutoffs of 0.001 or stricter are the common practice. The cost is real — genuinely small mismatches go undetected — but small mismatches are also the ones least likely to move a decision, and the alternative is an alert that has lost its authority. Make the number an explicit, documented platform decision rather than a default someone picked. ## Gate versus advisory A warning banner that anyone can scroll past is not a control. The stronger design is that a flagged experiment does not render its metric readout at all: the results page shows the mismatch and the diagnostics instead of the lift. This is uncomfortable, and that discomfort is the point — the single most common way an organisation ships a wrong decision is a team that read the number before the caveat. The override should exist, because genuine false alarms happen, but it should require an explanation attached to the experiment record and reviewed by someone outside the owning team. "We looked and it seemed fine" is not an explanation; "assignment and exposure counts both balanced, the gap is isolated to a bot-filter change deployed after the run, corrected counts pass" is. ## Segment-level checks Running the check within each browser, country, platform and device tier is genuinely useful for diagnosis, and disastrous as an alerting surface. With dozens of segments per experiment and hundreds of experiments, some segment somewhere breaches any fixed threshold constantly — that is what a fixed threshold applied many times does. Compute them, rank them, show them **inside the triage view once the overall check fires**, and do not page on them. The same reasoning argues for reporting the worst segments as evidence rather than as independent verdicts. ## Make triage cheap The binding constraint on a mismatch policy is investigation cost. A flag that takes a senior engineer three days to explain will be overridden; a flag that arrives with per-stage counts (assigned, exposed, analysed), the top deviating segments, and an overlay of deploys during the run can often be explained in an hour. Building that diagnostic bundle is usually a better investment than any refinement of the statistical threshold, because it changes the behaviour the policy actually produces. ## Prevention beats detection Policy should also reach into how experiments are built: - **Log assignment and exposure as separate events**, so the funnel can be measured at all. - **Require that any filter applied to experiment data depend only on information determined before assignment.** Bot removal, deduplication and outlier trimming applied afterward are the most common manufactured mismatches. - **Discourage designs that introduce an extra network hop in only one arm**, or require the same hop in both, so latency-driven loss is symmetric. - **Standardise where the exposure event fires** in the render path so it is not systematically later in one variant. ## Measure the policy Track the share of experiments flagged, the share of flags whose cause was identified, median time to diagnosis, and the override rate. A rising flag rate usually means an instrumentation regression somewhere in a shared library, not a run of bad luck. A high override rate with vague reasons means the gate has already failed and the organisation has quietly returned to reading broken experiments. These four numbers tell you whether the policy is working better than any argument about the threshold will. ## What a strong answer sounds like Name the threshold decision and justify it by portfolio scale rather than convention. Choose the hard gate and defend the discomfort. Acknowledge the false-alarm cost honestly and answer it with triage tooling rather than a softer rule. And close on prevention and measurement — the goal is fewer real mismatches, not more alerts.
- What is the argument against hard-blocking every flagged experiment?False alarms have a real cost: a team blocked on a genuinely healthy test loses days, and repeated unexplainable blocks train people to route around the platform entirely. The answer is not a softer gate but a stricter threshold plus fast diagnostics, so that flags are rare and the rare ones are cheap to resolve.
- How do you keep segment-level checks from producing constant false alarms?Compute them but never page on them. With dozens of segments per experiment, a fixed threshold applied that many times will breach somewhere by chance almost every time. Surface the worst-deviating segments as a ranked diagnostic inside the triage view, used as evidence once the overall check has fired.
- Which metrics tell you the policy is working?Share of experiments flagged, share of flags with an identified cause, median time to diagnosis, and override rate. A rising flag rate usually signals an instrumentation regression in shared code. A high override rate with thin justifications means the gate has effectively been abandoned even though it still exists on paper.
saying these in an interview costs you the question
- Uses a 0.05 cutoff for a check that runs on every experiment
- Makes the check advisory and lets owning teams self-certify
- Pages on every segment-level breach
- Judges the policy by alert volume rather than defects caught
- Invests in threshold tuning while triage still takes days