How would you design an experiment results dashboard so it does not invite peeking?
answer
- hide the verdict, show the progress
- colour is a verdict too
- bind the rule at creation time
- guardrails alert, results wait
- make the override visible, not impossible
basics
~20 sHide the verdict until the planned end date. Show progress toward the target sample size, the decision date and health diagnostics, but withhold p-values and significance badges until the test is due, then present the decision as read.
solid answer
~50 sThe core move is to separate what people legitimately need in-flight from what tempts them to decide early. In-flight, show enrolment progress against the planned sample size, the decision date, and health signals — instrumentation firing, both arms receiving traffic, errors by variant. Withhold the significance verdict: no p-value, no significance badge, no green-versus-red colouring on the effect estimate. At the horizon, unlock a single decision view whose headline is the pre-registered metric with its interval, and record the call. Two supporting details matter: the stopping rule and metric should be captured in the test plan at creation time so the dashboard renders a commitment rather than a preference, and harm guardrails should still alert continuously, since stopping for harm is a different decision. The point is to make the honest path the default one rather than relying on discipline.
go deeper
Know the basic idea: a dashboard that shows significance from day one makes early decisions almost inevitable, so results are usually withheld until the planned end date.
Explain which elements do the damage — the p-value, the significance badge, verdict colouring, effect-size leaderboards — and what can safely stay visible, such as enrolment progress and health checks.
Demonstrate the full design: commitments captured at creation, results locked until the horizon, guardrails alerting continuously, and a visible override. Name the responsiveness cost you are accepting.
Argue the governance case. Decide who owns the rule, how exceptions are granted and reviewed, and which behavioural metrics prove the change worked rather than merely annoyed people.
## Why tooling, not willpower Every team knows peeking is wrong and most peek anyway, because the dashboard makes it a one-click question and the organisation rewards speed. If early significance is visible, someone will eventually act on it — and the act will be defensible in the moment because the number really did say 0.04. Treating this as a training problem produces a memo; treating it as an interface problem produces a fix. The design goal is: **make the calibrated decision the path of least resistance, and make the uncalibrated one require a deliberate, visible override.** ## What people legitimately need mid-flight Suppressing everything is the wrong answer, and an interviewer will push on it. There are real reasons to open an experiment on day 2: - **Enrolment progress.** Is the test filling as expected, and is the decision date still realistic? - **Operational health.** Is the instrumentation firing, are events arriving, are errors or crashes concentrated in one variant? - **Harm signals.** Is anything going badly enough to justify a stop? None of these require the effect estimate on the primary metric to be dressed as a verdict. So the in-flight view can be rich and still safe. ## What to withhold, and how Withhold the things that turn a glance into a decision: - **The p-value and any significance badge.** This is the single highest-value suppression. - **Verdict colouring.** Green and red cells on the effect estimate are a verdict even without a number attached; people read colour faster than they read text. - **The point estimate presented as a result.** If the estimate is shown at all before the horizon, show it with its full interval and without a comparison to zero, so its instability is visible rather than implied. - **Sorting or ranking experiments by effect size.** A leaderboard is a peeking machine. A good pattern is a locked panel: the primary metric area shows "Results available <date> — 43% of planned exposure collected" instead of numbers. It communicates that the information exists and is not being withheld arbitrarily, only deferred. ## Binding the commitment at creation time The dashboard can only enforce a rule that exists. So the experiment-creation flow should require, before traffic starts, the primary metric, the planned sample size or end date, and the decision rule. Those become the fields the results view keys off. Two properties follow: - The rule is **pre-registered by construction** — nobody has to remember to write it down. - Changes to it after launch are **recorded and visible**, so extending a run or swapping the primary metric mid-flight leaves a trace instead of quietly happening. That audit trail is often more valuable than the lock itself. Most teams will not lie in an artefact their peers can read. ## The escape hatches, designed rather than improvised A system with no exit will be routed around, so design the exits: - **Harm guardrails alert continuously**, with a documented threshold, and stopping on one is logged as a harm stop rather than a result. This keeps the safety case out of the significance conversation entirely. - **A genuine early-decision need** is answered at planning time by choosing an appropriate design, not by unlocking the panel. The dashboard's job is to make that choice happen before launch rather than at 9am on day 3. - **A manual unlock exists**, requires a reason, and is visible to the team. If unlocking is common, that is data about the org's real constraints and should drive the next design change. ## Presenting the result at the horizon When the panel unlocks, present it as a decision rather than an invitation to interpret: the pre-registered metric, the estimate with its interval, the pre-agreed rule, and the resulting call. Secondary metrics belong lower and clearly labelled as exploratory. Then require the decision to be recorded. Locking the write-up to the same artefact stops the historical record from being rewritten later by whoever remembers the test most favourably. ## How to argue this in an interview Name the tradeoff honestly. Hiding results costs the team responsiveness and will annoy people who have legitimate urgency; you are buying error-rate integrity with patience. State that you would measure the fix — unlock frequency, how often experiments get extended past their planned end, how often a shipped win fails to reproduce in a follow-up — because a governance change that is never measured is indistinguishable from a memo.
- Isn't hiding results from your own team paternalistic?It is a constraint the team agrees to in advance, like a code review or a deploy gate, and it binds everyone equally including the person who built it. The alternative is not freedom but a decision procedure whose error rate nobody knows. I would keep the override available and visible rather than absent — the goal is that acting early is a deliberate, recorded choice, not a reflex.
- How would you know whether the dashboard change actually worked?Track behaviour, not sentiment: how often the results panel is unlocked early and by whom, how often experiments are extended past their planned end date, and how often shipped wins fail to reproduce when re-measured later. A drop in early unlocks with no rise in complaints about speed is the good outcome; frequent unlocks tell you the org has a real constraint the design has not addressed.
- What do you still show on day 2, given that people will open the page anyway?Enrolment against the planned sample size, the decision date, and operational health — events arriving, both arms receiving traffic, errors by variant. Those answer the legitimate day-2 questions, which are about whether the experiment is working rather than whether the variant won. Giving people a useful in-flight view is part of why the locked verdict is tolerable.
saying these in an interview costs you the question
- Relies on training and good intentions with no tooling change
- Hides everything, so broken experiments run for two weeks undetected
- Keeps green and red colouring while removing only the p-value
- Has no exception path, so the team routes around the system
- Lets the primary metric be chosen after results are visible