What is an Overall Evaluation Criterion (OEC) in an A/B test?
answer
- one number decides the launch
- written down before traffic starts
- stops after-the-fact metric shopping
- measurable, sensitive, aligned
- decides, versus explains, versus vetoes
basics
~20 sThe Overall Evaluation Criterion is the single metric, agreed before launch, that an experiment's ship-or-not decision is defined against. It expresses what the team means by the product getting better, so the result cannot be reinterpreted afterwards.
solid answer
~40 sThe OEC is the one number an experiment is read on: the metric, its direction, and the decision rule are written down before traffic is allocated, and the launch call follows from what that metric does. It exists because an experiment dashboard shows dozens of metrics, and if the decision rule is not fixed in advance, whoever wants to ship can always find a metric that moved their way. A good OEC is measurable inside the experiment window, sensitive enough to move detectably at realistic sample sizes, and aligned with the long-run outcome the business actually wants. Other metrics still get read — driver metrics to explain *why* the OEC moved, and separately owned checks that can veto a launch — but only the OEC defines success. One experiment, one criterion, decided first.
go deeper
Be ready to define the OEC in one sentence and say it is fixed before launch. Know that it is one metric with a stated direction, not the whole dashboard.
Explain the mechanics: why declaring the metric first protects the false-positive rate, and what makes a metric fit to be a criterion at all — measurable in the window, sensitive at the traffic you have, aligned with the outcome you want.
Show you can run the hierarchy in practice. Say which metric decides, which explains a movement, and which can object, and describe how you handled an experiment where the criterion and the story disagreed.
Own the standard. Interviewers expect a view on whether the org should have a default criterion, how you get teams to commit before launch, and how you keep a bad criterion from being ratified by a quarter of shipped experiments.
## What the OEC is The Overall Evaluation Criterion (OEC) is the single quantity an online experiment's decision is defined against. Concretely, before any user is assigned to a variant, the team writes down three things: the metric, the direction that counts as an improvement, and the rule that turns the observed result into a decision (for example, ship if the estimated treatment effect on the metric is positive and its confidence interval excludes zero at the pre-declared level). When the experiment ends, the decision is read off that rule. The word *overall* is doing real work. An experiment platform will compute hundreds of metrics; the OEC is the one that stands for the whole judgment of whether the change made the product better. It is the operational definition of *better* for this experiment. ## Why a single criterion, decided in advance Two distinct problems are solved by fixing one metric up front. **The garden of forking paths.** With enough metrics on a dashboard, some will look good and some will look bad by chance alone. If the decision rule is chosen after seeing the numbers, the analyst is effectively selecting the most favourable of many comparisons, and the reported evidence is far weaker than it appears. Declaring the OEC beforehand removes that freedom: the number that decides was named before it was seen. **Decision deadlock.** If three metrics are all called primary and they disagree — one up, one flat, one down — nothing in the design says what to do, so the call falls to whoever argues hardest in the review meeting. A single criterion forces the disagreement to happen earlier, while the experiment is being designed, which is when it is cheap to resolve. Note that these are different objections. The first is statistical; the second is organisational. Both point the same way. ## What makes a metric fit to be the OEC Three properties matter, and they pull against each other. **Measurable in the window.** The metric has to be observable for the users in the experiment during the time the experiment runs. An outcome that only materialises months later cannot be the criterion of a two-week test, however much the business cares about it. **Sensitive.** A metric can only serve as a decision rule if a real change of plausible size actually shows up in it. Sensitivity depends on how noisy the metric is relative to its mean — a metric with a large spread across users needs many more users to detect the same relative change than a quiet one does. A metric that would need ten times the available traffic to detect the effect you care about is not a usable criterion, no matter how meaningful it is. **Aligned.** Moving the metric has to correspond to the product genuinely improving. This is the property that is hardest to verify and easiest to lose, because the metrics that are easy to measure and sensitive are often shallow ones that a variant can move without helping anyone. The practical craft of choosing an OEC is mostly the negotiation between sensitivity and alignment: the most aligned metric is often too slow or too noisy, and the most sensitive metric is often a shallow proxy. ## The rest of the scorecard Choosing one OEC does not mean ignoring everything else. A well-run experiment reads three kinds of metric with three different jobs: - **The OEC** decides. One metric, one direction, declared first. - **Driver or diagnostic metrics** explain. When the OEC moves, these say through which mechanism — which surface, which step, which segment. They never override the decision; they inform the next hypothesis. - **Checks that can veto** protect against unacceptable side effects. These are a separate concern with their own thresholds and their own conventions, and they are the reason an OEC win is not automatically a launch. The hierarchy matters more than the vocabulary. Interviewers listen for whether a candidate can say which metric decides, which metric explains, and which metric objects — and whether they know that only the first is the OEC. ## Common shapes an OEC takes It is often a simple per-user average — an average of some outcome over the randomised users, so that the analysis unit matches the randomisation unit. It can also be a weighted combination of several components rolled into one number, which keeps the single-criterion discipline while acknowledging that success has more than one dimension; the cost is that the weights themselves become a contested design choice. What it must never be is a metric picked after the data came in, or a list of metrics with no rule for what to do when they conflict.
- Why not just declare three primary metrics and require all of them to improve?Requiring all three to move is a stricter rule than any single one, so it needs substantially more traffic and will reject genuinely good changes that help one dimension and are neutral on another. It also does not say what to do about the common case of one metric up and one flat. If several dimensions truly matter, combine them into one criterion with explicit weights rather than leaving the conflict to the review meeting.
- If the OEC is fixed in advance, what do you do when an unexpected metric moves sharply?You treat it as a finding, not a decision. The pre-declared rule still governs this experiment, so an unplanned movement does not by itself flip the call, though a serious unexpected harm is grounds to stop. The right response is to write the observation up as a hypothesis and, if it matters, design the next experiment with that metric declared up front.
- Does every experiment on a platform have to use the same OEC?No. A shared default is useful for comparability across a team, but the criterion should match what the change is trying to do. What must be constant is the discipline: one criterion, declared before launch, with the direction and decision rule written down. A team-wide default that nobody believes for a particular test is worse than a per-test choice that is justified in the design doc.
It is the scoring rule agreed before a match starts. Once the game is over, both sides already know what winning meant, so nobody gets to argue that possession should have counted instead of goals.
saying these in an interview costs you the question
- Calling the whole metric dashboard the OEC
- Choosing the decision metric after seeing results
- Declaring several co-equal primary metrics with no tie-break
- Confusing the criterion with the significance threshold
- Assuming any metric that moved is evidence of success