How would you set the weights in a composite OEC combining purchases, subscriptions and content consumption?
answer
- different scales, so normalise first
- weights are an exchange rate
- a value judgment, not a fitted output
- freeze before launch, review on a cadence
- report components beside the composite
basics
~20 sPut the components on a common scale first, then set weights from an exchange rate the business will defend — how much consumption one subscription is worth. Weights are a stated value judgment, fixed before launch, not a statistical output.
solid answer
~50 sFirst normalise: purchases, subscriptions and content consumption live on wildly different scales and variances, so a raw sum is dominated by whichever component has the largest numbers. Standardise each component or convert all three into one common unit before combining. Then the weights encode an exchange rate — what one subscription is worth in purchases, what an hour of consumption is worth against either — and that is a business value judgment, not something the data can hand you. Two routes: elicit the rate from leadership and document it, or fit weights to predict a long-run outcome the org already agrees on, then sanity-check the implied rates. Whichever route, the weights are declared before launch, published, and reviewed on a fixed cadence — never retuned after an experiment reads badly. And always report the components alongside the composite, because a flat composite can hide one component up and another down.
go deeper
Know that components measured on different scales must be put on a common footing before being added, and that the weights are chosen by people rather than computed from the experiment.
Explain normalisation options and what a weight means once components are standardised, plus why a flat composite can conceal one component rising while another falls.
Show how you would derive and validate weights against a long-run outcome, estimate the composite's variance rather than assuming it improves power, and keep the components on the report.
Own the governance and the incentives. Expect to argue who sets the exchange rate, why it is frozen before launch and revised only on a cadence, and how the weights steer what every team proposes next quarter.
## Why a composite at all The single-criterion discipline is about having one rule that decides. It does not require that success be one-dimensional. When a product genuinely succeeds along several dimensions at once — transactional revenue, recurring revenue, and engagement with content — collapsing them into one weighted number preserves the discipline (there is still exactly one thing that decides) while acknowledging the multi-dimensional reality. The alternative, declaring three co-equal primaries, does not resolve conflicts; it defers them to whoever argues best after the data lands. A composite forces the argument to happen up front, when it is about values rather than about a specific pending launch. That is the real reason to build one. ## Step one: normalisation, before any weight is discussed The components are on incompatible scales. Purchases might average a fraction of one per user in the window; consumption might average tens of minutes; subscriptions are rare events. Summing them raw means the component with the biggest numbers dominates regardless of intent, and the component with the largest variance dominates the composite's noise. Two defensible normalisations: **Standardise each component** by subtracting its mean and dividing by its standard deviation, computed on a fixed reference population — not on the experiment's own data, which would make the criterion depend on the results it is judging. The composite is then a weighted sum of standardised units, and the weights say how many standard deviations of one component trade for a standard deviation of another. **Convert to a common unit** — typically monetary. Each component is multiplied by an agreed per-unit value, and the composite is one interpretable number. This is more meaningful but requires the org to state what an hour of consumption is worth, which is exactly the hard part. Either is fine; doing neither is not. ## Step two: where the weights come from The weights are an exchange rate between dimensions of value. No statistical procedure can produce them from experiment data alone, because the data never says how much the organisation *should* care about consumption relative to subscriptions. Two legitimate routes: **Elicited weights.** Ask decision-makers directly, in the form of concrete trades: would we ship a change that lost this many purchases to gain one subscription? Iterate until the implied rates are ones people will still endorse in six months. The output is a documented value judgment, which is what it actually is. The strength of this route is transparency and accountability; the weakness is that stated preferences are unstable and often inconsistent. **Fitted weights.** Choose a long-run outcome the org already agrees is the real objective, and fit weights so the composite best predicts it. This grounds the rates in evidence. Two cautions. First, fitting to *user-level* correlation with the outcome is much weaker than it looks — a component can predict which users are valuable without predicting what an intervention does. Where possible, fit on *treatment effects* across a back catalogue of past experiments: which combination of component effects best predicts the effect on the long-run outcome. Second, always read the fitted weights back as exchange rates and check whether anyone would defend them; a fitted weight that implies one subscription is worth a thousand purchases is a signal of collinearity or a small sample, not a discovery. In practice a hybrid is common: fit to get a starting point, then round to numbers leadership will sign. ## Step three: governance Weights that can be changed after seeing a result are not a criterion; they are a dial for producing the desired verdict. Three rules make the composite honest: - **Declared before launch** and versioned, like any other part of the analysis plan. - **Reviewed on a schedule**, not on demand. Annual or semi-annual revision is healthy — strategy changes, and a weight set for a growth phase is wrong in a monetisation phase. Revision triggered by a disappointing experiment is not. - **Recomputed consistently.** When weights change, know that results before and after are not directly comparable, and say so. ## Step four: always report the components A composite compresses, and compression hides. A flat composite is consistent with all three components flat, and equally with a large gain in one exactly cancelling a large loss in another. Those are entirely different situations and often warrant different decisions. So the composite decides, but the component effects are always shown next to it, and a decision where a component moved sharply in the wrong direction gets human attention even when the composite says ship. ## Costs to acknowledge **Sensitivity is not automatic.** Combining metrics does not reliably reduce noise. The composite's variance depends on the components' variances and their correlations; a heavy weight on a high-variance component can make the composite noisier than the best single metric. Estimate the composite's variance directly on historical data rather than assuming averaging helps. **Interpretability drops.** A percentage change in a weighted sum of standardised components does not mean anything intuitive. Teams reason less well about it, which is a real organisational cost. **The weights become a target.** If consumption is heavily weighted, expect proposals optimised toward consumption. Setting weights is setting incentives across every team that reads them, which is why this is a leadership decision rather than an analyst's. ## What a strong answer sounds like Normalise first; derive weights from an exchange rate that is either elicited and documented or fitted against a long-run outcome and then sanity-checked; freeze and version them; review on a cadence; and always publish the components beside the composite. Above all, name clearly that the weights are a value judgment the organisation owns, not a number the data produces.
- Does combining three metrics into one composite reduce the noise and improve power?Not reliably. The composite's variance depends on the components' variances and their correlations, so a heavy weight on a volatile component can make the composite noisier than the quietest single metric. Averaging only helps when components are weakly correlated and comparably scaled. Estimate the composite's variance on historical data before assuming it buys sensitivity.
- A composite is flat but one component is sharply up and another sharply down. What do you do?Do not ship on the composite's flatness alone. Offsetting movements are a substantively different situation from genuine neutrality, and the composite cannot distinguish them. Report both components, treat the offset as the finding, and ask whether the exchange rate implied by the weights is one leadership actually endorses for this trade. That is a decision for people, not for the formula.
- Leadership asks to reweight the composite after an experiment came back negative. How do you respond?Decline for this experiment and offer a scheduled review. Changing weights after seeing a result turns the criterion into a dial that produces whichever verdict is wanted, and it destroys comparability with every earlier decision. If the argument for new weights is genuinely about strategy rather than this result, it will still be a good argument at the next scheduled revision, applied prospectively.
- Why fit weights on treatment effects across past experiments rather than on user-level correlation?Because the criterion has to predict what an intervention does, not which users are already valuable. A component can correlate strongly with the long-run outcome across users while its treatment effect carries no information about the outcome's treatment effect. Fitting across the back catalogue of experiments matches the quantity the criterion is used for, which is why it is the stronger evidence.
It is like a single index built from several currencies. Nothing is meaningful until you fix the conversion rates, and if you are allowed to revise the rates after the trade settles, the index proves nothing.
saying these in an interview costs you the question
- Summing components on wildly different raw scales
- Treating weights as something the data determines
- Retuning weights after an experiment reads badly
- Assuming a composite is automatically less noisy
- Reporting only the composite and hiding the components
- Standardising using the experiment's own data