skip to content

Two teams export an active-user label under different definitions — which one do you train on?

level: principalimportance: nice to knowfreq 32%

answer

  1. internally consistent, so not noise
  2. a step change on a single date
  3. same row, two answers
  4. pick from the downstream decision
  5. recompute history, or truncate to one regime

basics

~20 s

Neither, until the definition is settled. A label whose meaning changed is a specification defect, not noise: choose one definition tied to the decision the model serves, write it down with edge cases, and rebuild the target consistently.

solid answer

~40 s

The tell is that the same row is positive in one export and negative in the other while both exports are internally consistent -- so this is not random error, and no denoising method will find it. Treat it as a spec problem. First, pick the definition that matches the decision the model drives: if the score triggers a re-engagement campaign, use whichever notion of activity the campaign owner acts on, not whichever scores better in validation. Second, if the raw events allow it, recompute the target under that rule across the whole history. If they do not, restrict training to the window where the definition held and never blend the two -- mixing them teaches the model that identical behaviour is sometimes positive and sometimes negative. Then get the definition owned in writing.

go deeper

for a junior

Recognise that a label needs a written definition and that two exports of the same label can disagree. Knowing to ask what exactly counts as positive is the expected instinct here.

for a middle

Explain how a definition change shows up in the data -- a base-rate step on one date, the same key labelled both ways -- and why training on the mixture teaches the model contradictory targets.

for a senior

Show the repair path: recompute history from raw events where possible, otherwise truncate to a single regime, keep the evaluation set inside one definition, and quantify the data you give up.

for a principal

Own the choice itself. Tie the definition to the decision the prediction triggers, refuse validation score as the tiebreaker, and put a named owner, worked edge cases and a versioned change process behind it.

## Why this is not label noise Random annotator error is roughly unsystematic: individual rows are wrong, the mistakes scatter, and a model trained on the rest can often flag the offenders because they contradict their neighbours. A definition mismatch behaves in the opposite way. Every row in export A is correct *under A's rule*, and every row in export B is correct under B's. The labels are internally coherent, so out-of-fold loss ranking will not surface them and hiring more annotators will not fix anything. What you have is two different target variables wearing the same column name. ## Detecting it The symptoms are structural rather than statistical: - **A step change in the base rate.** The positive rate jumps from 12% to 31% on a single date. Real behaviour moves smoothly and seasonally; a cliff on one day is a rule change, a pipeline change, or a backfill. - **The same entity labelled both ways.** Join the two exports on the key. If the same user-day is positive in one and negative in the other, the disagreement is definitional. - **A backtest that flatters and a launch that disappoints.** If training and validation share a definition that production no longer uses, the offline number is measuring something the live system does not compute. - **Nobody can produce the written rule.** Ask two people to define the label independently. If the answers differ in the edge cases -- does a user who only opened a notification count? does a background sync count? -- the definition does not exist yet. ## Choosing the definition The decision rule is: **the definition that matches the action the prediction triggers.** A model whose score routes users into a win-back campaign should predict the notion of activity the campaign owner uses to judge success, otherwise the model and the metric it is graded on will forever disagree. Validation score is the wrong tiebreaker -- a looser definition usually produces a more balanced, easier target and a better-looking number while modelling something nobody asked for. If two consumers genuinely need different notions of activity, that is two targets, and the honest answer is two models or one model with two thresholds, not a compromise definition that serves neither. ## Repairing history Once the definition is chosen, there are three cases. 1. **Raw events survive.** Recompute the label from the underlying events across the whole history under one rule. This is the clean outcome and is worth real effort to reach, because it recovers all your data. 2. **Raw events are gone for the old period.** Truncate: train only on the window where the definition is consistent. You lose data, and that is the correct trade -- a model trained on contradictory targets is bounded by the contradiction. Quantify the loss before arguing about it; sometimes the recent window is plenty. 3. **You must use both periods.** Then be explicit: include a definition-regime indicator, keep the evaluation set inside a single regime, and check whether the indicator carries real weight. If it does, the two periods are teaching different things and you are back to case 2. What you must not do is average, union or intersect the two definitions to keep row count up. That manufactures a target that no consumer will act on and puts an artificial ceiling on the achievable score. ## Preventing the recurrence One definition, one written owner, and worked edge cases -- not a sentence but a list of decided cases, because the ambiguity always lives in the edges. The definition should serve both the model target and the number the business reads on a dashboard; when those come from different rules, the model will be blamed for a discrepancy it did not create. Add a monitor on the label's base rate so the next step change is caught in days rather than in a post-mortem. And when the definition genuinely must change -- products change, so it will -- treat it as a versioned change with a stated cutover date and a plan to recompute history, rather than a silent edit to a query. ## The judgement being tested An interviewer asking this wants to see three things: that you recognise a systematic definition problem rather than reaching for a denoising technique; that you pick the definition from the downstream decision rather than the validation score; and that you are willing to give up rows to keep the target coherent, and can say what that costs.

  • You cannot recompute the old labels because the raw events are gone. Now what?
    Truncate to the window where the definition held and accept the smaller dataset -- a model trained on contradictory targets is capped by the contradiction. Quantify the cost first: often the recent regime has enough data to settle the argument. If you must use both periods, add a regime indicator and keep the evaluation set entirely inside one regime, then check whether the indicator carries real weight.
  • How is this different from ordinary annotator noise, in practice?
    Annotator noise scatters, so contradicted rows stand out against their neighbours and out-of-fold loss ranking surfaces them. A definition mismatch is internally consistent within each export, so nothing flags it and extra annotators change nothing. You detect it structurally instead -- a base-rate step change on a date, the same key labelled both ways in two sources, or two colleagues defining the edge cases differently.
  • Who should own the label definition, and what does ownership actually mean?
    One named owner on the consuming side, because the definition encodes a business decision, not a modelling preference. Ownership means a written rule with worked edge cases, a versioned change process with a cutover date and a plan to recompute history, and a monitor on the label's base rate. The modelling team should never invent its own variant quietly -- the dashboard number and the training target must come from the same rule.

saying these in an interview costs you the question

  • Blends both definitions to keep the row count up
  • Calls it label noise and reaches for a denoising method
  • Picks whichever definition scores better in validation
  • Reads the base-rate jump as real user behaviour
  • Adds more annotators to fix a definition problem
  • Settles on a definition and leaves it undocumented

context