skip to content

An analyst cleared a checkout as legitimate and a dispute for unauthorised use settles against you six weeks later — which becomes the training label?

level: seniorimportance: should knowfreq 50%

answer

  1. several claims, arriving at different times
  2. settled outweighs provisional
  3. only unauthorised-use is a fraud label
  4. settled, not merely filed
  5. window plus label as-of date

basics

~10 s

The settled unauthorised-use dispute becomes the training label; the analyst verdict was a provisional stand-in. Keep both on an append-only record, because the disagreement rate measures the review policy rather than the model.

solid answer

~50 s

Label sources in this archetype form a ladder of latency against fidelity: a rule flag is instant and weak, an **analyst verdict** arrives in hours, a customer report in days, a **settled dispute** in weeks. When a settled unauthorised-use dispute contradicts an earlier analyst verdict, the settlement wins as the training label — but only for that reason class, and only once settled, since a dispute won at representment flips the row back to legitimate and a dispute over undelivered goods is not a fraud label at all. The verdict is not deleted: it stays on an append-only record beside the settlement, because a rising analyst-versus-settlement disagreement rate says the review guidance is drifting. And because labels change under you, a training set is identified by its time window *and* the as-of date the labels were read.

code

json · 19 lines
json
{
  "decisionId": "d-9f31c2",
  "decidedAt": "2026-07-02T11:04:19Z",
  "action": "allow",
  "scoreAtDecision": 0.62,
  "policyVersion": "risk-policy-2026-06-18",
  "labelClaims": [
    { "source": "analyst_review", "state": "provisional",
      "value": "legitimate", "observedAt": "2026-07-02T15:40:00Z" },
    { "source": "dispute", "state": "filed", "reasonClass": "unauthorised_use",
      "value": "fraud", "observedAt": "2026-08-01T08:22:00Z" },
    { "source": "dispute", "state": "settled", "reasonClass": "unauthorised_use",
      "value": "fraud", "observedAt": "2026-08-13T16:25:00Z" }
  ],
  "trainingLabel": {
    "asOf": "2026-09-15", "value": "fraud",
    "source": "dispute", "state": "settled"
  }
}

go deeper

for a junior

Remember that one transaction can collect several claims about what it was, arriving hours to months apart, and they will not always agree.

for a middle

Explain the ladder from rule flag to analyst verdict to settled dispute, and why only settled unauthorised-use cases become the fraud training label.

for a senior

Design the record: append-only claims with source, state and observation time, plus the as-of stamp that makes a dataset reproducible when labels keep moving.

for a principal

Decide when the business trains on provisional labels for speed, and who owns the review policy that the disagreement rate is actually measuring.

## A ladder of label sources One checkout can acquire several claims about what it was, arriving at different times and carrying different authority. | Source | Typical latency | Authority | What it is really saying | |---|---|---|---| | Rule or heuristic flag | Instant | Weakest | This matched a pattern somebody wrote | | Analyst verdict on a reviewed case | Hours | Provisional | A trained human, working from what was visible then | | Customer report of unauthorised use | Days | Strong but partial | The cardholder says they did not do it | | Filed dispute | Weeks | Not yet final | A case has been raised and may still be contested | | Settled dispute | Weeks to months | Final for training | The case resolved, and on what reason class | The ladder matters because a training pipeline that pools these into one boolean loses the ability to say why a row is positive, and loses the ability to notice when the cheap sources disagree with the expensive one. ## Which one wins, and the three qualifications The settled outcome wins as the training label. Three qualifications keep that from becoming an overstatement: 1. **Reason class matters.** A dispute filed because goods never arrived, or because the quality was wrong, is a service failure, not fraud. Only the unauthorised-use class belongs in the fraud positive class; pooling the rest teaches the model to predict logistics problems. 2. **Settled, not filed.** A filed dispute can be contested and won, and a win restores the row to legitimate. Training on filed disputes imports every reversal as a false positive. 3. **The verdict is not wrong, it is earlier.** The analyst decided from evidence available within hours. Losing to a settlement six weeks later is the normal outcome of an information gap, not a performance failure. ## What to do with the disagreement The disagreement is a signal in its own right, about the **review policy** rather than the model. Watch it as a rate, sliced by reviewer cohort and by the guidance version they were working under: - A rising rate of *cleared then disputed* means reviewers are being asked to clear too readily, or the evidence surfaced in the case view has stopped being sufficient. - A rising rate of *flagged then never disputed* is harder to read, because a flagged case is often actioned and so prevented — which is exactly the censoring problem again, in miniature. - A stable low rate is the healthy state, and it is what lets analyst verdicts be used as provisional labels at all. ## Recording it so the training set stays honest The record is append-only: the decision, then each label claim with its source, its state and the time it was observed. Nothing is overwritten, because the sequence is what lets you reconstruct what any past model was trained on. Two consequences follow that teams discover late: - **A dataset is a window plus an as-of date.** "Last July's checkouts" read in September and read again in November are different label sets. Two models compared without that stamp are not comparable. - **Provisional training is a choice with a cost.** Training on analyst verdicts lets you refresh weeks earlier than settlements allow, at the price of a label that can flip underneath a deployed model. A reasonable posture is to train on provisional labels only where the refresh urgency is real, and to keep the settled-label model as the reference the provisional one is measured against. ## The failure this prevents The failure mode is a pipeline that writes a single `isFraud` column and updates it in place. Everything downstream then sees one number with no source, no state and no time. A backfill silently rewrites history, a metric moves for reasons nobody can trace, an analyst's judgement and a settlement become indistinguishable, and nobody can answer the simplest audit question about a model in production: which labels was it actually trained on, and when were they true?

  • Should a filed dispute be used as a training label to save a few weeks?
    Only with its state recorded, and knowing it can reverse. A contested dispute won at representment means the transaction was legitimate after all, so training on filed disputes imports every reversal as a positive. If the weeks genuinely matter, train on filed disputes as a separate provisional source and measure the model against the settled-label version.
  • Why does a training set need a label as-of date as well as a time window?
    Because labels keep changing for a window that has already closed. The same month of checkouts read two months apart carries different positives, so a dataset identified only by its window is ambiguous and two models built on it are not comparable without the as-of stamp.
  • What does a rising rate of analyst-cleared-then-disputed cases tell you?
    That the review side is drifting, not the model. Either reviewers are being pushed to clear faster, the evidence in the case view no longer distinguishes the current attack, or the guidance is out of date. It is fixed in review policy and tooling, and only shows up if both label claims are kept.

saying these in an interview costs you the question

  • Overwriting the label in place instead of appending a new claim
  • Treating any dispute as a fraud label regardless of reason class
  • Training on filed disputes without recording that they can reverse
  • Discarding the analyst verdict once a settlement arrives
  • Identifying a training set by its time window alone