Why does a card-not-present fraud model retrained on the last 30 days of checkouts learn that recent traffic is almost fraud-free?
answer
- the clock has not run out yet
- absence of a dispute is not innocence
- labels mature; rows split by age
- measure the maturity curve, then cut
- three states: fraud, legitimate, unknown
basics
~10 sA settled dispute confirms fraud weeks after the checkout, so recent rows carry no dispute yet. Joining absence-of-dispute to a legitimate training label marks immature fraud as good and deflates the recent fraud rate.
solid answer
~50 sThe truth about a checkout arrives on the cardholder's timetable: a dispute is filed and settled weeks to months after the decision. So in any snapshot, rows are split by **age**, not by quality — an old row has had time to be disputed, a young one has not. If the training job treats "no dispute row exists" as a legitimate training label, every immature fraud in the recent window enters as a confirmed good example, and the newest weeks — the ones carrying the current attack pattern — are the most corrupted. The fix is a **training cut** at the measured maturation window: only decisions older than that age get a final label, younger ones get an explicit `UNKNOWN` state rather than a negative. The price is that the model always trains on a world several weeks old.
code
pseudocode · 15 linesMATURATION_DAYS = 45
for each decision in decisionLog:
age = trainingCutDate - decision.decidedAt
if age < MATURATION_DAYS:
label = UNKNOWN // clock still running: exclude, never call it good
else if decision.hasSettledDispute(reason = UNAUTHORISED_USE):
label = FRAUD
else if decision.action == DECLINE:
label = UNKNOWN // blocked: no outcome can arrive, at any age
else:
label = LEGITIMATE
trainingRows = rows where label != UNKNOWNgo deeper
Remember the shape: the decision happens now, the confirmation arrives weeks later, so recent rows are unfinished rather than clean.
Explain the mechanics — the maturity curve, the training cut at the measured age, and the third label state that keeps immature rows out of the negative class.
Show you have operated it: measuring the curve on your own history, spotting a fraud rate that fell only because the window moved, and paying the staleness the cut costs.
Own the tradeoff between label completeness and freshness, and decide how much staleness the business will carry before it buys labels another way.
## Why the newest rows look clean A card-not-present checkout is decided in milliseconds; the truth about it arrives on the cardholder's timetable. The cardholder notices the charge on a statement, disputes it with the issuer, the issuer raises the case, and it eventually settles for or against the merchant. Most disputes land weeks after the checkout, and a thin tail lands months after. That gap is the **maturation window**: the age at which a decision's outcome can be treated as final. Any table of past checkouts is therefore split by age rather than by quality. Old rows have had time to be disputed and are informative either way. Young rows have not, and the absence of a dispute on them means only that the clock has not run out. Joining "no dispute row exists" to a legitimate training label puts every immature fraud from the recent window into the training set as a confirmed good transaction. The damage is concentrated exactly where it hurts most: - The **measured fraud rate** on recent weeks is deflated, and it reads as a win rather than as a censoring artefact. - The mislabelled rows carry the *current* attack pattern — the rows you most wanted the next model to learn from. - The model is taught that the newest behaviour is safe, which is the lesson an attacker would happily write himself. - Everything downstream that reads the same table — a base-rate estimate, a loss forecast, a reviewer's benchmark — inherits the deflation. ## The maturity curve, and where to cut The window is not a constant to be guessed; it is a curve you measure from your own history. For decisions old enough to be finished, plot the share of their eventual disputes that had been filed by each age. The curve rises steeply and then flattens, and the cut goes where it flattens. | Age of the decision | Share of its eventual disputes already filed | What the row may be used for | |---|---|---| | 0-7 days | a small minority | input-distribution work only; no training label | | about 30 days | a clear majority on many merchant curves | confirmed positives are usable; negatives are not | | 45-60 days | the large majority | usable, ideally weighted for residual censoring | | 90 days and older | effectively all of them | a settled label, usable as it stands | Two things move that curve and both are worth re-measuring rather than assuming: the mix of goods you sell (delivery delay pushes disputes later) and the mix of card issuers, since issuers differ in how quickly they raise a case. ## Three label states, not two The structural fix is to stop modelling the label as a boolean. A decision carries one of three states: 1. **Fraud** — a dispute for unauthorised use has settled against you. 2. **Legitimate** — the decision is older than the maturation window and nothing arrived. 3. **Unknown** — either it is younger than the window, or it was declined, so no outcome can arrive at all. The training job selects rows where the state is not `UNKNOWN`. That single rule prevents both the immature-fraud deflation and the separate problem of blocked traffic, which never earns an outcome no matter how long you wait. ## What the cut costs, and the honest mitigations The cut is not free. With a 45-day window, the freshest example a supervised model can learn from describes a world 45 days old, and against an adapting attacker that is a real handicap rather than a rounding error. Teams reach for three mitigations, and only the first two are safe: - **Inverse-maturity weighting.** Keep younger rows, but weight the observed positives up by the reciprocal of the share of disputes expected to have arrived by that age. This corrects the *rate*, not the identity of which specific rows will turn out to be fraud, so it helps a base-rate estimate more than it helps the classifier. - **Faster proxy signals.** An analyst verdict or a customer report arrives in hours or days and can stand in as a provisional label, as long as the record keeps it distinguishable from a settled one and the provisional value is allowed to flip later. - **Shortening the window because retraining is weekly.** This is the wrong direction — it lets the retraining schedule choose the statistics, and it reintroduces exactly the deflation the cut exists to prevent. ## The two ways teams get it wrong The first is silence: no cut at all, immature rows labelled negative, and a fraud rate that appears to be improving every week. The second is over-correction: a 180-day window chosen for safety, which buys a few percent of extra label completeness at the cost of half a year of staleness. Both are decided by measuring the curve, and by revisiting it when the product mix changes.
- How do you choose the maturation window rather than guessing it?Take decisions old enough to be finished, and plot the share of their eventual disputes filed by each age. Cut where the curve flattens — typically the age holding the large majority of eventual disputes. Re-measure when the goods mix or delivery times change, since delayed delivery pushes disputes later.
- Can rows younger than the window be used for anything at all?Yes, for input-distribution work, and for base-rate estimates with inverse-maturity weighting that scales observed positives by the reciprocal of the share expected so far. What they must never supply is confirmed negatives, because a young row with no dispute is censored rather than clean.
- The same transaction is labelled legitimate in one snapshot and fraud in the next. Is that a pipeline bug?No — it is the normal life of a delayed label. It does mean a training set is identified by its time window *and* the as-of date its labels were read at, so two models trained on the same month weeks apart genuinely saw different labels and are not comparable without that stamp.
It is like judging a restaurant's hygiene from this week's complaint log. Nobody has complained yet about last night's dinner, and that silence is the incubation period, not a clean kitchen.
saying these in an interview costs you the question
- Treating a recent undisputed checkout as a confirmed legitimate example
- Reading a falling fraud rate on the last two weeks as a real improvement
- Choosing the training cut from the retraining schedule, not the maturity curve
- Assuming a longer maturation window is free because it is safer
- Calling a later relabel of the same transaction a data-quality defect