skip to content

When a dataset has both repeated customers and a time order, how do you design the validation split?

level: principalimportance: nice to knowfreq 28%

answer

  1. two axes, four kinds of held-out row
  2. what does inference actually face?
  3. breaking one axis leaves the other
  4. unseen entities in a later window
  5. one protocol, versioned and named

basics

~20 s

Start from what the model faces at inference. Break the axis production does not repeat: hold out unseen customers if it scores strangers, hold out later time if it re-scores known customers, and hold out both when it does both.

solid answer

~40 s

The split is a claim about deployment, so decide it from deployment rather than from convention. Three regimes exist. If the model scores customers it has never met, held-out customers must be unseen. If it re-scores an existing base each month, seeing a customer in training is realistic and what must be honoured is time - train on their earlier rows, validate on their later ones. If both happen, the strictest cell is later time for customers absent from training, and it usually reports a markedly lower number than either single-axis split. Breaking only one axis leaves the other inflating the estimate, so a grouped split on time-ordered data can still be optimistic. Where the business genuinely serves both populations, report both numbers rather than averaging them into one figure nobody can act on.

go deeper

for a junior

Know that entities and time are two separate ways rows can overlap, and that a random split respects neither. Be able to ask who the model will actually score before choosing.

for a middle

Explain why breaking one axis leaves the other free to inflate the score, and describe the three deployment regimes - known customers forward, unseen customers, and both - with the split each implies.

for a senior

Show you would diagnose which axis is doing the inflating by comparing scores under each split, and that you can defend a low strict-cell number as possibly a population shift rather than a model failure.

for a principal

Own the protocol as an organisational asset: one versioned definition, named by every reported metric, so numbers stay comparable across teams and the more permissive split cannot quietly win arguments.

## Two axes, four cells When rows carry both an entity and a timestamp, any held-out row falls into one of four cells relative to training: seen entity and earlier time, seen entity and later time, unseen entity and earlier time, unseen entity and later time. A random row-level split draws mostly from the first cell - the easiest one - which is why it reports the highest and least useful number. Designing the split means choosing which cell your headline metric comes from. That choice is not a matter of rigour for its own sake; it is a claim about who the model will serve. ## Let deployment pick the cell **Only known customers, forward in time.** A monthly churn refresh, a next-best-action score for an existing base, a demand model for stores that already exist. The model genuinely will have historical rows for the entity it is scoring, so excluding them produces a pessimistic estimate of a capability the system really has. The requirement is temporal: every training row for a customer must precede the validation rows for that customer, otherwise the model sees a customer's future while predicting their past. **Only new customers.** Onboarding risk, cold-start ranking, a model shipped to a clinic or a market it has never touched. Here entity identity is worthless at runtime, so it must be worthless at evaluation time: held-out entities must be absent from training. A time-ordered split alone does not deliver this, because the same customers appear on both sides of a date cutoff. **Both, in proportion.** Most real products. The honest design evaluates the strict cell - later time, unseen entities - as the floor, and optionally reports the seen-entity forward cell alongside it. Two numbers, labelled, beat one number that silently averages populations with very different difficulty. ## The cost of the strict cell Holding out both axes shrinks the usable data twice over. Take the latest window and remove from it every customer who appears in training: what remains is a small, recency-limited, possibly unrepresentative slice, and the resulting metric is noisy. Model comparisons made on it need care, because the difference between two candidates can easily be smaller than the fold-to-fold spread. There is also a composition trap. Customers new in the final window are not a random sample of customers - they arrived through whatever acquisition channel was active then, and may be systematically different. A low score on that cell can reflect a population shift rather than a modelling weakness, and reading it as the latter sends a team off optimising the wrong thing. ## Which single axis, when you can only afford one When data is too thin for the strict cell, choose the axis whose violation would flatter you most. If entity signal is strong - long histories, distinctive behaviour, a heavy tail of power users - entity overlap is the bigger inflator and the entity axis must be broken. If the target drifts hard with time - seasonality, policy changes, a shifting base rate - temporal overlap dominates and time must be broken. Diagnose it: score once with each split and compare against a random split. The gap each one opens tells you which leak was doing the work. ## Making the choice stick The failure mode at organisational scale is not a bad split; it is many splits. Two teams quoting the same metric under different protocols produce numbers that cannot be compared, and the more permissive protocol always wins the comparison. The durable fix is to define the evaluation protocol once as a shared artifact - the axis, the boundary dates, the entity key, any gap - version it, and require that every reported metric names it. Then a change to the protocol is a reviewed event with a visible level shift, rather than an invisible drift in what the number means. Be explicit, too, about what the split does not cover. A split cannot detect a feature that reads across the whole timeline, and it cannot repair labels whose windows straddle a boundary. Those are separate defences; the split is only the one that gets the row assignment right. ## What a strong answer sounds like Name the inference-time population first, derive the split from it, state which cell the headline number comes from, and be candid that the strict cell costs precision. Volunteering that you would report two numbers when the product serves two populations is the mark of someone who has had to defend a metric to a business owner, not just compute one.

  • The strict cell leaves too little data to compare models. What do you do?
    Break the axis whose violation would flatter you most and say so. Long entity histories with distinctive behaviour mean entity overlap dominates; strong drift or seasonality means time does. Diagnose empirically by scoring under each split and against a random one - the gap each opens shows which leak carried the score. Then quote the weaker number as the estimate.
  • Why can a split that breaks only entity overlap still be optimistic on time-ordered data?
    Because held-out entities can still be drawn from the same period as training, so the model is evaluated on a world it has already seen the shape of. Drift, seasonality and shifting base rates all remain hidden. Only pushing the held-out entities into a later window tests both the population change and the calendar change together.
  • How do you stop two teams reporting the same metric under different splits?
    Define the protocol as a shared, versioned artifact - axis, boundary dates, entity key, any gap - and require every reported figure to name it. Otherwise the more permissive protocol quietly wins every comparison. Changing the protocol then becomes a reviewed event with an explained level shift rather than an unexplained jump.

saying these in an interview costs you the question

  • Picks a split by convention instead of from deployment
  • Assumes grouping by customer also handles time
  • Averages seen and unseen populations into one figure
  • Reads a weak strict-cell score as pure model weakness
  • Lets each team define its own evaluation protocol

context