skip to content

Temporal and Group Overlap

Shuffled folds on time-ordered rows let a model train on the future, and duplicated users put the same entity on both sides. Interviewers probe the gap you leave between train and validation.

on this pageshow

questions

4

Why does a random row-level split inflate the score when one user contributes 400 sessions?

level: middleimportance: must knowfreq 68%

answer

  1. rows from one person are not independent
  2. who else appears in both sides?
  3. the model can memorise an individual
  4. which question does the score answer?
  5. split on the unit, not the row

basics

~20 s

That user's sessions land on both sides of the split, so the model can memorise the individual rather than learn a general pattern. The score then measures recall of users already seen, not performance on new ones.

solid answer

~50 s

Rows from one entity are not independent. If a power user's 400 sessions are scattered at random, the model trains on some of them and is validated on the rest, so it can lean on person-specific signal - their device, their baseline activity level, their habitual timing - instead of anything that transfers. The validation number then answers "how well do we predict a user we have already seen?", which is only the right question if production works that way. Two side effects compound it: that one user carries an outsized share of the metric, so the score partly reports how well the model fits one person; and the folds are no longer independent, so the spread across folds understates the true uncertainty. The remedy is to make the split respect the entity, keeping every row of a user wholly on one side.

go deeper

for a junior

Know that many rows can come from one person and that splitting rows at random puts that person on both sides. Be able to say why the resulting score is not evidence the model handles someone new.

for a middle

Explain the mechanism: features act as a fingerprint, the model fits the individual, and the validation number answers the within-entity question. Mention that heavy entities also dominate the aggregate metric.

for a senior

Show you check fold membership for shared identifiers as a routine assertion, that you reconstruct groups when no id exists, and that you decide the split from what production actually faces rather than by habit.

for a principal

Own the choice of what the evaluation unit is for the organisation, and require that reported metrics state whether they are seen-entity or unseen-entity numbers so teams cannot compare across incompatible protocols.

## The independence assumption the split relies on A held-out validation set estimates generalisation only if the held-out rows are, in the relevant sense, new. Random row-level splitting assumes rows are exchangeable draws. Behavioural data almost never is: one user generates hundreds of sessions, one patient many visits, one device many readings, one merchant many transactions. Those rows share a hidden cause - the entity itself. ## What the model actually learns With 400 sessions from one user scattered across folds, roughly four fifths of them sit in training whenever that user's remaining rows are being scored. The model does not need to identify the user explicitly to exploit this. Ordinary features act as a fingerprint: a rare city plus a specific device type plus a characteristic session length can single a person out almost uniquely. Even a linear model will tilt its coefficients toward whatever combination that heavy user shows, because those rows are numerous and mutually consistent. A high-capacity model does it far more aggressively, effectively storing "rows that look like this behave like that". So the model learns *this user*, and is then rewarded for it at validation time. Nothing generalises. ## Three distinct damages **1. Optimistic point estimate.** The validation score measures within-entity prediction while you intend to report cross-entity prediction. The two can differ enormously - it is entirely normal for a model that looks strong on seen users to be near-useless on new ones, because the useful signal was identity, not behaviour. **2. A metric dominated by a few entities.** Those 400 rows out of, say, 5,000 are 8% of the metric from one person. Aggregate performance now partly reports how well the model fits a handful of heavy users, and improvements that only help them look like general improvements. Reporting the metric averaged per entity as well as per row exposes this. **3. Understated uncertainty.** Fold scores agree suspiciously well, because folds share entities and are therefore correlated. The fold-to-fold spread, which people routinely use as an error bar, becomes far too tight, and a difference between two candidate models looks significant when it is not. ## When entity overlap is actually fine This is the part weaker candidates miss. The split should mirror deployment. If the model exists to score the *next* session of customers you already have - a churn score refreshed monthly on an existing base, a recommender for logged-in users - then seeing a user in training is exactly what happens in production, and forbidding it produces a needlessly pessimistic estimate. What still must hold in that case is the time direction: train on their earlier rows, validate on their later ones. "Same user, later time" is a legitimate evaluation; "same user, random rows" is not, because it lets the model see a user's future while predicting their past. If instead the model will meet strangers - onboarding risk scoring, cold-start ranking, a diagnostic deployed to a new clinic - the held-out entities must be unseen, and a row-level split answers the wrong question entirely. ## Detecting it The direct check is trivial and worth automating: intersect the entity identifiers of each pair of folds; the intersection should be empty when it is meant to be. Alongside it, count rows per entity and look at the tail - a Zipf-like distribution where the top 1% of entities hold a large share of rows is a warning that the estimate is entity-sensitive. The hard case is when there is no identifier column at all. Near-duplicates still create the same coupling: the same marketplace listing reposted four times with a tweaked title and price yields four rows whose targets are essentially one observation. If they straddle folds the model is scored on a row it has already read. Here you have to reconstruct the grouping - hash or fuzzy-match the stable fields, cluster rows that are near-identical on their features, or exploit a surrogate key such as an image fingerprint or a normalised text body - and then treat each cluster as the unit that must stay intact. ## The remedy, stated carefully The rule is: the unit of splitting must be the unit of independence. Whatever the entity is - user, patient, listing, merchant - all of its rows go to one side. Deduplicating instead is usually wrong: it discards real repeated observations that carry information, and it does not help when the duplicates are merely similar rather than identical. Nor does a different random seed help; the problem is structural, not a bad draw. One consequence to expect and accept: the honest score is lower, sometimes dramatically. That number is the one worth optimising, because it is the only one production can reproduce.

  • When is it legitimate for the same user to appear in both training and validation?
    When production works that way - a churn or recommendation model refreshed on an existing customer base sees those users again by design, so excluding them understates real performance. The requirement then shifts to time: train on the user's earlier rows and validate on their later ones. Same user with random rows is still invalid, because it shows the model a user's future.
  • There is no user id column. How do you find entity overlap anyway?
    Reconstruct the grouping from the data. Hash or fuzzy-match stable fields to find near-duplicate rows - the same listing reposted with a tweaked title, the same document resubmitted - and cluster rows that are near-identical on their features. Surrogate keys such as a normalised text body or an image fingerprint often work. Treat each cluster as one unit and keep it intact across the split.
  • How does entity overlap distort the fold-to-fold spread you quote as an error bar?
    It shrinks it. Folds sharing entities are correlated rather than independent samples, so their scores cluster far more tightly than genuine resampling variability would produce. Teams then read that narrow spread as precision and declare a small difference between two models real when it is well inside the true noise.
  • Why is deduplicating repeated rows a poor substitute for splitting by entity?
    It throws away real information - a user's hundred sessions genuinely describe their behaviour better than one does - and it fails on the common case where rows are merely similar, not identical, so the coupling survives deduplication. It also leaves the metric weighting unaddressed. Splitting on the entity fixes the leak without discarding data.

Testing a student on questions they already saw in the homework. The exam is genuinely new paper, but not new content.

saying these in an interview costs you the question

  • Says shuffling is safe because every row is independent
  • Treats a different random seed as the fix
  • Deduplicates rows instead of splitting on the entity
  • Reads a high score as proof of generalisation to new users
  • Ignores that heavy users dominate the aggregate metric

context

open as a page

Why does a 5-day-ahead return label force a gap between train and validation folds?

level: middleimportance: should knowfreq 40%

basics

~20 s

A 5-day-ahead label dated day t is computed from outcomes through day t+5, so the last training rows overlap the validation window and share its outcomes. Dropping those final five days of training removes the overlap.

open as a page

Why does a per-city average demand computed over the full three-year history leak, even with forward-chaining folds?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The aggregate summarises the entire timeline, so a row from year one already carries year three's demand. Every training row sees the future and every validation row sees its own outcomes, no matter how correctly the folds are ordered.

open as a page

When a dataset has both repeated customers and a time order, how do you design the validation split?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Start from what the model faces at inference. Break the axis production does not repeat: hold out unseen customers if it scores strangers, hold out later time if it re-scores known customers, and hold out both when it does both.

open as a page