Why does a random row-level split inflate the score when one user contributes 400 sessions?
answer
- rows from one person are not independent
- who else appears in both sides?
- the model can memorise an individual
- which question does the score answer?
- split on the unit, not the row
basics
~20 sThat user's sessions land on both sides of the split, so the model can memorise the individual rather than learn a general pattern. The score then measures recall of users already seen, not performance on new ones.
solid answer
~50 sRows from one entity are not independent. If a power user's 400 sessions are scattered at random, the model trains on some of them and is validated on the rest, so it can lean on person-specific signal - their device, their baseline activity level, their habitual timing - instead of anything that transfers. The validation number then answers "how well do we predict a user we have already seen?", which is only the right question if production works that way. Two side effects compound it: that one user carries an outsized share of the metric, so the score partly reports how well the model fits one person; and the folds are no longer independent, so the spread across folds understates the true uncertainty. The remedy is to make the split respect the entity, keeping every row of a user wholly on one side.
go deeper
Know that many rows can come from one person and that splitting rows at random puts that person on both sides. Be able to say why the resulting score is not evidence the model handles someone new.
Explain the mechanism: features act as a fingerprint, the model fits the individual, and the validation number answers the within-entity question. Mention that heavy entities also dominate the aggregate metric.
Show you check fold membership for shared identifiers as a routine assertion, that you reconstruct groups when no id exists, and that you decide the split from what production actually faces rather than by habit.
Own the choice of what the evaluation unit is for the organisation, and require that reported metrics state whether they are seen-entity or unseen-entity numbers so teams cannot compare across incompatible protocols.
## The independence assumption the split relies on A held-out validation set estimates generalisation only if the held-out rows are, in the relevant sense, new. Random row-level splitting assumes rows are exchangeable draws. Behavioural data almost never is: one user generates hundreds of sessions, one patient many visits, one device many readings, one merchant many transactions. Those rows share a hidden cause - the entity itself. ## What the model actually learns With 400 sessions from one user scattered across folds, roughly four fifths of them sit in training whenever that user's remaining rows are being scored. The model does not need to identify the user explicitly to exploit this. Ordinary features act as a fingerprint: a rare city plus a specific device type plus a characteristic session length can single a person out almost uniquely. Even a linear model will tilt its coefficients toward whatever combination that heavy user shows, because those rows are numerous and mutually consistent. A high-capacity model does it far more aggressively, effectively storing "rows that look like this behave like that". So the model learns *this user*, and is then rewarded for it at validation time. Nothing generalises. ## Three distinct damages **1. Optimistic point estimate.** The validation score measures within-entity prediction while you intend to report cross-entity prediction. The two can differ enormously - it is entirely normal for a model that looks strong on seen users to be near-useless on new ones, because the useful signal was identity, not behaviour. **2. A metric dominated by a few entities.** Those 400 rows out of, say, 5,000 are 8% of the metric from one person. Aggregate performance now partly reports how well the model fits a handful of heavy users, and improvements that only help them look like general improvements. Reporting the metric averaged per entity as well as per row exposes this. **3. Understated uncertainty.** Fold scores agree suspiciously well, because folds share entities and are therefore correlated. The fold-to-fold spread, which people routinely use as an error bar, becomes far too tight, and a difference between two candidate models looks significant when it is not. ## When entity overlap is actually fine This is the part weaker candidates miss. The split should mirror deployment. If the model exists to score the *next* session of customers you already have - a churn score refreshed monthly on an existing base, a recommender for logged-in users - then seeing a user in training is exactly what happens in production, and forbidding it produces a needlessly pessimistic estimate. What still must hold in that case is the time direction: train on their earlier rows, validate on their later ones. "Same user, later time" is a legitimate evaluation; "same user, random rows" is not, because it lets the model see a user's future while predicting their past. If instead the model will meet strangers - onboarding risk scoring, cold-start ranking, a diagnostic deployed to a new clinic - the held-out entities must be unseen, and a row-level split answers the wrong question entirely. ## Detecting it The direct check is trivial and worth automating: intersect the entity identifiers of each pair of folds; the intersection should be empty when it is meant to be. Alongside it, count rows per entity and look at the tail - a Zipf-like distribution where the top 1% of entities hold a large share of rows is a warning that the estimate is entity-sensitive. The hard case is when there is no identifier column at all. Near-duplicates still create the same coupling: the same marketplace listing reposted four times with a tweaked title and price yields four rows whose targets are essentially one observation. If they straddle folds the model is scored on a row it has already read. Here you have to reconstruct the grouping - hash or fuzzy-match the stable fields, cluster rows that are near-identical on their features, or exploit a surrogate key such as an image fingerprint or a normalised text body - and then treat each cluster as the unit that must stay intact. ## The remedy, stated carefully The rule is: the unit of splitting must be the unit of independence. Whatever the entity is - user, patient, listing, merchant - all of its rows go to one side. Deduplicating instead is usually wrong: it discards real repeated observations that carry information, and it does not help when the duplicates are merely similar rather than identical. Nor does a different random seed help; the problem is structural, not a bad draw. One consequence to expect and accept: the honest score is lower, sometimes dramatically. That number is the one worth optimising, because it is the only one production can reproduce.
- When is it legitimate for the same user to appear in both training and validation?When production works that way - a churn or recommendation model refreshed on an existing customer base sees those users again by design, so excluding them understates real performance. The requirement then shifts to time: train on the user's earlier rows and validate on their later ones. Same user with random rows is still invalid, because it shows the model a user's future.
- There is no user id column. How do you find entity overlap anyway?Reconstruct the grouping from the data. Hash or fuzzy-match stable fields to find near-duplicate rows - the same listing reposted with a tweaked title, the same document resubmitted - and cluster rows that are near-identical on their features. Surrogate keys such as a normalised text body or an image fingerprint often work. Treat each cluster as one unit and keep it intact across the split.
- How does entity overlap distort the fold-to-fold spread you quote as an error bar?It shrinks it. Folds sharing entities are correlated rather than independent samples, so their scores cluster far more tightly than genuine resampling variability would produce. Teams then read that narrow spread as precision and declare a small difference between two models real when it is well inside the true noise.
- Why is deduplicating repeated rows a poor substitute for splitting by entity?It throws away real information - a user's hundred sessions genuinely describe their behaviour better than one does - and it fails on the common case where rows are merely similar, not identical, so the coupling survives deduplication. It also leaves the metric weighting unaddressed. Splitting on the entity fixes the leak without discarding data.
Testing a student on questions they already saw in the homework. The exam is genuinely new paper, but not new content.
saying these in an interview costs you the question
- Says shuffling is safe because every row is independent
- Treats a different random seed as the fix
- Deduplicates rows instead of splitting on the entity
- Reads a high score as proof of generalisation to new users
- Ignores that heavy users dominate the aggregate metric