Leave-one-store-out CV scores 0.71 where random 10-fold scores 0.94 — which do you trust?
answer
- the gap is information, not a bug
- who will the model score tomorrow?
- new branches or the same forty?
- report both, label each clearly
- read the per-store spread, not the mean
basics
~10 sThe two schemes answer different questions: random 10-fold describes new rows from branches already in training, leave-one-store-out describes a branch never seen. Report whichever matches who the model will actually score.
solid answer
~50 sThe 23-point gap is a measurement, not a bug: it says most of the model's apparent accuracy comes from branch-specific patterns it learned from that branch's own history. Which estimate is honest depends on the deployment population. If the model will only ever score the same 40 branches, all of which contributed training rows, the random estimate is the closer analogue of production and the grouped one is a pessimistic stress test. If the company opens new branches and expects the model to work on day one there, leave-one-store-out is the honest headline and 0.94 is a number you must not quote. In practice I report both, labelled by the question each answers, and I look at the per-store held-out scores rather than only their mean, because a bimodal spread usually means a few atypical branches are carrying the whole gap.
go deeper
Know that a grouped split usually scores lower than a random one, and that this is expected rather than a mistake. Being able to say the two measure different things is enough at this level.
Explain what each scheme estimates — new rows from known entities versus a wholly new entity — and why the size of the gap tells you how much the model depends on entity-specific signal.
Show you drive the choice from the deployment population, report both estimates with explicit labels, examine the per-entity score distribution, and re-run the grouped scheme after adding entity-describing features to test whether the gap actually narrowed.
Own what the organisation is promised: which estimate goes into the business case, how a rollout to unseen entities is gated, and how model selection is standardised so that nobody picks a winner under a scheme that does not match production.
## Two schemes, two questions A cross-validation scheme is a claim about what 'new data' means. Random 10-fold over rows from 40 retail branches holds out a random 10 percent of rows; every held-out row comes from a branch whose other rows are in training. Its estimate answers: how well do we predict a new row from a branch we already know? Leave-one-store-out holds out an entire branch at a time. Its estimate answers: how well do we predict for a branch we have never trained on? Both are legitimate estimates of genuinely different quantities. Reading the gap as 'one of them is broken' is the mistake. When 0.94 falls to 0.71, that 23-point drop is quantifying how much of the model's performance rests on branch-specific structure — a branch's baseline volume, its assortment, its local seasonality, the quirks of how its staff enter data — which is available for known branches and unavailable for a new one. ## Deciding which is the headline The deciding question is who the model scores in production, and it is a product question rather than a statistical one. If the answer is 'the same 40 branches, indefinitely, and we retrain monthly on their fresh data', then every scored row at inference time comes from a branch richly represented in training. The random estimate mirrors that. Publishing 0.71 as the expected performance would understate the system by a wide margin and might kill a project that would in fact work. If the answer is 'we open eight branches a year and the model must work there from week one', then a new branch's rows look to the model exactly like a held-out store in leave-one-store-out. 0.71 is the number, and quoting 0.94 to a stakeholder is a promise the system cannot keep. Most real answers are mixed: mostly known branches, occasionally a new one. The response is to report both with explicit labels — 'known-branch performance' and 'new-branch performance' — rather than to average them into a single meaningless figure. Averaging two estimates of different quantities produces an estimate of nothing. ## Look at the distribution, not the mean With 40 branches, leave-one-store-out gives 40 held-out scores. The mean hides everything interesting. Two very different worlds produce a mean of 0.71: a tight cluster of branches all scoring near 0.71, meaning the model is uniformly weaker on unseen branches; or 33 branches near 0.90 and 7 near 0.20, meaning the model transfers fine except to a distinct minority. The second case is actionable — inspect what those seven have in common (size, region, format, data quality, a recent refit) and either add features that encode that variation or exclude that segment from the model's remit. Unequal branch sizes matter here too. If the largest branches also score best, the unweighted mean of per-store scores and the metric pooled over all out-of-fold predictions will disagree, and you should be able to say which you are quoting. ## What the gap changes about the model The gap is also a design signal. A large one says the model leans on identity-level signal. If new branches matter, the response is to make the branch representable by its properties rather than by its identity — floor area, format, region, catchment, weeks since opening — so a new branch inherits behaviour from similar known branches instead of being a blank. Then re-run leave-one-store-out and see whether the gap narrowed; that is the direct test of whether the change did what it was meant to. Model choice should be made under the scheme that matches deployment. A candidate that wins under random 10-fold and loses under leave-one-store-out is exploiting branch-specific memorisation. That is a perfectly sensible thing to do when the branch set is fixed, and a trap when it is not — and it is why running the selection under the wrong scheme quietly picks the wrong model. ## What not to do Do not quote whichever number is higher. Do not treat the grouped scheme as automatically more rigorous and therefore always correct — when deployment genuinely only sees known entities, grouped CV answers a question nobody asked. Do not blend the two. And do not report the grouped mean without the spread, because with entity-level folds the spread is where the risk lives.
- How do you check whether a few stores are driving the whole gap?Look at the 40 individual held-out scores rather than their mean. A tight cluster near 0.71 means uniformly weaker transfer; 33 branches near 0.90 and 7 near 0.20 means a distinct minority is carrying it. Then find what those seven share — size, format, region, data quality, recent opening — and decide whether to add features encoding that variation or to exclude the segment from the model's remit.
- Deployment only ever covers the existing 40 branches. Is the grouped estimate then useless?Not useless, but not the headline. It becomes a stress test: it bounds how the model would behave if a branch changed character or a new one opened, which is a real robustness question. The primary estimate should mirror deployment, so the known-branch scheme leads and the grouped number is reported alongside as a downside scenario rather than as the expected performance.
- How should the gap influence which of two candidate models you pick?Select under the scheme that matches deployment. If new branches must be served, choose on the leave-one-store-out ranking even when a rival wins by a wide margin under random folds — that rival is winning by memorising branches it will not have. If the branch set is fixed, the reverse holds and penalising branch-specific signal throws away accuracy you are entitled to.
saying these in an interview costs you the question
- Reports whichever of the two numbers is higher
- Calls the grouped score a bug in the splitting code
- Assumes grouped CV is always the correct estimate to report
- Averages the grouped and random estimates into one figure
- Quotes the grouped mean without the per-store spread