How should dataset size change your train/validation/test split ratios?
answer
- counts, not percentages
- how tight does the estimate need to be?
- standard error shrinks with the square root
- small data loses twice over
- millions of rows: one percent is plenty
basics
~20 sRatios are a heuristic; what matters is the absolute count in each partition. With a few hundred rows the estimate is too noisy to trust, so resample instead; with millions, a one-percent hold-out is plenty.
solid answer
~50 sI size the evaluation partitions by the precision I need on the metric, then give everything else to training. Take a 900-row employee-attrition table split 60/20/20: the test set is 180 rows, so an accuracy of 0.85 carries a standard error of about `sqrt(0.85*0.15/180)` = 0.027, which is roughly plus or minus 5 points at 95%. You cannot tell an 84% model from an 88% one, and you have also thrown a third of your rows out of training, so the model is worse *and* the verdict is vaguer. There I would keep a test set and run cross-validation over the rest rather than reserve a fixed validation block. Now take a 4-million-row ad-impression log: a 1% hold-out is 40,000 rows, the same standard error falls to about 0.002, and ten-fold cross-validation would cost ten fits to buy precision you already had.
go deeper
Be ready to say that common ratios such as 80/20 or 60/20/20 are starting points, and that a few hundred rows in a test set gives a much shakier number than tens of thousands do.
Compute the standard error of the metric on the partition size you propose and use it to justify the ratio. Explain why removing rows from training hurts most when the dataset is small.
Tie the split to the decision the number supports: state the smallest difference that would change a call, size the evaluation partitions to resolve it, and defend the compute cost of resampling versus one hold-out.
Own the standing convention for the team - what evaluation precision the organisation buys by default, when a project is allowed to deviate, and how that convention trades reporting confidence against model quality and pipeline cost.
## The ratio is downstream of the count 60/20/20 and 80/10/10 are conventions from an era of small tables, not laws. What actually determines whether a partition does its job is how many rows - and, for classification, how many rows *of the class you care about* - land in it. Two projects with the same ratio can be in completely different regimes. ## The small-data regime Consider a 900-row employee-attrition table split 60/20/20. Training gets 540 rows, validation 180, test 180. Those 180 test rows carry the entire verdict. Work out the precision. For a proportion-style metric such as accuracy, the standard error is `sqrt(p*(1-p)/n)`. At `p = 0.85` and `n = 180` that is `sqrt(0.1275/180)` = about 0.027, so a 95% interval spans roughly plus or minus 5 percentage points. Any improvement smaller than that is invisible. Meanwhile the validation partition is equally noisy, so the *selection* it drives is close to a coin flip between similar candidates. The small-data regime hurts twice, and this is the part candidates usually miss: - **The estimate is vague.** Fewer evaluation rows means a wider interval. - **The model is worse.** Every row moved out of training is a row the learner never sees, and on the steep part of the learning curve those rows are worth real accuracy. There is no ratio that fixes both, which is why the standard answer at this size is structural rather than numeric: hold out a test set once, and get the selection signal by resampling over everything that is left, so no row is permanently exiled from training. The test set still has to exist, and it still has to be reported with an interval rather than as a bare decimal. ## The large-data regime Now a 4-million-row ad-impression log. Reserve 1% for validation and 1% for test: 40,000 rows each. The same accuracy calculation gives `sqrt(0.85*0.15/40000)` = about 0.0018, so the interval is a few tenths of a percentage point wide. That is far tighter than any decision you are going to make. At this size a single hold-out is not a compromise, it is the right tool. Ten-fold cross-validation would cost ten fits, and on a four-million-row log a fit is the expensive thing in the loop. It would buy a slightly lower-variance estimate of a quantity you already know to a fraction of a point. Meanwhile the 2% you removed from training is invisible on the flat part of the learning curve - the model trained on 3.92 million rows is indistinguishable from the model trained on 4 million. This is the general shape of the rule: **cross-validation buys precision with compute, and you only want that trade when precision is scarce and compute is not.** ## Sizing by precision instead of percentage A workable procedure: 1. Decide the smallest difference in the metric that would change a decision - say one percentage point of recall. 2. Solve for the `n` that makes the standard error small enough to see it. For a proportion, `n` is about `p*(1-p)/se^2`. 3. Give that many rows to test, the same or fewer to validation, and everything else to training. 4. Sanity-check the residual training size against a learning curve if you have one. For an imbalanced problem, count the minority rows rather than the rows. A test set of 40,000 rows with 100 positives estimates recall from 100 observations, and the standard error follows that 100, not the 40,000. ## Things worth saying out loud in an interview - A bigger *percentage* is not the same as a better estimate; 20% of 900 is a worse test set than 1% of 4 million. - Reporting a test score without an interval hides which regime you are in. - The validation partition and the test partition do not have to be the same size; validation only has to rank candidates, test has to pin down a number. - If the metric is not a simple proportion - an area under a curve, a mean absolute error - the arithmetic changes but the logic does not: estimate the spread, then decide whether the partition is big enough to resolve the difference you care about.
- Your test set has 180 rows and the model scores 0.85 accuracy - how precise is that number?The standard error is `sqrt(0.85*0.15/180)`, about 0.027, so a 95% interval runs roughly plus or minus 5 accuracy points. A rival scoring 0.88 is not measurably better, and any decision resting on a three-point gap is reading noise. Report the interval alongside the point estimate.
- When is a single hold-out preferable to cross-validation?When the hold-out is already large enough to be precise and fits are expensive. On a four-million-row log a 1% hold-out is 40,000 rows and pins accuracy to a few tenths of a point; ten-fold would cost ten fits for no decision-relevant gain. With a few thousand rows the trade reverses.
- Does a rare outcome change how you size the test partition?Yes - size it by the number of minority-class rows, not total rows. Recall, precision and their intervals are driven by how many positives you actually observe, so a huge test set with 60 positives still gives a wobbly recall estimate. Work backwards from the positive count you need.
saying these in an interview costs you the question
- Applies 80/20 regardless of dataset size
- Reports a test score with no sense of its uncertainty
- Assumes a larger percentage always means a better estimate
- Reserves half a million rows for a test set that needed forty thousand
- Ignores minority-class counts when sizing an imbalanced test set