Why does a model's held-out error flatten above zero no matter how many training rows you add?
answer
- not every error is fixable
- same inputs, different labels
- noise in the label given the features
- only the finite-sample part shrinks with rows
- floor is relative to the inputs you have
basics
~20 sPart of the error is irreducible: the label is not fully determined by the features, so near-identical inputs carry different labels. Data cannot remove that. The rest of the floor is the model's own bias, which rows also cannot fix.
solid answer
~40 sHeld-out error splits into three parts. **Irreducible error** is the noise in the label given the features — two radiologists shown the same ambiguous scan disagree, so no function of that scan can be right for both. **Approximation error** is the bias of the chosen model family and feature set: the best function it can express is still not the true one. **Estimation error** is the part that comes from having only a finite sample. Adding rows shrinks the third term and nothing else, so the curve descends until estimation error is negligible and then flattens on the sum of the first two. That asymptote, not zero, is the target you should be comparing against.
go deeper
Recall that some error can never be removed by data, because identical inputs can carry different labels. Be able to give one concrete example of an outcome the available inputs simply do not determine.
Explain the three-way split — irreducible noise, model-family bias, finite-sample error — and state which single term shrinks when rows are added. That mapping is what the question is really testing.
Show you measure the floor rather than assert it: double-labelling a sample, hunting conflicting near-duplicates, or benchmarking a strong human reference, and then reporting model quality relative to that number.
Own the consequence for planning: whether remaining error is unavoidable decides whether a team keeps investing or redefines the problem, and that call should be backed by a measured floor, not an intuition.
## The decomposition behind the flat line When a learning curve stops descending, it is not because the algorithm gave up. It is because only one of the three sources of error is a function of sample size, and that one has run out. **1. Irreducible error (also called the Bayes error).** For a given set of inputs, the label may not be a deterministic function of those inputs. Two records with identical feature values can carry different labels — because the outcome genuinely varies, because the measurement is noisy, or because the labelling itself is a judgement call. The canonical illustration is a triage classifier trained on chest scans: on the ambiguous ones, two qualified radiologists label differently. Whatever they disagree on, no classifier reading only that image can get right in both cases. If independent experts disagree on 6% of cases, roughly 6% error is baked in before you write a line of modelling code. **2. Approximation error (the model family's bias).** Even with infinite data and zero label noise, a model can only express the functions its form allows. A strictly additive model asked to represent an interaction, or a linear boundary asked to separate a curved one, has a best achievable function that is still wrong. This is a property of the family plus the features, not of the sample, so more rows do not touch it. **3. Estimation error.** With a finite sample you do not recover the best function in the family; you recover a noisy estimate of it. This term shrinks as rows are added — that is the entire descent you see on a learning curve. Add rows and only term 3 moves. Once it is small relative to the others, the curve flattens on `irreducible + approximation`, and every further row is bought for nothing. ## "Irreducible" is relative to your inputs The most useful nuance, and the one interviewers probe: irreducible error is **not** a fixed property of the outcome you are predicting. It is a property of the pair (outcome, available inputs). If two scans look identical to the model but the patients differ in ways recorded elsewhere — prior imaging, presenting symptoms, age — then adding those inputs makes the previously indistinguishable cases distinguishable and lowers the floor. Saying "this error is irreducible" is therefore always shorthand for **"irreducible given this representation"**. Enriching the representation is a different move from adding rows, and the learning curve, which only varies rows, cannot see it. ## Estimating the floor You can put a number on it rather than guessing: - **Double-label a random sample.** Have two independent qualified labellers annotate the same cases and count how often they differ. Their disagreement rate is a practical estimate of how much of the outcome the inputs plus human judgement fail to pin down. - **Look for near-duplicate inputs with conflicting labels.** If the same or nearly the same feature vector appears with different targets, the fraction of the data involved is a direct lower bound on unavoidable error. - **Use a strong reference performer.** How well an expert or an existing well-tuned system does on the same inputs is an empirical ceiling you can measure without theory. All three give an estimate of the *sum* of what data cannot fix — they do not, on their own, separate irreducible error from your model's bias. Separating those two requires changing the model, not the sample. ## Why this matters when reading a curve A plateau at 8% means something completely different depending on whether the floor is 1% or 7.5%. In the first case there is real headroom and the model is leaving it on the table; in the second the model is already close to as good as the inputs allow and effort belongs elsewhere. Without an estimate of the floor, a flat curve tells you only that *rows* are exhausted, not whether the *problem* is. ## Common traps - **"Enough data drives any error to zero."** False for any problem where the label is not a deterministic function of the features, which is nearly all of them. - **"Zero training error means the floor is zero."** No — a flexible model can drive in-sample error to zero by memorising, on data with any amount of noise. Only held-out error speaks to the floor. - **"Label noise just means the data is dirty."** Some of it is cleanable (typos, mis-joins, stale records). The part that comes from genuine ambiguity is not; cleaning cannot resolve a case two experts read differently. - **Confusing the floor with the model's bias.** Both are flat with respect to sample size, which is exactly why a curve alone cannot tell them apart.
- Is irreducible error a property of the problem or of your feature set?Of the pair. It is the part of the outcome the *available inputs* leave undetermined, so it is only irreducible given that representation. Add an input that distinguishes cases that previously looked identical and the floor drops. This is why the phrase is always shorthand for "irreducible given these features", and why a data-collection decision about breadth is different from one about volume.
- How would you actually estimate the irreducible error on a real task?Have two independent qualified labellers annotate the same random sample and measure how often they disagree; that rate approximates what the inputs plus human judgement cannot pin down. As a cross-check, look for near-identical feature vectors carrying conflicting labels. Both give the combined floor of what data cannot fix, not a split between label noise and model bias.
- A model reaches zero training error. Does that mean the task has no irreducible error?No. A sufficiently flexible model can drive in-sample error to zero by memorising the sample, including its noise, on a task whose floor is high. Only held-out error carries information about the floor, and a model with zero training error and poor held-out error is the classic high-variance picture rather than evidence of a clean problem.
Predicting a coin flip from a photo of the coin before it is tossed: however many photos you collect, half the outcomes stay unpredictable, because the answer was never in the picture.
saying these in an interview costs you the question
- Claims enough data drives any error to zero
- Confuses irreducible error with the model's bias
- Assumes all label noise can be cleaned away
- Treats a flat curve as proof the model is optimal
- Calls the floor fixed regardless of which features are available