A speech model shows 9% training and 10% validation word error rate, humans 6% - what next?
answer
- compare against something other than zero
- three buckets, not two
- training error minus the human floor
- the gap here is only one point
- bias remedies, and pick the best humans
basics
~20 sMost of the error is avoidable bias: three points sit between the model's 9% training error and the 6% human benchmark, against a one-point train-validation gap. Attack bias first - more capacity, richer features, weaker regularization - not more data.
solid answer
~50 sI split the 10% into three pieces against the human floor. Human transcribers reach 6% on this audio, so roughly six points are the task's irreducible difficulty - noisy call-centre lines, crosstalk, accents. The model's own training error is 9%, so three points are **avoidable bias**: error the model could remove and has not. The train-validation gap is one point, so variance is almost a rounding error. That ranking dictates the plan: increase capacity, add or improve features, loosen regularization, train longer, and run error analysis on the worst slices to find structure the model is missing. Buying more labelled audio is the wrong first purchase here - it attacks variance, and there is barely a point of variance to win. I would re-diagnose after each round, because closing the bias gap typically opens a variance gap.
go deeper
Recall that a human benchmark gives you something to compare training error against, and that error close to what humans achieve leaves little to win. Know that a tiny train-validation gap means overfitting is not the problem.
Be ready to do the arithmetic out loud: floor, training-minus-floor as avoidable bias, validation-minus-training as variance, and name which remedies attack which bucket.
Demonstrate the loop: diagnose, act on the biggest bucket, re-measure, expect the diagnosis to flip. Lead with error analysis on real failures rather than a generic list of remedies, and state the target and the floor before you start.
Own the call about whether the headroom justifies the investment at all. If the product needs 4% and the floor is 6%, the answer is to change the inputs, narrow the domain or renegotiate the promise - not to fund another modelling quarter.
## Three buckets, not two With a human benchmark in hand you stop thinking in terms of bias and variance alone and decompose the held-out error into three parts: ``` validation error 10% = irreducible floor ~6% (what expert humans also get wrong) + avoidable bias 3% (training error 9% minus the 6% floor) + variance 1% (validation 10% minus training 9%) ``` The floor is not zero. On noisy call-centre audio - overlapping speakers, compressed lines, unfamiliar accents, unclear proper nouns - expert transcribers disagree with the reference too. Chasing a 0% word error rate is chasing noise. The number worth optimising is the **avoidable** part. **Avoidable bias** is the distance from the model's *training* error to the floor: error the model makes on data it has already seen, which a capable human would not make. **Variance** is the distance from training to held-out error: how much the fit fails to transfer. Here avoidable bias is three times variance. Every remedy should be pointed at bias. ## Why human level is a usable proxy for the floor The theoretical floor is the Bayes error - the lowest error any predictor could achieve given the information in the inputs. It is unobservable. For tasks humans are good at - perception tasks especially, such as transcription, image labelling and reading a scan - human performance sits close to it and is measurable, which makes it the standard stand-in. Which humans matter. Use the **best available** benchmark, not the average one: - One typical transcriber: a floor for *replacing that person*, not for the task. - The best individual expert: a tighter, more honest floor. - A panel that discusses and adjudicates: usually the tightest, and the right number when you are deciding whether headroom still exists. Picking the average annotator inflates the floor and lets you declare victory early - a model that beats the median annotator can still be three points off what the task allows. Once a model *surpasses* human level, this whole tool stops working: you no longer have a credible estimate of the floor, progress slows, and you must fall back on held-out curves and error analysis alone. ## The bias-first plan Ranked, cheapest and fastest first: 1. **Error analysis on the failures.** Sample a hundred misrecognised utterances and label the failure mode - background noise, a specific accent, domain vocabulary, short utterances. This is hours of work and routinely rewrites the rest of the plan, because a large share of avoidable bias usually concentrates in one or two identifiable slices. 2. **Relax the constraints you imposed.** If a penalty, a depth cap or an early stop is currently active, it is holding training error up. Loosening it is free and directly reduces bias. 3. **More capacity or a richer representation.** A more expressive model family, or input features that expose information the current representation throws away. 4. **Train longer / optimise harder.** If the fit has not converged, part of the three points is optimisation error masquerading as bias - check that training error has actually plateaued before concluding the model class is too weak. 5. **Targeted data, only where analysis says so.** Not "more of everything": if the failures cluster on one accent or one noise condition, more audio *of that kind* changes the training distribution and can reduce bias on the slice that matters. ## Why buying general labelled data first is a mistake Adding representative rows attacks variance. Variance is one point of a ten-point error. Even a remedy that eliminated variance completely would leave 9%. A useful thing to say out loud in an interview: with training error at 9% and the floor at 6%, more data of the same kind cannot take you below 9% - the model already fails on data it has seen. ## Re-diagnose after every round The decomposition is not stable. Raise capacity to close bias and the train-validation gap usually widens - and *then* the data purchase becomes the right call. So the deliverable is not one action but a loop: measure the three buckets, act on the largest, re-measure. State the stopping condition too. If the product needs 8% and the floor is 6%, there are two points of headroom to fight for; if it needs 4%, no amount of modelling gets there and the honest answer is to change the input - better microphones, a narrower domain, a constrained vocabulary - or to change the product's promise.
- Whose performance would you use as the human benchmark, and why does the choice matter?The best available expert, or a panel that adjudicates disagreements - not the average annotator. The benchmark is standing in for the achievable floor, so a weak one inflates the floor and hides real headroom. A model that beats the median transcriber can still be several points short of what the task permits.
- What happens to this framework once the model beats human performance?It stops working. Human level was only ever a proxy for the unobservable floor, so once you pass it you no longer know how much avoidable bias remains, and the usual signals - error analysis, targeted slices, held-out curves - become the only guide. Progress typically slows sharply at that point.
- You close the bias gap and training error falls to 6.5%, but validation error is now 11%. What now?The diagnosis has flipped: avoidable bias is half a point, variance is 4.5. Now the data purchase you declined earlier is the right move, along with constraining the enlarged model. This is exactly why the plan is a loop - each remedy shifts error between the buckets rather than removing it uniformly.
saying these in an interview costs you the question
- Calls a one-point train-validation gap overfitting
- Buys more labelled data to fix avoidable bias
- Assumes the achievable error floor is zero
- Uses the average annotator rather than the best as the benchmark
- Never re-diagnoses after applying a remedy
- Blames distribution shift for a one-point gap without evidence