skip to content

A threshold picked to hit 85% precision on validation delivers less in production. Why?

level: seniorimportance: should knowfreq 48%

answer

  1. you chose among many noisy numbers
  2. few items clear a strict cut
  3. the winner of a sweep is flattered
  4. on the bar is a coin flip
  5. bound the estimate, then add margin

basics

~20 s

Picking the cut that first clears 85% in a sweep selects the point where sampling noise flattered you, and precision at a strict cut rests on few items. The estimate is optimistic; aim above the target on untouched data.

solid answer

~50 s

Two things compound. First, selection: you evaluated many candidate cuts and kept the one that first reached 85%, and the maximum of a set of noisy estimates is biased upward -- you picked the threshold where the noise helped. Second, sample size: precision at a strict cut is computed over only the items that clear it, so if 300 items clear it the standard error on an 85% estimate is about two percentage points, and choosing the point that just touches the bar leaves you roughly a coin flip above or below it in truth. Fixes are all about honesty of estimation: choose the cut on one slice and measure it on an untouched one, or cross-validate the cut; require a lower confidence bound of 85% rather than a point estimate; build in margin by targeting 88%; and audit live precision on a sample after launch.

go deeper

for a junior

Be ready to say that a metric measured on a sample carries error, and that precision above a strict cut is computed on only the handful of items that clear it, so the figure is shakier than it looks.

for a middle

Explain both mechanisms in your own words: selecting the best of many candidate cuts biases the winner upward, and a small number of predicted positives makes each estimate noisy. Compute a rough standard error when given the counts.

for a senior

Demonstrate the working procedure -- choose the cut on one slice and measure it on another or cross-validate it, constrain the lower confidence bound rather than the point estimate, target above the requirement, and audit live precision on sampled reviews.

for a principal

Own how commitments are made. Decide whether the organisation promises a point estimate or a bound, who pays the recall that margin costs, and what monitoring makes a broken precision promise visible in days rather than at the next quarterly review.

## The setup An auto-refund approver decides whether a customer's refund request is granted without human review. The product requirement is stated as a constraint rather than a cost: **at least 85% precision, take whatever recall that gives**. The obvious procedure is to sweep thresholds on the validation set, plot precision against the cut, and take the lowest cut whose precision reaches 0.85 -- lowest, because that maximises recall subject to the constraint. Live precision then comes in at 78%, and the constraint the product promised is broken. Nothing about the model has changed. The failure is in how the threshold was estimated. ## Cause one: the winner's curse Sweeping produces hundreds of precision estimates, one per candidate cut, and each carries sampling error. Selecting the one that first crosses a bar is a maximum-like operation over noisy quantities, and the maximum of noisy estimates is **optimistically biased**: you disproportionately select points where the noise happened to be positive. The chosen cut is therefore, on average, the cut whose true precision is below its measured precision. This is the same mechanism that makes the best hyperparameter's validation score an over-estimate of its test score. It is not overfitting the model's parameters -- the model was never touched -- it is overfitting one scalar, the threshold, to the validation sample. One scalar is far less dangerous than a million weights, which is why the effect is often mild, and it is not zero, which is why it bites at exactly the point where the constraint is tight. ## Cause two: the estimate is thin where it matters Precision at a strict cut is computed **only over the items that clear the cut**. Push the cut up and the denominator collapses. If 300 items clear it and 255 are true positives, the standard error on 0.85 is roughly `sqrt(0.85 * 0.15 / 300) = 0.021` -- two percentage points, so a 95% interval runs about 81% to 89%. Reporting "85% precision" as if it were exact hides a range that straddles the requirement. Combine this with the constraint being an *equality* at the chosen point: you deliberately selected the cut where the estimate sits right on 85%. Even with a perfectly unbiased estimator, the truth is then about equally likely to be above or below. A procedure that lands you on a coin flip is not a procedure that meets a promise. ## Cause three: the cut was tuned on data that was not clean If the threshold was chosen on the same split used for model selection, or on data that also fed feature selection, the whole estimate inherits that optimism. The threshold must be treated like any other tuned quantity: chosen on one slice, reported from another. (A fourth possibility exists -- the production population differs from validation, which moves precision for reasons that have nothing to do with how the cut was chosen. Diagnose it by checking whether the flagged volume and the score distribution also moved; if they did, you have a distribution problem rather than an estimation problem, and it needs a different fix.) ## Fixing it **Separate choosing from measuring.** Pick the cut on a tuning slice; report its precision on a slice untouched by the choice. Better on limited data, do it inside cross-validation: on each fold choose the cut that meets the constraint on the in-fold data and measure the resulting precision on the held-out part, then look at the distribution of those out-of-fold precisions. If the median is 85% but the spread runs to 79%, the constraint is not safe. **Constrain a bound, not a point.** Require the *lower* end of a 95% confidence interval on precision to clear 85%. Mechanically this shifts the cut higher and costs recall, which is precisely the price of actually keeping the promise. **Add explicit margin.** Target 88% offline for an 85% commitment, and state the margin as a deliberate choice rather than hoping. **Get more labels where the cut lives.** The estimate is thin because few items clear the cut. Over-sample the top of the score range for human labelling; a few thousand labels concentrated above the candidate region shrinks the interval far more than the same effort spread uniformly. **Smooth rather than take the first crossing.** Fitting a smooth curve of precision against the cut and reading off the crossing is more stable than taking the first raw point that pokes above the line, which is by construction a noise-selected point. **Audit live.** Sample flagged items for human review continuously and track realised precision with an interval. A promise that is only checked at launch is a promise nobody is keeping. ## What the interviewer is listening for The weak answer is "the model overfit" or "we need more data". The strong answer separates the two mechanisms -- selection bias from choosing among many cuts, and variance from a small denominator at strict cuts -- and then proposes remedies that follow from each: clean separation and cross-validation for the first, confidence bounds, targeted labelling and margin for the second.

  • How would you decide how much margin to add above the 85% target?
    From the width of the interval, not intuition. Estimate precision's standard error at the candidate cut, and set the offline target so the lower confidence bound clears 85%. With 300 items above the cut that is roughly four points of margin; with 5,000 it is under one. That also makes the cost of the promise visible -- margin is paid for in recall, and the way to buy it back is more labels near the cut.
  • Does the same selection bias apply to a threshold chosen by maximising F1?
    Yes, and more strongly, because taking an argmax over a sweep is the purest form of selecting on noise. The reported F1 at the chosen cut is optimistic, and the cut itself is displaced toward wherever the sample was lucky. Cross-validating the choice -- select on the in-fold data, score out of fold -- gives an honest number and shows how unstable the chosen cut is across resamples.
  • How do you tell this apart from the production population simply being different?
    Look at what else moved. An estimation problem shows up as precision below the offline figure while the flagged volume and the score distribution look as expected. A population change shows up as the score distribution or the flag rate shifting too. The first is fixed by re-estimating the cut honestly; the second is fixed by re-deriving it on recent data and monitoring continuously.

It is like picking the sunniest-looking week from a hundred weather forecasts and then promising a picnic: the week you picked is the one where the forecast erred optimistically.

saying these in an interview costs you the question

  • Blames model overfitting when the model never changed
  • Reports precision at a strict cut without an interval
  • Picks the first threshold that crosses the target
  • Chooses and reports the cut on the same data
  • Assumes an unbiased estimate means the promise is safe
  • Never audits realised precision after launch

context