skip to content

A precision-recall curve built on only 60 positives is jagged. How much do you trust its average precision?

level: seniorimportance: nice to knowfreq 30%

answer

  1. count the positives, not the rows
  2. each positive is one recall step
  3. the left end rests on few items
  4. it is an estimate with a spread
  5. resample the test set for an interval

basics

~20 s

Not much as a point value. The curve's resolution comes from the positive count, not the row count: with 60 positives each is a 1.7% recall step, so average precision carries sampling variance. Report a resampled interval, not three decimals.

solid answer

~50 s

The jaggedness is the honest picture, not a rendering artefact. Every one of the 60 positives is a single visible step of 1/60 in recall, and the precision at low recall is computed from a handful of top-ranked rows, so one item moving a few ranks swings the left end of the curve hard. Average precision is a statistic with a sampling distribution whose width is driven by the number of positives, so a 0.02 to 0.03 gap between two models here is usually noise. I would bootstrap the test set — resample rows, recompute AP, take percentiles — and report AP with an interval and the positive count beside it. For two models I would use a paired bootstrap on the difference, over a fixed test set. The structural fix is more positives: extend the evaluation window or pool repeated cross-validation folds. Adding negatives does not help.

go deeper

for a junior

Remember that a curve drawn from very few positive examples is unreliable, and that the number of positives, not the number of rows, is what to check first.

for a middle

Explain why recall moves in steps of one over the positive count, and why precision at the top of the ranking is computed from so few items that it swings on a single row.

for a senior

Show the working habits: resample for an interval, compare models with a paired procedure on a fixed evaluation set, and refuse to call a small gap an improvement until the interval on the difference excludes zero.

for a principal

Set the bar for what counts as evidence. Decide the minimum positive count an evaluation must reach before a model change can be approved, and require intervals rather than point metrics in every review.

## Where the noise comes from A precision-recall curve is built by walking down the ranked scores. Recall only changes when a **positive** is admitted, and it changes by exactly `1 / (number of positives)`. With 60 positives, recall moves in jumps of about 0.017 and the curve has at most 60 vertical resolution points, no matter whether the test set holds 6,000 rows or 6,000,000. The row count buys precision resolution; the positive count buys the curve. On top of the coarse grid, the left end is intrinsically unstable. At recall 0.05, three positives have been caught; precision there is `3 / (number flagged)`, and one negative slipping above one positive in the ranking changes it by a large fraction. Small denominators produce large relative swings. So the region of the curve people care most about — the high-precision top of the ranking — is precisely the region estimated from the fewest observations. ## Average precision is an estimate, not a measurement It helps to name the thing properly. There is a true average precision for this model on the population it will run against; what you computed is an **estimate** of it from one finite sample. Like any estimate it has a standard error, and here that standard error is governed mainly by the 60 positives. The consequence is blunt: quoting `AP = 0.412` implies a precision the data cannot support, and a leaderboard ranking two models 0.412 against 0.389 is very likely ranking noise. A second, subtler effect: with few positives the estimate is not merely noisy but skewed by which positives happen to be easy. If two of your 60 positives are near-duplicates of training rows and both land in the top ten, the left end of the curve — and therefore a disproportionate share of the AP sum — is carried by two rows. ## Quantifying the uncertainty The standard move is resampling. Draw bootstrap resamples of the test rows with replacement — stratified by label so each resample keeps a comparable positive count — recompute AP on each, and read the 2.5th and 97.5th percentiles of the resulting values as a 95% interval. With 60 positives that interval is routinely wide, often plus or minus 0.05 or more, and seeing it is the point: it converts an argument about decimals into an argument about overlap. For comparing two models, resample the **difference**, not the two numbers independently. Both models are scored on the same resampled rows in each iteration, which removes the shared difficulty of the sample and gives a much tighter interval on the delta than comparing two separately-computed intervals. If the interval on the difference straddles zero, you have not shown an improvement. A cheaper sanity check costs nothing: recompute AP with the single highest-scored true positive removed. If the number collapses, the result rests on one row. ## Getting more signal Uncertainty this large is usually a data problem, and the fixes are about positives: - **Extend the evaluation window.** Six months of history instead of one multiplies the positive count directly. Weigh this against drift: an older window may no longer describe the population you will run on. - **Pool repeated cross-validation.** Rather than one held-out split, run repeated stratified splits and either pool all out-of-fold scores into one curve or average AP across folds and report the spread across repeats. Pooling gives one curve built on every positive in the dataset. - **Do not add negatives.** Growing the test set with more negatives makes the sample bigger and the curve no smoother, because recall resolution never improves. It will, however, lower AP if the extra negatives change the prevalence — which is a reason to keep the evaluation set's base rate matched to production. - **Keep the test set fixed** across model iterations so comparisons stay paired, and accept that a fixed set gradually gets overfit by repeated selection against it. ## What to actually report A defensible report for a rare-positive detector is: average precision with a resampling interval, the number of positives and the positive rate of the evaluation set, and the curve itself drawn raw rather than smoothed. Smoothing a 60-positive curve into an elegant line communicates a confidence that does not exist. And when a stakeholder asks whether the new model beat the old one by 0.02, the correct senior answer is that the test set cannot tell — followed by a plan to get more positives.

  • How would you attach a confidence interval to an average precision value?
    Bootstrap the evaluation set: draw resamples of the rows with replacement, stratified by label, recompute average precision on each, and take the 2.5th and 97.5th percentiles of the resulting distribution. Report that interval alongside the point estimate and the number of positives the curve was built from.
  • Two candidate models differ by 0.03 average precision on this test set. Is that a real improvement?
    Probably not demonstrable. Run a paired bootstrap: on each resample score both models on the same rows and record the difference in average precision, then look at whether that distribution straddles zero. With 60 positives it usually does. Also check whether the gap comes from one or two highly-ranked items.
  • Would enlarging the test set with more negative rows make the curve smoother?
    No. Recall steps are 1 divided by the number of positives, so the curve's resolution is unchanged by adding negatives. Extra negatives can only lower precision at each recall, and if they change the base rate they make the average precision value non-comparable with the earlier one.

Sixty positives make the curve a dot-to-dot drawing rather than a photograph: the outline may be right, but any single dot in the wrong place visibly bends the shape.

saying these in an interview costs you the question

  • Quotes average precision to three decimals with no interval
  • Blames the model rather than the sample size
  • Proposes adding negative rows to smooth the curve
  • Picks the higher average precision regardless of overlap
  • Smooths the curve to make it presentable

context