Per-class AP on a 12-instance defect class swings 8 points on one missed box - how do you report it?
answer
- recall can only move in twelfths
- sampling noise, not model variance
- the mean weights all classes equally
- report counts and an interval
- more thresholds is not more evidence
basics
~20 sWith 12 instances, recall moves in steps of about 8 points, so that class's AP is a coarse, high-variance statistic. Publish per-class AP beside its instance count with an uncertainty estimate, and never let the headline mean carry a ship decision it cannot support.
solid answer
~50 sThe number is not wrong, it is unresolvable: with 12 ground-truth instances, recall can only take the values 0/12, 1/12, ... so a single box flipping moves AP by roughly 8 points, and no amount of careful modelling changes that quantisation. Two consequences drive the reporting. First, mAP averages classes with equal weight, so this class contributes the same as a class with thousands of instances and injects 8/N points of noise straight into the headline number. Second, comparisons between runs on that class are meaningless below the noise floor. So: always publish per-class AP with instance counts alongside it, attach an interval from a bootstrap over images, define a minimum evaluation-set size before a class is allowed to gate a release, and if the rare defect genuinely matters, give it its own larger evaluation slice and its own acceptance criterion rather than burying it in a mean.
go deeper
Recall that a metric computed from a dozen examples is unreliable, and that per-class scores should always be read next to how many instances of that class the evaluation set contains.
Explain the mechanics: recall is quantised in units of one over the instance count, so with 12 instances the smallest possible move is about 8 points, and an unweighted mean over classes passes that noise into the headline number.
Show the working practice - per-class tables with counts, bootstrap over images for intervals, paired comparisons between runs - and be able to say when a reported improvement sits below the noise floor of the evaluation set.
Own the evaluation contract. Decide which classes carry business risk, size their evaluation slices so their floors are measurable, set the effect size that counts as real before the experiment, and resist a headline mean being used for decisions it cannot support.
## Why the number is coarse Average precision summarises a precision-recall curve, and recall has denominator equal to the number of ground-truth instances of the class. With 12 instances, recall is confined to the 13 values 0, 1/12, 2/12, ... 1. There is no such thing as a small change in recall for this class - the smallest possible move is 8.3 points. Precision at each of those points is estimated from a similarly small number of detections. A single borderline box that slips below the IoU threshold, or a single annotation the labeller drew a little differently, moves the class's AP by several points. The key framing for an interview: **this is sampling noise, not model variance.** Retraining with a different seed will move it. So will re-splitting the data. So will a slightly different annotation of one image. Reporting a point estimate to two decimal places implies a precision the evaluation set cannot deliver. ## How it contaminates the headline mAP is an unweighted mean over classes. That choice is deliberate - it stops a few common classes from hiding total failure on the rest - but it has a direct consequence here: a class with 12 instances carries exactly the same weight as one with 5,000. If a change flips one box on the rare class, the mean moves by roughly 8/N points for N classes. On a 20-class problem that is a 0.4-point swing in the headline metric from a single box, which is the same order as the improvements teams routinely celebrate. So when someone reports that mAP went up 0.4 after a change, the first question is: did the common classes move at all, or is this one rare class's coin flip? ## What to actually do **Publish the breakdown, always.** Per-class AP next to that class's instance count in the evaluation split, in the same table. A reader who sees `n = 12` calibrates instantly; a reader who sees only the mean cannot. **Attach uncertainty.** Bootstrap over *images* (not over detections - detections within an image are not independent), recomputing the whole matching and AP on each resample, and report an interval. It will be embarrassingly wide for the rare class, and that is the point: the interval is the argument you will use to stop a bad decision later. **Set a decision rule before the experiment.** Define, in advance, the minimum number of instances a class needs before its AP is allowed to gate a release, and the effect size that counts as real for the classes above that bar. Use paired comparisons - evaluate both models on the same images and compare per-image differences - which removes most of the split-to-split variance and is far more sensitive than comparing two independent point estimates. **Fix the measurement, not just the report.** If this defect class matters to the business, 12 instances is not an evaluation set; it is an anecdote. Options in rough order of cost: mine and annotate more instances specifically for evaluation, even if training data stays as it is; collect a targeted slice from production; or, where instances are irreducibly scarce, change what you measure - report a fixed-operating-point recall with an exact confidence interval, or count misses directly and set an absolute limit, which is easier to reason about with tiny counts than an area under a curve. **Consider whether the mean is the right contract at all.** For a defect-detection product, the rare class is usually the expensive one, and hiding it inside an unweighted average across classes is a governance failure regardless of the statistics. A common resolution is a two-part contract: a headline metric for tracking overall progress, plus explicit per-class floors on the classes that carry business risk, each with its own evaluation slice sized so that the floor is actually measurable. ## The trap to name out loud Averaging AP over ten IoU thresholds does *not* fix this. The ten evaluations re-use the same 12 objects, so they are strongly correlated; averaging smooths the curve a little but adds no independent evidence. Sample size is the binding constraint, and only more annotated instances relax it. Similarly, adding more classes to the benchmark dilutes each class's contribution to the mean but does nothing for the rare class's own estimate - it just hides the noise better, which is worse.
- Does averaging AP over ten IoU thresholds reduce the noise on that class?Barely. The ten evaluations score the same 12 objects, so their results are strongly correlated rather than independent samples; averaging smooths the curve slightly but adds no new evidence about the class. The binding constraint is the number of annotated instances, and only labelling more of them narrows the interval.
- The headline mAP rose 0.4 points after a change. How do you decide whether it is real?Look at the per-class table first: on a 20-class benchmark, one flipped box on a 12-instance class moves the mean by that much on its own. Then run a paired comparison - both models on the same images, differences per image, bootstrapped - and check whether the common classes moved at all. If the gain lives entirely in the rare classes, it is a coin flip, not a result.
- Would weighting classes by instance count in the mean be an improvement?It would stabilise the number and destroy its purpose. An instance-weighted mean is dominated by the common classes and can look healthy while the rare, expensive defect is undetected. The better answer keeps the unweighted mean for tracking and adds explicit per-class floors, each with an evaluation slice large enough for the floor to be measurable.
- What would you measure instead of AP for a class with barely a dozen instances?Something with an interpretable small-sample interval: recall at a fixed operating point, reported as a count of misses out of the total with an exact binomial interval, or a hard cap on missed instances. Counts are honest at n = 12 in a way an area under a precision-recall curve is not, and stakeholders can reason about them directly.
saying these in an interview costs you the question
- Reports only the headline mean with no per-class table
- Treats a fraction-of-a-point mAP change as a real improvement
- Believes more IoU thresholds means more statistical evidence
- Says adding classes to the benchmark fixes the variance
- Ignores instance counts when comparing per-class scores