Why can a detector win on mAP at IoU 0.5 but lose on mAP averaged over 0.50 to 0.95?
answer
- one hidden parameter behind every mAP
- ten evaluations, not one
- nine of ten are stricter
- classification good, boxes loose
- plot AP against the threshold
basics
~20 smAP at IoU 0.5 forgives loose boxes, while averaging over the ten thresholds from 0.50 to 0.95 in 0.05 steps scores localisation quality directly. A model that classifies confidently but draws sloppy boxes tops the first and collapses on the second.
solid answer
~40 sThe two numbers reward different skills. [email protected] only asks whether a box roughly covers the object, so it is dominated by whether you found the object at all and ranked it confidently. The averaged form recomputes the whole matching at IoU 0.50, 0.55, ... 0.95 and means the results, so nine of its ten terms are stricter than 0.5 and it is mostly a localisation-precision score. A detector with a strong classifier and a weak box head detects everything, ranks it well, and wins at 0.5 - then its true positives evaporate above 0.7 and the average craters. The reverse also happens: a model with tight boxes but noisier confidence ordering loses the loose comparison and wins the strict one. Plotting AP against the threshold separates the two failure modes immediately.
go deeper
Know that mAP is always computed at some IoU threshold, that 0.5 is the lenient classic, and that a stricter threshold demands tighter boxes and can only lower the score.
Explain that the averaged form re-runs the whole matching at ten thresholds and means the results, and that per-class AP is non-increasing in the threshold so the gap between the two numbers is a localisation measure.
Diagnose from the shape: sweep AP against threshold to separate loose boxes from missed objects or bad ranking, and check whether the strict thresholds are measuring your annotators' own disagreement rather than the model.
Set the evaluation contract before the team optimises against it. Tie the reported threshold to how boxes are actually consumed downstream, insist both conventions are published when they disagree, and refuse comparisons that quietly switch protocols.
## Two ways to aggregate Average precision for one class is computed *at a fixed IoU threshold*: the threshold decides which predictions match ground truth, and from those true/false positives a precision-recall curve is traced by sweeping the confidence cut-off. mAP is that averaged over classes. So mAP is really a function of one hidden parameter, and reporting it without saying the threshold is meaningless. The two conventions in wide use: - **mAP at IoU 0.5** - a single, forgiving threshold. Classic and easy to read. - **mAP averaged over 0.50 to 0.95 in steps of 0.05** - ten separate evaluations, meaned. This is the headline number of the COCO-style protocol. Because a match at a high threshold is automatically a match at every lower one, per-threshold AP is non-increasing in the threshold. The averaged number is therefore always at or below [email protected], and the *size* of the gap is the diagnostic. ## What the gap measures Suppose your model finds nearly every object and orders its detections well, but each box is drawn a little too big, or is consistently shifted. At threshold 0.5, that costs nothing: a box that overlaps two thirds of the object still passes. At 0.75 a fair share of those boxes fall out of the match and turn into false positives *and* leave their objects unclaimed as misses, so precision and recall both drop. At 0.9 almost nothing survives. Average the ten terms and you get a small number even though [email protected] looked excellent. That is the confident-but-sloppy detector: strong classification head, weak or under-trained box regression. It wins any comparison run at 0.5 and loses badly on the averaged metric, and if your team ships on the 0.5 number you will discover the problem only when a downstream consumer starts cropping images at those boxes. The mirror image exists too. A model whose boxes are extremely tight but whose confidence scores are poorly ordered - some true detections scored below some false ones - suffers on the precision-recall curve at every threshold, but its loss is roughly constant across thresholds, whereas its rival's grows. Comparing the two curves rather than the two numbers is what tells you this. ## Reading it in practice The move is to plot AP against IoU threshold for each model, per class: - **High at 0.5, steep fall by 0.75** - localisation problem. Look at box-loss choice, feature resolution, label tightness, and whether your objects are small relative to the stride of the features the box head reads. - **Low everywhere, flat-ish shape** - detection or ranking problem: the objects are being missed or scored badly, and tightening boxes will not help. - **A cliff between two adjacent thresholds for one class only** - suspect systematic annotation convention, for example whether labels include an object's shadow, handle or occluded part. If humans drew the boxes to a different convention than your model learned, the two disagree by a constant margin that a threshold sweep exposes exactly. That last point is worth stressing: a strict metric measures agreement with the *annotation convention*, not with physical truth. If annotators are themselves inconsistent to within IoU 0.85, then AP at 0.9 is measuring label noise and no model can win it. Knowing your labels' own agreement level tells you which thresholds are informative on your data. ## Which one to report This is a product decision, and the honest answer names the consumer: - If a box triggers an alert, feeds a counter, or is drawn for a human who then looks at the region, 0.5 is a defensible operating criterion. - If a box is used to crop for a downstream reader, to measure a physical dimension, or to plan a grasp, loose boxes are failures and the strict average is closer to the truth. - If you are comparing to published work, report the same convention it used, and report both when they disagree. The worst outcome is a team that tunes for months against a threshold nobody chose deliberately. Pick the threshold that matches how the boxes are consumed, state it beside every number, and keep the sweep as a diagnostic so that a localisation regression cannot hide behind an unchanged headline figure.
- How many IoU thresholds does the averaged form use, and what is the step?Ten: 0.50, 0.55, 0.60 and so on up to 0.95, in steps of 0.05. Each threshold is a full re-evaluation - matching, precision-recall curve and per-class AP are all recomputed - and the ten results are averaged with equal weight. Only one of the ten is the forgiving 0.5 criterion, which is why the averaged number is so sensitive to box tightness.
- [email protected] is unchanged after a release but the averaged mAP dropped. What do you check?Something moved boxes without changing what was detected. Check for a change in box-loss weighting or objective, a resolution or crop change in preprocessing, and any relabelling that shifted the annotation convention. Plot AP versus threshold before and after: a widening gap that starts around 0.7 confirms pure localisation drift rather than a detection or ranking regression.
- Can AP at IoU 0.75 ever exceed AP at 0.5 for the same model?No. Any prediction that matches at 0.75 also matches at 0.5, so the stricter threshold can only remove true positives, turning them into false positives and misses. Per-class AP is non-increasing in the IoU threshold, and a reported number that violates this points to a bug in the evaluation code rather than an unusual model.
saying these in an interview costs you the question
- Thinks a higher IoU threshold makes the task easier
- Says mAP@[.50:.95] averages over classes only, not thresholds
- Assumes the two conventions always rank models identically
- Blames a drop at strict thresholds on the classifier
- Reports a mAP number without stating the IoU threshold