In object detection, how does IoU between a predicted and ground-truth box decide a detection is correct?
answer
- area of agreement over area covered
- scale free, always between 0 and 1
- compared against a fixed cut-off
- confidence-sorted, greedy, strictly one-to-one
- duplicates on a claimed object count against you
basics
~20 sIoU is the overlap area of a predicted and a ground-truth box divided by the area they jointly cover. A prediction counts as correct when its IoU with an unclaimed ground-truth box of the same class clears a threshold, commonly 0.5.
solid answer
~40 sIoU, intersection over union, is `intersection_area / union_area` for two boxes, so it runs from 0 for no overlap to 1 for a perfect match, and it is scale free: the same relative sloppiness scores the same on a tiny box and a huge one. Scoring a detector is a matching problem, not a per-box test. Within one class, predictions are sorted by confidence and matched greedily from the top: each prediction claims the highest-IoU ground-truth box it clears the threshold against, and that ground truth is then taken. A matched prediction is a true positive; a prediction that matches nothing is a false positive; a ground-truth box nobody claimed is a false negative. Because the matching is one-to-one, a second good box on an already-claimed object is a false positive, not a bonus.
code
python · 16 linesdef iou(a, b):
ax1, ay1, ax2, ay2 = a
bx1, by1, bx2, by2 = b
iw = max(0.0, min(ax2, bx2) - max(ax1, bx1))
ih = max(0.0, min(ay2, by2) - max(ay1, by1))
inter = iw * ih
area_a = (ax2 - ax1) * (ay2 - ay1)
area_b = (bx2 - bx1) * (by2 - by1)
return inter / (area_a + area_b - inter)
# half-overlapping boxes score 1/3, not 1/2
print(iou((0, 0, 10, 10), (5, 0, 15, 10)))
# disjoint boxes: 0 here, and 0 however far apart they are moved
print(iou((0, 0, 10, 10), (20, 0, 30, 10)))
print(iou((0, 0, 10, 10), (90, 0, 100, 10)))go deeper
Be ready to state the formula in one line - overlap area divided by combined area - give its 0 to 1 range, and say that a detection is normally called correct at IoU 0.5 or above.
Explain the matching protocol, not just the ratio: confidence-sorted, one prediction to one ground truth, leftovers become false positives and misses. Expect to be asked why the union, not the ground-truth area, is the denominator.
Show you know the threshold is a product decision. Say what your downstream consumer does with the box and argue for the IoU cut-off that reflects it, and flag greedy matching as a source of scoring noise on crowded scenes.
Own the choice of evaluation contract. Decide which IoU threshold the team is judged on, whether crowded or ambiguous objects are excluded from scoring, and make sure the metric the org optimises actually tracks the value the product delivers.
## What IoU measures For two axis-aligned boxes A and B, intersection over union is ``` IoU = area(A and B) / area(A or B) ``` The numerator is the overlapping rectangle; the denominator is everything either box covers, counted once. Because a union always contains its intersection, IoU lies in [0, 1]: 0 when the boxes are disjoint or merely touch along an edge, 1 only when they coincide exactly. Two properties make it the standard criterion: - **It is scale invariant.** Multiply both boxes' coordinates by ten and IoU is unchanged. A ten-pixel error on a twenty-pixel object is a catastrophe; the same ten pixels on a thousand-pixel object is nothing, and IoU already reflects that, whereas a raw distance between corners does not. - **It is symmetric and joint.** It penalises a box that is too big, too small, or offset, in one number, rather than treating four coordinates as four independent errors. A useful sanity anchor: two identical boxes offset by half their width have intersection 1/2 and union 3/2 of one box's area, so IoU = 1/3, not 1/2. People consistently overestimate IoU by eye. ## From a number to a verdict A detector emits many boxes, each with a class and a confidence. Evaluation does not ask *is this box good* in isolation; it asks *which predictions can be paired with which ground-truth objects*. The standard protocol, per class and per image: 1. Fix an IoU threshold t (0.5 is the classic choice). 2. Sort that class's predictions by confidence, highest first. 3. Walk down the list. For each prediction, look at the still-unclaimed ground-truth boxes of that class and take the one with the highest IoU. If that IoU is at least t, the prediction is a **true positive** and that ground truth is now claimed. Otherwise the prediction is a **false positive**. 4. Ground-truth boxes still unclaimed at the end are **false negatives**. The matching is strictly one-to-one, and that is the rule candidates most often miss. If a model puts three tight boxes on the same car, one is a true positive and two are false positives, even though all three would have passed the threshold on their own. Precision therefore punishes duplication directly, which is why detectors deduplicate their raw output before it is ever scored. The ordering matters too. Because matching runs highest-confidence-first, the assignment is greedy, not globally optimal. When two objects of the same class touch - two adjacent bottles on a shelf, two overlapping industrial parts - a confident prediction sitting between them can clear the threshold against both, and it claims whichever it overlaps more. If that was the *other* object's better match, the second ground truth may end up unclaimed and counted as a miss, even though a different pairing would have scored two hits. This is not a bug being papered over; it is a deliberate choice to make the metric cheap and order-dependent in the same way a user reading a confidence-sorted list would be. ## Why the threshold is a knob, not a constant t = 0.5 is forgiving: a box can be noticeably loose and still pass. Stricter thresholds - 0.75, 0.9 - demand progressively tighter localisation, and the same model's true-positive count falls monotonically as t rises, since a match at a higher threshold is always a match at a lower one. Which t is right is a product question. If a downstream step crops the image at the predicted box and feeds it to a reader, or a robot uses the box to plan a grasp, loose boxes are useless and a high threshold is the honest score. If the system only counts objects or triggers an alert, 0.5 is fine. ## What this feeds By sweeping the confidence cut-off from high to low, the true/false positive counts above trace a precision-recall curve for the class, whose summarised area is that class's average precision at threshold t; averaging over classes gives mAP. Everything downstream inherits the matching rule: change the IoU threshold, and every true positive, false positive and false negative is recomputed from scratch. ## Common traps Dividing the intersection by the ground-truth area instead of the union produces a score of 1 for a huge box that swallows the object - the union is exactly what stops over-large boxes from cheating. And the IoU threshold is not the confidence threshold: the first decides whether a box is close enough geometrically, the second decides whether it is emitted at all.
- Two predicted boxes both clear IoU 0.5 against the same object. How are they scored?The higher-confidence one is matched first and becomes the true positive; the object is then claimed, so the second box matches nothing and is a false positive. Its geometry is irrelevant - it could be the tighter box of the two. This is why duplicate boxes are removed before scoring rather than being tolerated.
- Two same-class objects touch and one prediction clears the threshold against both. Which ground truth does it claim?The one it overlaps more, since each prediction takes its highest-IoU unclaimed ground truth. Because matching runs highest-confidence-first and is never revisited, that choice can strand the other object: if no remaining prediction clears the threshold against it, it is counted as a miss even though a different global pairing would have scored two hits.
- Why divide by the union rather than by the ground-truth box's area?Dividing by the ground-truth area rewards over-large boxes: a box covering the whole image contains the object entirely and would score 1. The union counts the predicted box's excess area as well, so a box is penalised for being too big and for being too small. That two-sided penalty is exactly what makes IoU a localisation measure.
IoU is like laying two stickers on a page and asking what fraction of the ink-covered area both stickers cover: identical placement gives 1, edge-to-edge gives 0.
saying these in an interview costs you the question
- Divides the intersection by the ground-truth area instead of the union
- Says any overlap at all counts as a detection
- Assumes several predictions on one object all count as true positives
- Thinks IoU depends on how many pixels the object spans
- Confuses the IoU match threshold with the confidence threshold