Why does class-wise non-maximum suppression delete correct boxes on a shelf of densely packed cartons?
answer
- post-processing, no gradient reaches it
- greedy, sorted by score, class-wise
- assumes big overlap means duplicate
- genuine neighbours really do overlap
- decay the score instead of deleting
basics
~20 sSuppression assumes heavily overlapping boxes of the same class are duplicates of one object. Identical cartons packed side by side genuinely overlap above the threshold, so the lower-scoring box is deleted even though it is a real, distinct object.
solid answer
~50 sGreedy suppression sorts detections by score, keeps the top one, deletes every remaining same-class box overlapping it above a threshold, and repeats. The whole method rests on one assumption: two same-class boxes that overlap a lot are two guesses at one object. On a shelf of 200 identical cartons that assumption is false — neighbouring cartons really do overlap above 0.5, so each deletion costs a true positive, and that recall is gone before anything is measured. Retraining will not fix it, because the loss never sees the suppression step. The levers are: raise the overlap threshold, which keeps the neighbours but lets genuine duplicates through and hurts precision; use soft suppression, which decays an overlapping box's score instead of deleting it, so a confident true positive survives while a weak duplicate falls under the score threshold; or use a head trained with one-to-one assignment, which needs no suppression.
code
python · 25 linesdef iou(a, b):
ax1, ay1, ax2, ay2 = a
bx1, by1, bx2, by2 = b
ix = max(0, min(ax2, bx2) - max(ax1, bx1))
iy = max(0, min(ay2, by2) - max(ay1, by1))
inter = ix * iy
area_a = (ax2 - ax1) * (ay2 - ay1)
area_b = (bx2 - bx1) * (by2 - by1)
return inter / (area_a + area_b - inter)
def nms(boxes, scores, thr):
order = sorted(range(len(boxes)), key=lambda i: -scores[i])
keep = []
while order:
i = order.pop(0)
keep.append(i)
order = [j for j in order if iou(boxes[i], boxes[j]) <= thr]
return keep
# two adjacent cartons, each 100 px wide, genuinely overlapping by 70 px
cartons = [(0, 0, 100, 100), (30, 0, 130, 100)]
scores = [0.90, 0.85]
print(round(iou(cartons[0], cartons[1]), 3)) # 0.538 - real overlap, two real objects
print(nms(cartons, scores, 0.5)) # [0] - the second carton is deleted
print(nms(cartons, scores, 0.6)) # [0, 1] - both survivego deeper
Know that a detector outputs many overlapping boxes for one object and that suppression keeps the highest-scoring one and drops near-duplicates. Be able to state that it runs after the network, not inside it.
Explain the greedy loop precisely: sort by score, keep the top, delete same-class boxes above an overlap threshold, repeat. Then state the assumption it makes and give one concrete situation where that assumption is false.
Diagnose it. Recognise the signature of recall that is fine on sparse scenes and collapses on crowded ones, rule out training as the cause, and argue the threshold change or soft suppression with both sides of the cost stated.
Frame it as a product decision. Decide whether a missed object or a double count is more expensive in this application, whether one operating point can serve every scene density, and whether the recurring tuning justifies moving to a head that needs no suppression.
## Why suppression exists A detection head predicts densely: many neighbouring locations see the same object and all fire. Raw output for one carton might be a dozen boxes within a few pixels of each other. Non-maximum suppression is the classical fix, and it is pure post-processing — it has no parameters, no gradient, and is not part of training. The greedy algorithm is: 1. Drop everything below a score threshold. 2. Sort what is left by score, descending. 3. Take the top box, move it to the output. 4. Delete every remaining box whose overlap with it exceeds an IoU threshold. 5. Repeat from step 3 until nothing is left. In practice it is run *class-wise*: boxes only suppress other boxes of the same predicted class. That is what lets a person box and a bicycle box occupy nearly the same pixels and both survive, and it is why a cross-class overlap is not the failure mode here. ## The assumption that breaks Step 4 encodes a belief: *high same-class overlap means duplicate*. That is true for a single carton seen by twelve nearby locations. It is false for two different cartons standing shoulder to shoulder on a shelf. Make it concrete. Two cartons each 100 pixels wide, whose visible boxes overlap by 70 pixels because one occludes the other. The intersection is 7,000 pixels and the union 13,000, giving an overlap of about 0.54. With a threshold of 0.5 the second carton is deleted. There was nothing wrong with the prediction — the detector found both objects, scored both correctly, and post-processing threw one away. At scale this compounds. On a densely packed pallet the same deletion happens along every row, and a count that should read 200 reads 130. The characteristic signature is that recall is fine on sparse images and collapses only on crowded ones, and that increasing training data does not move it, because the loss never observes the suppression step. ## The levers, and what each one costs **Raise the overlap threshold.** Going from 0.5 to 0.65 keeps the neighbouring cartons. The cost is symmetric: genuine duplicates on one object also now survive, so each real object may emit two or three boxes. Precision falls, and downstream consumers that count objects or trigger actions get inflated numbers. Where you set it depends entirely on which error is more expensive to you. **Soft suppression.** Instead of deleting an overlapping box, multiply its score by a decay factor that grows with the overlap — a linear factor `1 - IoU` or a Gaussian `exp(-IoU^2 / sigma)`. A duplicate of an already-kept object usually had a mediocre score to begin with; after decay it drops below the score threshold and disappears. A genuinely distinct neighbouring carton usually had a strong score; after decay it is weakened but survives and is emitted with a lower confidence. This trades a hard, irreversible decision for a graded one, at the price of one more hyperparameter and detections whose scores are no longer directly comparable to the raw ones. **Change the head.** Set-prediction heads train with a one-to-one assignment between predictions and ground-truth objects, so the network itself learns not to emit duplicates and suppression is dropped entirely. This removes the failure mode at its root, but it is an architecture decision, not a knob you turn on a Friday. **Do not reach for the wrong lever.** Lowering the score threshold does not help — the deleted carton was not low-scoring, it was suppressed. Neither does adding training data of crowded shelves; the model may already be right. ## The greedy detail worth knowing Suppression is greedy and score-ordered, not globally optimal. A high-scoring *false* positive sitting between two cartons can suppress both true positives, and nothing later in the pipeline can recover them. This is why the choice of threshold interacts with how well calibrated your scores are: suppression trusts the ranking absolutely. ## The other knob in the same place The score threshold applied *before* suppression is the other post-processing decision and it is often set by habit. It does two things: it controls how many boxes suppression has to consider, which is a real latency lever on crowded frames; and it moves the operating point, trading recall for precision, before any evaluation happens. On a crowded-shelf system the sensible order is to fix the suppression behaviour first, then choose the score threshold against the cost of a miss versus a false alarm in that application. ## What an interviewer is listening for Name the assumption, not just the algorithm. State that this is a post-processing failure that training cannot fix. Give the threshold trade in both directions rather than only one. And be specific that suppression is class-wise, so the answer is about same-class neighbours, not about different classes overlapping.
- You raise the suppression threshold from 0.5 to 0.7 and recall improves. What should you check before shipping?Precision on ordinary, sparse images. The higher threshold lets duplicates on a single object survive, so anything downstream that counts objects or triggers per detection now over-reports. Check duplicates-per-object on a sparse validation slice, and confirm the crowded and sparse slices are still acceptable together — a single threshold has to serve both unless you are willing to branch on scene density.
- How does the score threshold applied before suppression differ from the suppression overlap threshold?They act on different quantities. The score threshold discards low-confidence boxes and moves the precision/recall operating point; it also caps how many boxes suppression must compare, which is a real latency lever on crowded frames. The overlap threshold decides which surviving boxes are treated as duplicates of each other. A box can be deleted by suppression despite a high score, so lowering the score threshold cannot recover it.
- Why can a high-scoring false positive be especially damaging under greedy suppression?Greedy suppression processes boxes in score order and trusts that ranking absolutely. A confident false positive placed between two real objects is kept first and then suppresses both of them if it overlaps each above the threshold. Two true positives are lost to one bad prediction, and nothing downstream can recover them — which is why score calibration matters more in crowded scenes than in sparse ones.
Suppression is a proofreader deleting duplicate lines. It works fine on an accidental copy-paste, but on a packing list where the same item legitimately appears twenty times it silently shortens the list.
saying these in an interview costs you the question
- Believes more training data fixes suppression-caused misses
- Thinks suppression compares boxes across different classes
- Says lowering the score threshold recovers suppressed boxes
- Describes the threshold trade in only one direction
- Treats suppression as a learned, differentiable part of the model