How do you turn raw cv2.dnn detection output into final boxes on a frame?
answer
- Layout is a property of the model's head
- Coordinates come back normalised, not in pixels
- Centre-form and corner-form need different maths
- Two different thresholds do two different jobs
- Suppression returns indices, not boxes
basics
~20 sRead the model's output layout, keep rows above a confidence threshold, rescale the normalised coordinates by the original frame's width and height, convert them to x, y, w, h, then call cv2.dnn.NMSBoxes to drop overlapping duplicates and draw only the returned indices.
solid answer
~50 sThere is no universal parser — the layout is a property of the model's head. An SSD-style output is `(1, 1, N, 7)` where each row is `[batch_id, class_id, confidence, x1, y1, x2, y2]` with corner coordinates **normalised to 0..1**; a Darknet YOLO output row is `[cx, cy, w, h, objectness, class scores...]`, also normalised, with a **centre-form** box. In both cases you multiply by the *original* frame width and height, not the blob's input size, because the normalisation is relative to whatever was fed in. Then filter by confidence, build a flat list of `[x, y, w, h]` boxes plus scores, and call `cv2.dnn.NMSBoxes(boxes, scores, score_threshold, nms_threshold)`, which returns the indices to keep. Skipping suppression leaves a cluster of near-duplicate boxes on each object; forgetting the coordinate rescale gives boxes crammed into the top-left few pixels.
code
python · 16 linesimport cv2
def decode_ssd(out, frame_w, frame_h, conf_thr=0.5, nms_thr=0.4):
boxes, scores = [], []
for i in range(out.shape[2]):
confidence = float(out[0, 0, i, 2])
if confidence < conf_thr:
continue
x1 = int(out[0, 0, i, 3] * frame_w)
y1 = int(out[0, 0, i, 4] * frame_h)
x2 = int(out[0, 0, i, 5] * frame_w)
y2 = int(out[0, 0, i, 6] * frame_h)
boxes.append([x1, y1, x2 - x1, y2 - y1])
scores.append(confidence)
keep = cv2.dnn.NMSBoxes(boxes, scores, conf_thr, nms_thr)
return [boxes[int(i)] for i in keep]go deeper
Know that raw output is normalised numbers, not pixels, and that a suppression step is needed before drawing anything.
Be able to describe both common row layouts, convert centre-form to top-left plus size, and explain what NMSBoxes takes and returns.
Demonstrate the production details: per-class suppression, undoing letterbox padding, and validating a new model's output layout at startup rather than trusting a pasted parser.
Own postprocessing as a contract surface — one shared, tested decode path per model family beats per-service reimplementations, since every bug in this stage is silent and geometric rather than an exception.
## The output tensor is model-specific The DNN module returns whatever the network's final layers produce. It does not decode, threshold, or suppress anything for you when you use the raw `Net` API. So the first job on any new model is to establish the output layout — from the model card, from the exporting code, or by printing `out.shape` and inspecting a few rows. Two layouts dominate in interview material: **SSD / Caffe-style detection output** — shape `(1, 1, N, 7)`. Each of the `N` rows is `[image_id, class_id, confidence, x_left, y_top, x_right, y_bottom]`. Coordinates are corner-form and normalised to `[0, 1]`. **Darknet YOLO output** — shape `(num_boxes, 5 + num_classes)` per head. Each row is `[center_x, center_y, width, height, objectness, class_0 ... class_k]`. Coordinates are centre-form and normalised. The class score is typically `objectness * max(class scores)` or simply the max class score depending on the variant, and the class id is the argmax over the class slice. Newer single-output exports (for example transposed YOLO ONNX heads) may be shaped `(1, 4 + num_classes, num_boxes)` and need a transpose before the same logic applies. This is exactly why you print the shape first instead of pasting a parser. ## Rescaling coordinates Because the values are normalised, you multiply x-components by the frame's width and y-components by its height: ``` x1 = int(row[3] * frame_width) y1 = int(row[4] * frame_height) ``` Use the **original frame's** dimensions, not the network input size. If you resized with `crop=False`, the network saw a stretched version of the whole frame, so normalised coordinates map back onto the whole frame exactly. If you letterboxed the frame yourself before `blobFromImage`, you must undo the padding and the scale before you draw — a step that is easy to forget and produces boxes that are consistently offset and slightly small. Centre-form needs one more conversion, because both the drawing call and the suppression call want top-left plus size: ``` x = int(cx * W - w * W / 2) y = int(cy * H - h * H / 2) ``` ## Thresholding Two thresholds do different jobs and are often confused: - **Confidence threshold** — discard rows whose score is below it. This is a recall/precision dial. - **NMS threshold** — the IoU above which two surviving boxes are considered the same object. Lower is more aggressive suppression. Raising the NMS threshold when you mean to raise the confidence threshold is a classic mix-up; it makes duplicates worse rather than reducing false positives. ## Non-maximum suppression `cv2.dnn.NMSBoxes(bboxes, scores, score_threshold, nms_threshold)` takes a list of `[x, y, w, h]` boxes and a parallel list of float scores and returns the **indices** of the boxes to keep, sorted by score. Two habits matter: - The boxes must be in `x, y, w, h` form. Feeding corner-form silently computes wrong overlaps, so suppression either keeps too much or too little with no error. - For multi-class detection, run suppression **per class** (or use the batched variant), otherwise a person box and an overlapping bicycle box suppress each other, and the model looks like it cannot see two things at once. Then you index back into your candidate list and draw only the kept boxes with `cv2.rectangle` and `cv2.putText`. ## The wrapper alternative `cv2.dnn.DetectionModel` wraps a `Net`, takes `setInputParams(scale, size, mean, swapRB)` once, and exposes `detect(frame, confThreshold, nmsThreshold)` returning class ids, confidences and pixel-space boxes. It performs the blob construction, the decode and the suppression internally. For the common model families this removes every failure mode above. The reason to still know the manual path is that the wrapper only supports the output layouts it recognises — an unusual head means you decode by hand, and that is when this question separates candidates who have parsed a real detector from those who have only called a wrapper. ## What a good answer emphasises That the shape and semantics of the output are part of the model contract, that the failures here are silent and geometric rather than exceptions, and that suppression is a required step rather than a refinement.
- Why should non-maximum suppression be applied per class rather than over all detections?Because IoU alone cannot tell two overlapping objects of different classes apart. A person standing beside a bicycle produces overlapping boxes; class-agnostic suppression deletes one of them, so the model appears blind to the second object. Run suppression within each class id, or use the batched variant that takes class ids into account.
- Your boxes are all bunched in the top-left corner of the frame. What went wrong?The normalised coordinates were drawn without rescaling — values in `[0, 1]` cast to int collapse to 0 or 1 pixels. Multiply x-components by the original frame width and y-components by its height before converting to int, and remember to undo any letterbox padding you added yourself.
- What is the difference between the score threshold and the NMS threshold in cv2.dnn.NMSBoxes?The score threshold drops low-confidence candidates outright; it trades recall for precision. The NMS threshold is an IoU cutoff deciding when two surviving boxes describe the same object — lower suppresses more aggressively. Turning up the IoU threshold to reduce false positives is the common mix-up and makes duplicates worse.
- When would you use cv2.dnn.DetectionModel instead of parsing the output yourself?Whenever the model is one of the families it recognises. `setInputParams` plus `detect` handles blob construction, decoding and suppression, eliminating the coordinate and threshold bugs. Hand-parse only when the head has an unusual layout the wrapper does not decode, or when you need intermediate outputs it does not expose.
saying these in an interview costs you the question
- Drawing normalised coordinates directly without scaling to the frame
- Skipping non-maximum suppression and blaming the duplicates on the model
- Passing corner-form boxes to NMSBoxes, which expects x, y, w, h
- Confusing the confidence threshold with the IoU threshold
- Assuming every detector shares one output layout