Why does OpenCV's net.forward() with no argument miss a detector's outputs?
answer
- No-argument form returns one blob only
- Some detectors have several heads
- Ask the network which layers are unconnected
- A helper method exists so you avoid index arithmetic
- Suppression belongs after merging all heads
basics
~20 sCalled with no argument, cv2.dnn Net.forward returns a single blob — the first output. A multi-head detector such as a Darknet YOLOv3 net has several output layers, so you must pass net.getUnconnectedOutLayersNames() to forward and receive a list of blobs.
solid answer
~40 s`forward()` with no arguments runs the whole graph but hands back one blob, the first output. That is fine for a classifier or an SSD head, and quietly lossy for any network with multiple detection heads — a Darknet YOLOv3 or v4 graph has three output layers at three scales, and taking only one means you silently lose every detection at the other two scales. The fix is `outs = net.forward(net.getUnconnectedOutLayersNames())`, which returns a list of blobs in the order the names were given. `getUnconnectedOutLayersNames()` is the safe API: the older idiom indexed `net.getLayerNames()` with `net.getUnconnectedOutLayers()`, whose indices are 1-based, which is the source of a very common off-by-one in copied sample code. Passing a single explicit layer name is also how you pull an intermediate activation out for debugging.
code
python · 9 linesimport cv2
import numpy as np
def run_all_heads(net, frame):
blob = cv2.dnn.blobFromImage(frame, 1 / 255.0, (416, 416), swapRB=True, crop=False)
net.setInput(blob)
names = net.getUnconnectedOutLayersNames()
outs = net.forward(names) # list, one blob per head
return list(zip(names, [np.asarray(o).shape for o in outs]))go deeper
Know that forward() without arguments gives you one output blob, and that a network can have more than one output layer.
Be able to write the two-line fix with getUnconnectedOutLayersNames, and explain why the missing scale looks like a model-accuracy problem rather than an API mistake.
Show that you merge candidates across all heads before suppression, and that you would catch this class of bug by asserting the number of output blobs against the model's expected head count at startup.
Argue for the higher-level DetectionModel wrapper as a default across a team's codebase — fewer hand-rolled postprocessing paths means fewer silent variants of this same defect in production.
## What "output" means to the DNN module After import, an OpenCV `Net` is a directed graph of layers. Some layers feed other layers; the ones whose output nothing consumes are the network's **unconnected output layers**. A classifier has exactly one. An SSD detection head has one. A Darknet YOLOv3 or YOLOv4 graph has three, one per detection scale, and each produces predictions for a different range of object sizes. ## The overload that catches people There are two shapes of the call: - `net.forward()` or `net.forward(outputName)` — returns a **single** blob. With no name, it computes the whole network and returns the first output. With a name, it computes up to that layer and returns its blob. - `net.forward(outBlobNames)` — takes a sequence of layer names and returns a **list** of blobs, one per name, in that order. The single-blob form is the one every tutorial opens with, so it is the one that gets pasted into a multi-head pipeline. Nothing raises. You get a valid array, you parse it correctly, you draw correct boxes — and small objects, or large ones, simply never appear. The bug reads as a model-quality problem, which is why it can survive a long time. ## Getting the names right The direct API is: ``` names = net.getUnconnectedOutLayersNames() outs = net.forward(names) ``` Older sample code predates that helper and does this instead: take `net.getLayerNames()`, take `net.getUnconnectedOutLayers()`, and index the first with the second. The catch is that the layer indices returned there are **1-based**, so the correct lookup subtracts one. Copies of this snippet float around both with and without the subtraction, and the version without it either throws an index error or, worse, picks the wrong layers. `getUnconnectedOutLayersNames()` exists precisely to remove that class of bug — prefer it and the question disappears. ## Parsing a list of outputs With three YOLO heads you get three arrays, each `(num_boxes, 5 + num_classes)` rows for that scale. The standard postprocessing loops over all three, thresholds by confidence, converts the normalised centre-form boxes to pixel corner-form, and accumulates one flat candidate list — and only then applies non-maximum suppression **across the combined list**. Running NMS per output blob is a second, subtler version of the same bug: the same object detected at two scales produces two overlapping boxes that never get compared, so duplicates survive into the final result. ## Where else the multi-output form matters - **Multi-task networks** — a model with both a segmentation mask and a classification head. - **Text detection models** with separate score and geometry outputs, where you need both blobs to reconstruct boxes. - **Debugging a port** — passing one intermediate layer name lets you compare that activation against the original framework and bisect where the two diverge. ## The higher-level escape hatch If the ceremony is not adding value, OpenCV ships wrapper classes over `Net` — `cv2.dnn.DetectionModel` is the relevant one — that take `setInputParams(scale, size, mean, swapRB)` and expose `detect(frame, confThreshold, nmsThreshold)` returning class ids, confidences and boxes. They handle the multiple output blobs, the coordinate conversion and the suppression internally. The tradeoff is the usual one: less code and fewer places to get it wrong, against less control when a model's output layout does not match what the wrapper assumes. Knowing both paths, and why the low-level one has this trap, is what the question is really probing.
- After collecting three YOLO output blobs, where exactly should NMS run?Once, over the merged candidate list from all three scales. The same object is often detected at two scales; suppressing per blob compares each detection only against others at its own scale, so cross-scale duplicates survive into the final output. Threshold first, flatten into one list of boxes and scores, then call `cv2.dnn.NMSBoxes`.
- Why is getUnconnectedOutLayersNames preferred over indexing getLayerNames?`getUnconnectedOutLayers` returns 1-based layer indices, so indexing the name list with them directly is off by one. Copied sample code exists in both corrected and uncorrected forms. The names helper returns the strings directly, removing the arithmetic and the failure mode with it.
- How do you use forward to inspect a single intermediate layer?Pass its name: `net.forward("layer_name")` computes the graph only up to that layer and returns its blob. Use `net.getLayerNames()` to find the exact string. This is the standard bisection technique for a ported model whose final output disagrees with the original framework.
saying these in an interview costs you the question
- Assuming forward() always returns every output of the network
- Running non-maximum suppression separately per output blob
- Indexing getLayerNames with 1-based indices without subtracting one
- Thinking a missing detection scale means the weights are bad
- Believing forward with a layer name still executes the whole graph