How do you load an ONNX model with cv2.dnn and run one inference?
answer
- Four calls: read, blob, setInput, forward
- Extension picks the importer for you
- Unsupported operators fail at load, not inference
- Caffe and TensorFlow importers order arguments differently
- One Net object per thread
basics
~10 sCall cv2.dnn.readNet("model.onnx") (or readNetFromONNX), build a blob with cv2.dnn.blobFromImage, pass it in with net.setInput(blob), then call net.forward() to get the output array. No training framework is needed at runtime.
solid answer
~40 sThe pipeline is four calls: `net = cv2.dnn.readNet(path)` imports the graph, `blob = cv2.dnn.blobFromImage(...)` turns a frame into an NCHW tensor, `net.setInput(blob)` binds it to the input layer, and `out = net.forward()` runs the graph and returns a NumPy array. `readNet` infers the framework from the file extension and delegates to the specific importer — `readNetFromONNX` for `.onnx`, `readNetFromTensorflow` for a frozen `.pb`, `readNetFromCaffe` for a `.prototxt` plus `.caffemodel` pair. Two-file formats have an argument-order trap worth remembering: `readNetFromCaffe(prototxt, caffeModel)` takes the text config first, while `readNetFromTensorflow(model, config)` takes the binary weights first. The net object is stateful and not thread-safe, so a multi-threaded server needs one `Net` per worker, and the first `forward()` is slower because the module allocates and optimises the layer graph on that call.
code
python · 10 linesimport cv2
import numpy as np
def infer(model_path, frame):
net = cv2.dnn.readNet(model_path)
blob = cv2.dnn.blobFromImage(frame, 1 / 255.0, (224, 224), (0, 0, 0), swapRB=True)
net.setInput(blob)
out = net.forward()
_, timings = net.getPerfProfile()
return out, timingsgo deeper
Memorise the four-call sequence — readNet, blobFromImage, setInput, forward — and be able to say that OpenCV runs models but never trains them.
Explain that the importer translates operators into OpenCV layers, so unsupported models fail at load, and know why the first forward pass is slower than the rest.
Talk about running this inside a real service: one Net per worker, warm-up before benchmarking, and per-layer timing via getPerfProfile when a frame budget is missed.
Frame the choice: a single already-linked dependency and identical APIs across C++, Python and Android, weighed against operator coverage that lags the training frameworks and no vendor accelerator path beyond the build flags.
## The minimal pipeline OpenCV's DNN module is an inference-only engine. There is no training, no autograd, no optimiser — you bring a model somebody else trained and run it. The end-to-end flow is short enough to memorise: 1. **Import the graph.** `net = cv2.dnn.readNet(model, config, framework)`. With one argument it sniffs the framework from the file extension. 2. **Preprocess.** `blob = cv2.dnn.blobFromImage(frame, scale, size, mean, swapRB, crop)` produces the 4-D NCHW input. 3. **Bind.** `net.setInput(blob)`. An optional layer-name argument targets a specific input for multi-input graphs. 4. **Run.** `out = net.forward()` returns a NumPy array whose shape depends entirely on the model's head. ## Choosing the importer `readNet` is a convenience wrapper. The concrete importers are the ones to know by name: - `cv2.dnn.readNetFromONNX(onnxFile)` — one file, weights and topology together. This is the path most teams use today because every major training framework can export ONNX. - `cv2.dnn.readNetFromTensorflow(model, config)` — a frozen `.pb` graph, optionally with a `.pbtxt` text graph that the importer uses to resolve subgraphs it cannot infer (object-detection APIs typically need it). - `cv2.dnn.readNetFromCaffe(prototxt, caffeModel)` — the architecture text file first, the weights second. - `cv2.dnn.readNetFromDarknet(cfgFile, darknetModel)` — the classic YOLOv3/v4 path. The argument ordering is genuinely inconsistent between the Caffe and TensorFlow importers, and swapping them produces a parse error rather than anything descriptive. `readNet` papers over it because it dispatches on extension, which is one reason to prefer it. ## What import actually does The importer does not just deserialise; it translates every operator into OpenCV's own layer implementations. That is why import is the step where unsupported models fail. If the graph contains an operator the module has no layer for, `readNet` throws `cv2.error` naming the layer type, and it throws at load time rather than at inference time. This is actually a useful property: model compatibility is a startup check, not a runtime surprise. Fixing it means either simplifying the export (fusing or removing exotic ops, exporting with static shapes) or choosing a different runtime. ## Running the forward pass `forward()` executes the graph. Some things about it are worth knowing before you benchmark: - **The first call is expensive.** The module allocates internal buffers, fuses layers, and prepares the chosen backend on the first `forward()`. Any latency measurement that includes it is wrong; warm up with a dummy frame first. - **The Net is stateful.** `setInput` stores the blob, `forward` writes internal buffers. Sharing one `Net` across threads is a data race. Give each worker its own instance, or serialise access. - **Output shape is model-specific.** A classifier returns `(1, num_classes)`; an SSD detector returns `(1, 1, N, 7)`; a segmentation head returns a spatial map. There is no universal parsing. - **You can inspect intermediates.** `net.forward(layerName)` runs only up to that layer and returns its blob, which is the standard way to debug where a ported model diverges from its original framework. ## Timing and introspection `net.getPerfProfile()` returns the total inference time in tick counts along with a per-layer breakdown; divide by `cv2.getTickFrequency()` for seconds. `net.getLayerNames()` lists every layer in execution order. Together they answer "is it slow, and which layer is slow" without any external profiler — a real advantage of the module in embedded contexts where you cannot install tooling. ## When this is the right tool The pipeline above needs exactly one dependency, which is often already linked because you are decoding video with OpenCV anyway. The Mat you got from `VideoCapture` goes into `blobFromImage` with no conversion, and the same four calls exist in the C++, Java and Android bindings. For a small vision model at the edge, that is a genuinely strong story. The cost is limited operator coverage and no access to vendor accelerators beyond what OpenCV was built with.
- Why is your first measured inference far slower than the ones after it?The module lays out buffers, fuses layers and initialises the selected backend on the first `forward()` call rather than at `readNet` time. Always run a warm-up pass on a dummy blob before timing, and measure with `net.getPerfProfile()` or a wall-clock loop over many frames.
- What happens if the ONNX file contains an operator OpenCV does not implement?`readNet` raises `cv2.error` at load time, naming the unsupported layer type. Because it fails on import rather than mid-inference, it works as a startup compatibility check. The fixes are to re-export with simpler ops or static shapes, or to move that model to a runtime with broader operator coverage.
- Can you share one Net across a thread pool to serve concurrent requests?No. `setInput` and `forward` mutate internal state, so concurrent use is a data race with corrupted outputs, not just contention. Create one `Net` per worker thread, or guard a single instance with a lock and accept the serialisation.
- How would you inspect an intermediate activation to debug a ported model?Pass a layer name to `forward`: `net.forward("layer_name")` runs the graph up to that layer and returns its blob. Combine it with `net.getLayerNames()` to find the name, then compare the tensor against the same activation in the original framework to locate where the port diverges.
saying these in an interview costs you the question
- Thinking OpenCV can fine-tune or train the loaded network
- Calling forward() without setInput and expecting a result
- Reusing one Net object across threads for throughput
- Assuming imread output can be passed straight to setInput
- Believing an unsupported operator surfaces only at inference time