In OpenCV DNN, why might DNN_TARGET_CUDA still run inference on the CPU?
answer
- The setters are named preferable, not required
- Wheels from PyPI are not built the way you assume
- Backend and target are two axes, not one
- Unsupported layers do not stay on the device
- Prove it with the profiler, not with belief
basics
~20 ssetPreferableBackend and setPreferableTarget express a preference, not a requirement. If the binary was not built with CUDA support — pip's opencv-python wheels are not — OpenCV logs a warning and silently falls back to the default CPU implementation, so timings stay flat.
solid answer
~50 sThe two setters are named *preferable* for a reason: an unavailable backend or an unsupported backend/target pairing falls back rather than throwing. The most common cause is the build itself — the standard `opencv-python` wheels from PyPI ship without the CUDA DNN backend, so `setPreferableBackend(cv2.dnn.DNN_BACKEND_CUDA)` and `setPreferableTarget(cv2.dnn.DNN_TARGET_CUDA)` do nothing until you build OpenCV from source with `WITH_CUDA` and `OPENCV_DNN_CUDA` enabled against cuDNN. Check with `cv2.getBuildInformation()` before you trust any GPU number. Second cause: the two settings must be consistent — a CUDA target only makes sense under the CUDA backend, and `DNN_TARGET_CUDA_FP16` additionally needs a GPU with fast half precision or it will be slower than FP32. Third: even on a working CUDA build, layers the backend does not implement fall back per layer, and the resulting host/device transfers can erase the speedup. Confirm with `net.getPerfProfile()` rather than by belief.
code
python · 12 linesimport cv2
info = cv2.getBuildInformation()
has_cuda = "NVIDIA CUDA" in info
print("cuda dnn backend available:", has_cuda)
def profile(net, blob, warmups=3, runs=20):
for _ in range(warmups + runs):
net.setInput(blob)
net.forward()
total, layer_times = net.getPerfProfile()
return total / cv2.getTickFrequency(), layer_timesgo deeper
Know that backend and target are separate settings, and that asking for a GPU does not guarantee you get one.
Explain the silent fallback: an unavailable backend logs a warning and reverts to CPU, so identical timings before and after are the diagnostic signal.
Show the full diagnosis — check getBuildInformation, warm up, compare profiled totals across targets, and read the per-layer vector for fallback layers and their transfer costs.
Treat accelerator capability as a deployment contract: assert it at startup, pin the build that provides it, and require a measurement on the real target hardware before any GPU claim reaches a capacity plan.
## What the two setters actually do ``` net.setPreferableBackend(cv2.dnn.DNN_BACKEND_CUDA) net.setPreferableTarget(cv2.dnn.DNN_TARGET_CUDA) ``` **Backend** is the *implementation* that executes layers — OpenCV's own (`DNN_BACKEND_OPENCV`), the CUDA backend (`DNN_BACKEND_CUDA`), and others depending on build flags. **Target** is the *device* it executes on — `DNN_TARGET_CPU`, `DNN_TARGET_OPENCL`, `DNN_TARGET_OPENCL_FP16`, `DNN_TARGET_CUDA`, `DNN_TARGET_CUDA_FP16`. They are two dimensions, not one, and only certain combinations are valid: a CUDA target belongs with the CUDA backend, the OpenCL targets with the default OpenCV backend. Crucially, both are *preferences*. Ask for something the binary cannot provide and you get a log message and a silent fall back to the default CPU path. Your code runs. Your frames get processed. Your latency does not move, and if you never checked, you conclude the GPU is not worth it. ## Cause one: the build has no CUDA This is the answer in most real cases. The `opencv-python` and `opencv-contrib-python` wheels on PyPI are built for portability and do **not** include the CUDA DNN backend. Getting it means compiling OpenCV from source with CUDA and cuDNN available and the relevant CMake options turned on (`WITH_CUDA`, `OPENCV_DNN_CUDA`), then installing that build. The verification step, before anything else: dump `cv2.getBuildInformation()` and look at the section listing NVIDIA CUDA and cuDNN. It is a plain string, always available, and it settles the question in one line. Do this in the service's startup logs, not just once at your desk — the failure mode where a container image quietly ships the pip wheel instead of the custom build is exactly the one this catches. ## Cause two: an inconsistent or unsupported pairing Setting a CUDA target while leaving the backend at the default, or asking for an OpenCL target on a machine with no usable OpenCL runtime, both fall back. `DNN_TARGET_CUDA_FP16` deserves its own warning: half precision is only a win on GPUs with high FP16 throughput. On hardware without it, the conversions cost more than the arithmetic saves and the FP16 target is *slower* than FP32. Measure both rather than assuming smaller is faster. ## Cause three: per-layer fallback Even with a correct CUDA build, the backend does not implement every layer type. Unsupported layers execute on the CPU, which means the intermediate tensors move device→host and back around them. A network with one such layer in the middle can end up slower than pure CPU because it pays the transfer twice per frame. This is not visible in wall-clock totals alone — you need the per-layer breakdown. ## How to actually confirm `net.getPerfProfile()` returns the total inference time in ticks plus a per-layer timing vector; divide by `cv2.getTickFrequency()` for seconds. Three things to do with it: 1. **Warm up first.** The first `forward()` includes graph preparation and backend initialisation, so it is not representative. Discard it. 2. **Compare totals across targets.** Run the same warmed-up loop under CPU and under the target you want. If the numbers are identical, you are on the CPU either way. 3. **Read the per-layer vector.** A handful of layers dominating the time on a supposedly-GPU run points at fallback layers and their transfers. ## The judgment part On a small network at modest resolution, the CPU backend is often competitive, because OpenCV's CPU implementations are well optimised and the fixed costs of dispatch and transfer do not amortise over so little work. OpenCL targets in particular frequently lose to CPU on integrated GPUs. The senior instinct here is not "turn on CUDA", it is "prove the change with a warmed-up measurement on the target hardware, and make the build capability an asserted startup check rather than an assumption".
- How do you verify at runtime that the installed OpenCV has the CUDA DNN backend?Print `cv2.getBuildInformation()` and inspect the section covering NVIDIA CUDA and cuDNN. It is a plain string available in every build, so it works as a startup assertion in a service — which is what catches a container image that quietly shipped the PyPI wheel instead of your custom build.
- Why can DNN_TARGET_CUDA_FP16 be slower than DNN_TARGET_CUDA?Half precision only pays off on GPUs with high FP16 throughput. Elsewhere the conversion work around each layer costs more than the reduced arithmetic saves, so the FP16 target loses. It is an empirical question on the exact deployment hardware, not a general rule that smaller precision is faster.
- A correct CUDA build still shows a small speedup. How do you find out why?Take `net.getPerfProfile()` after a warm-up pass and read the per-layer timing vector. Layers the CUDA backend does not implement run on the CPU, forcing device-to-host transfers around them. A few layers dominating the profile identifies the fallback points; the fix is usually re-exporting the model without those operators.
- Why is the first forward pass excluded from any of these measurements?Buffer allocation, layer fusion and backend initialisation all happen on the first `forward()` rather than at load. Including it inflates the measurement by a fixed cost that never recurs, which can make a genuinely faster target look slower. Run a dummy frame through, discard it, then measure a loop.
saying these in an interview costs you the question
- Assuming the pip opencv-python wheel includes the CUDA DNN backend
- Expecting an exception when the requested target is unavailable
- Treating backend and target as a single setting
- Believing FP16 is always faster than FP32 on any GPU
- Benchmarking with the first forward pass included