skip to content

What does cv2.dnn.blobFromImage do to an image before inference?

level: middleimportance: must knowfreq 74%

answer

  1. One call replaces the transform pipeline
  2. Four knobs: size, scale, mean, channel swap
  3. Order matters between the two arithmetic steps
  4. Output is channels-first, not channels-last
  5. swapRB defaults to off; imread gives BGR

basics

~10 s

cv2.dnn.blobFromImage resizes the image to the network's input size, subtracts a per-channel mean, multiplies by scalefactor, optionally swaps the red and blue channels, and returns a 4-D NCHW float array ready for setInput.

solid answer

~40 s

`blobFromImage(image, scalefactor, size, mean, swapRB, crop, ddepth)` is OpenCV's whole preprocessing pipeline in one call. It resizes the BGR image to `size`, applies `scalefactor * (pixel - mean)` per channel, optionally reorders channels when `swapRB=True`, and packs the result into a 4-D blob shaped `(1, C, H, W)` — channels-first, unlike the `(H, W, C)` Mat you started with. Two details cause most silent bugs. First, the mean is subtracted **before** the scale factor is applied, so a model trained with `pixel/255` then ImageNet normalisation cannot be reproduced by mean+scale alone. Second, `swapRB` defaults to `False`, and the mean triple is interpreted in the post-swap order — so a model trained on RGB needs `swapRB=True` and RGB-ordered means. Getting either wrong does not raise; the network just returns confident nonsense.

code

python · 14 lines
python
import cv2
import numpy as np

frame = np.zeros((480, 640, 3), dtype=np.uint8)

blob = cv2.dnn.blobFromImage(
    frame,
    scalefactor=1 / 255.0,
    size=(224, 224),
    mean=(0, 0, 0),
    swapRB=True,
    crop=False,
)
print(blob.shape)  # (1, 3, 224, 224)

go deeper

for a junior

Know the four things the call does — resize, mean subtract, scale, optional channel swap — and that its output goes straight into net.setInput.

for a middle

Be ready to state the arithmetic as scalefactor times (pixel minus mean), name the NCHW output shape, and explain why swapRB defaulting to false bites anyone using imread with an RGB-trained model.

for a senior

Show how you verify preprocessing against the model's training recipe rather than guessing, and describe debugging a silently wrong pipeline by inspecting the blob's value range and channel statistics.

for a principal

Own the convention: preprocessing parameters are part of the model contract and should ship as data beside the weights, not as constants scattered through inference code, so a model swap cannot silently break the numbers.

## What the function is for Every deep network expects its input in a very specific form: a fixed spatial size, a particular channel order, a particular numeric range, and a particular memory layout. In PyTorch or TensorFlow you express that as a chain of transforms. OpenCV's DNN module compresses it into a single call, `cv2.dnn.blobFromImage`, because the module's whole selling point is that you can run inference with nothing in the build but OpenCV itself. ## The signature and the defaults ``` blobFromImage(image, scalefactor=1.0, size=Size(), mean=Scalar(), swapRB=False, crop=False, ddepth=CV_32F) ``` - **scalefactor** — a single multiplier applied to every channel. For a network trained on `[0, 1]` inputs this is `1/255.0`. - **size** — the spatial size `(width, height)` the network expects. Note the OpenCV convention: width first, even though the resulting blob is `(N, C, height, width)`. - **mean** — a scalar triple subtracted per channel. - **swapRB** — whether to exchange the first and third channels. - **crop** — how the aspect ratio is handled (below). - **ddepth** — output depth, `CV_32F` or `CV_8U`. ## The order of operations The arithmetic is `output = scalefactor * (input - mean)`. **Mean subtraction happens first, then the multiply.** This matters because the two most common normalisation recipes are not the same shape: - Caffe-era models: subtract `(104, 117, 123)` in BGR, scale `1.0`. This maps cleanly onto mean+scale. - Torch/ImageNet models: divide by 255, then subtract `(0.485, 0.456, 0.406)` and divide by `(0.229, 0.224, 0.225)`. The subtraction here comes *after* the division and the divisor is per-channel, which `blobFromImage` cannot express. You reproduce it either by folding the maths yourself (`mean * 255` with a matching single scale, accepting the per-channel std is approximated) or by doing the normalisation in NumPy and passing an already-normalised float image with `scalefactor=1.0, mean=(0,0,0)`. ## Channel order `cv2.imread` and `VideoCapture` hand you **BGR**. Most published ONNX and TensorFlow models were trained on **RGB**. `swapRB=False` is the default, so if you say nothing you feed BGR to an RGB model. The failure is silent: shapes match, the forward pass succeeds, confidences look plausible, and the predictions are wrong in a way that gets blamed on the model. The mean triple is interpreted **after** the swap. The documentation states the values are intended to be in `(mean-R, mean-G, mean-B)` order when the input is BGR and `swapRB` is true. So flipping `swapRB` without reordering your mean introduces a second bug that partially masks the first. ## Resizing and crop With `crop=False` (the default) the image is resized directly to `size`, distorting the aspect ratio if it does not match. With `crop=True` the image is resized so the *smaller* side reaches the target and then centre-cropped, preserving aspect ratio at the cost of throwing pixels away at the edges. Detection pipelines almost always want `crop=False`, because a centre crop can delete the object you were asked to find, and because normalised box coordinates then map straight back onto the full original frame. Classification pipelines that were evaluated with a centre crop want `crop=True`. Neither mode letterboxes. If your model was trained with letterbox padding (common for YOLO-family models), you must pad the frame yourself before calling `blobFromImage`, or use the parameterised variant of the call that exposes a padding mode. ## The output blob The return value is a 4-D NumPy array of shape `(N, C, H, W)` — batch, channels, height, width. This is NCHW, the layout the DNN module works in internally, regardless of whether the original model was a channels-last TensorFlow graph; the importer handles the transposition. `cv2.dnn.blobFromImages` takes a list and produces `N > 1`, which is how you batch. `cv2.dnn.imagesFromBlob` reverses the packing, which is useful when debugging what actually went into the network. ## Why this is an interview question Because it is the single most common source of "the model works in Python but not in my OpenCV app". Everything about the failure is quiet — no exception, no shape mismatch, no warning. A candidate who has actually shipped an OpenCV DNN pipeline will immediately reach for the four knobs (size, scale, mean, swapRB) and say that they must be copied from whatever the model card or the training preprocessing specified, not guessed.

  • How would you reproduce ImageNet normalisation with per-channel standard deviations here?
    You cannot express per-channel divisors with a single `scalefactor`. Either normalise in NumPy first and pass the float image with `scalefactor=1.0` and a zero mean, or accept an approximation by folding a single averaged std into the scale. Most teams do the NumPy path because it is exact and the cost is negligible next to the forward pass.
  • What changes if you pass a list of frames to blobFromImages instead?
    You get the same per-image preprocessing but a blob with `N` equal to the list length, so one `forward()` runs the whole batch. All images must share the target `size`. Batching helps throughput on GPU targets; on CPU with a small network the gain is usually modest because the module already parallelises across the layer.
  • Your detector's boxes are systematically squashed horizontally. What preprocessing setting would you check?
    The aspect-ratio handling. With `crop=False` the frame is stretched to the network's input size, which is fine if the model was trained that way but wrong if it was trained on letterboxed input. Either pad the frame to the training aspect ratio yourself before the call, or switch to the crop behaviour the model expects.

saying these in an interview costs you the question

  • Assuming blobFromImage scales first and then subtracts the mean
  • Thinking swapRB defaults to true because imread returns BGR
  • Expecting an NHWC blob because the source Mat is HxWxC
  • Believing a wrong mean or scale will raise an error
  • Passing size as (height, width) instead of (width, height)

context