skip to content

Filtering and Transforms

Blurring with Gaussian, median, and bilateral filters, edge detection via Sobel, Laplacian, and Canny, and geometric warps through affine and perspective matrices. Interviewers ask why bilateral preserves edges and which interpolation mode a resize should use.

on this pageshow

questions

6

Which cv2.resize interpolation flag should you use when downscaling, and why?

level: juniorimportance: must knowfreq 72%

answer

  1. direction decides the flag
  2. shrinking is averaging, not sampling
  3. most source pixels never get read
  4. moire on fine texture
  5. masks never get averaged

basics

~20 s

Use cv2.INTER_AREA. It averages over the source pixel area that maps to each output pixel, so detail is integrated rather than skipped. The default cv2.INTER_LINEAR samples only a few points and produces aliasing and moire on large reductions.

solid answer

~40 s

For shrinking, pass `interpolation=cv2.INTER_AREA` to `cv2.resize`. It resamples using the pixel-area relation: each destination pixel is the average of the whole source region that maps onto it, which is the antialiasing you need when many source pixels collapse into one. The default, `cv2.INTER_LINEAR`, interpolates from only the four nearest source pixels regardless of the scale factor, so on a 10x reduction it is effectively sampling and most of the image never contributes — you get aliasing, moire on fine textures, and unstable results frame to frame. For enlarging, `INTER_AREA` behaves roughly like nearest-neighbour, so switch to `cv2.INTER_LINEAR` (fast) or `cv2.INTER_CUBIC` / `cv2.INTER_LANCZOS4` (slower, sharper). For a label or segmentation mask use `cv2.INTER_NEAREST` in either direction, because averaging class ids invents ids that mean nothing.

code

python · 11 lines
python
import cv2, numpy as np

# fine vertical stripes: the classic aliasing target
img = np.zeros((1000, 1000), np.uint8)
img[:, ::2] = 255

area   = cv2.resize(img, (100, 100), interpolation=cv2.INTER_AREA)
linear = cv2.resize(img, (100, 100), interpolation=cv2.INTER_LINEAR)

print("area  std:", round(float(area.std()), 2))    # near 0: stripes averaged to grey
print("linear std:", round(float(linear.std()), 2))  # high: aliased pattern survives

go deeper

for a junior

Memorise the two rules: INTER_AREA when making an image smaller, INTER_LINEAR or INTER_CUBIC when making it bigger, and remember dsize is (width, height).

for a middle

Explain why — decimation needs averaging over the mapped source region while interpolation estimates between samples — and know the mask exception where only INTER_NEAREST is safe.

for a senior

Watch for the cross-library seam: state that preprocessing at inference must match training, and be able to point at aliasing artefacts in a pipeline as the cause of unstable downstream results.

for a principal

Treat resize policy as part of the data contract between capture, training and serving, and make it explicit and tested rather than a per-call flag anyone can change.

## What resizing actually has to do `cv2.resize(src, dsize, fx=..., fy=..., interpolation=...)` builds an output grid and has to decide what value each output pixel takes. There are two very different regimes: - **Upscaling**: output pixels are denser than input pixels. Each output pixel falls between known samples, and the job is *interpolation* — estimate a plausible value between neighbours. - **Downscaling**: many input pixels map into one output pixel. The job is *decimation*, and the correct operation is to integrate (average) over the region, not to interpolate at a point. Sampling a signal more sparsely than its detail warrants folds high frequencies down into fake low frequencies — aliasing — which is what makes a shrunken picket fence or striped shirt shimmer with moire patterns. That asymmetry is why one flag is not right for both directions. ## The flags - `cv2.INTER_NEAREST` — copies the closest source pixel. Fast, blocky, invents no new values. The right choice for label maps, index images and any data where the numbers are categories rather than intensities. `cv2.INTER_NEAREST_EXACT` is a variant with different rounding at exact half-pixel positions. - `cv2.INTER_LINEAR` — bilinear from the 4 nearest pixels. This is the **default**, a good general choice for upscaling and mild changes, and the wrong choice for big reductions. `cv2.INTER_LINEAR_EXACT` uses a bit-exact computation. - `cv2.INTER_CUBIC` — bicubic over a 4x4 neighbourhood. Sharper than bilinear when enlarging, slower, and it can overshoot into slight ringing at strong edges. - `cv2.INTER_AREA` — pixel-area resampling, the documented preferred method for image decimation. When enlarging it behaves similarly to nearest-neighbour, so it is not a general-purpose default. - `cv2.INTER_LANCZOS4` — an 8x8 windowed sinc, the sharpest and slowest of the common options for enlargement. ## The practical rules 1. **Shrinking?** `INTER_AREA`. 2. **Enlarging and speed matters?** `INTER_LINEAR`. **Enlarging and quality matters?** `INTER_CUBIC` or `INTER_LANCZOS4`. 3. **Masks, label maps, id images?** `INTER_NEAREST`, always, in both directions. Bilinearly resizing a mask whose values are class ids 0, 1, 7 produces values like 3.5 that correspond to no class at all, and after rounding you get phantom regions of a class that was never there. 4. **Feeding a neural network?** Match whatever the training preprocessing used, including the library. Resize implementations differ in filter support and pixel-centre conventions, and a mismatch is a real, measurable accuracy loss that is invisible in the code. ## The dsize gotcha `dsize` is `(width, height)` — the opposite order of `img.shape`, which is `(rows, cols)` = `(height, width)` for a grayscale image and `(rows, cols, channels)` for colour. Writing `cv2.resize(img, img2.shape[:2])` silently transposes the target dimensions unless the target happens to be square. The safe idioms are `cv2.resize(img, (w, h))` with named variables, or `cv2.resize(img, dsize=(0, 0), fx=0.5, fy=0.5)` to scale by factors and let OpenCV compute the size. Note that when you use `fx`/`fy`, `dsize` must be `(0, 0)` or `None`. ## Extreme reductions Even `INTER_AREA` in a single step is not always the best answer for very large factors on very noisy input; a classic alternative is repeated halving with `cv2.pyrDown`, which applies a Gaussian before each 2x decimation, giving a smooth multi-step reduction. In practice `INTER_AREA` is close enough for a single large step and is far simpler. ## The interview signal The weak answer is "INTER_CUBIC, because it is the highest quality". Quality is direction-dependent: bicubic on a 10x reduction is *worse* than area averaging because it still ignores most of the source. The strong answer names the direction first, explains aliasing versus interpolation, and adds the mask exception unprompted.

  • Why is cv2.INTER_NEAREST the right choice for a segmentation mask?
    Mask values are class ids, not intensities. Bilinear or area interpolation averages them, so ids 0 and 7 produce an in-between value that names no class, and rounding turns it into a phantom region. INTER_NEAREST copies an existing label, guaranteeing the output values are a subset of the input values.
  • What is the argument order of dsize in cv2.resize?
    (width, height), which is the reverse of img.shape's (rows, cols). Passing a shape tuple directly transposes the result for any non-square image and is a frequent bug. Either build the tuple explicitly from named width and height variables, or pass dsize=(0, 0) with fx and fy scale factors instead.
  • Someone downscales with INTER_CUBIC because it is the highest-quality flag. What do you tell them?
    Quality depends on direction. Bicubic still samples a 4x4 neighbourhood around one point, so at a 10x reduction the vast majority of source pixels contribute nothing and fine detail aliases into moire. INTER_AREA averages the full source region per output pixel, which is both cheaper and visibly better for decimation.
  • Does the choice matter when resizing inputs for a trained model?
    Yes. Match the preprocessing used at training time, including which library did the resize. Implementations differ in filter support and pixel-centre convention, so a model trained on area-downscaled crops and served bilinear-downscaled ones sees a subtly different distribution — a measurable accuracy drop with no error anywhere in the code.

saying these in an interview costs you the question

  • Uses INTER_CUBIC for everything because it is highest quality
  • Passes img.shape as dsize and transposes the image
  • Bilinearly resizes a label mask
  • Thinks the default INTER_LINEAR is safe at any scale
  • Ignores which resize the model was trained with

context

open as a page

Why does cv2.bilateralFilter preserve edges when cv2.GaussianBlur does not?

level: middleimportance: must knowfreq 62%

basics

~20 s

cv2.GaussianBlur weights neighbours only by distance, so it averages straight across edges. cv2.bilateralFilter multiplies that spatial weight by a second weight based on intensity difference, so pixels on the far side of an edge contribute almost nothing.

open as a page

What do the two threshold arguments of cv2.Canny control?

level: middleimportance: must knowfreq 70%

basics

~20 s

They drive hysteresis. Pixels with gradient magnitude above the upper threshold are accepted as strong edges; below the lower they are discarded; in between they are kept only if they connect to a strong edge. A common ratio is 1:2 to 1:3.

open as a page

Why is cv2.medianBlur better than cv2.GaussianBlur for salt-and-pepper noise?

level: juniorimportance: should knowfreq 52%

basics

~10 s

Salt-and-pepper noise is isolated extreme pixels. cv2.medianBlur takes the middle value of each neighbourhood, so an outlier is discarded outright. cv2.GaussianBlur averages the outlier in, spreading one bad pixel into a visible grey smudge.

open as a page

Why does cv2.Sobel with ddepth=cv2.CV_8U lose half of the edges?

level: middleimportance: should knowfreq 48%

basics

~10 s

A derivative is signed: bright-to-dark transitions give negative values. CV_8U cannot hold negatives, so they saturate to 0 and one polarity of every edge vanishes. Compute into CV_16S/CV_32F/CV_64F, then convert with cv2.convertScaleAbs.

open as a page

Why chain resize, warpAffine and warpPerspective into one warp per frame?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Each warp resamples the previous warp's output, so interpolation blur compounds and you pay three passes over the image. Compose the matrices into a single 3x3 and warp once, or precompute cv2.remap maps when the transform is fixed.

open as a page