skip to content

Why chain resize, warpAffine and warpPerspective into one warp per frame?

level: seniorimportance: should knowfreq 35%

answer

  1. each pass resamples the last pass
  2. interpolation error compounds
  3. all three are 3x3 matrices
  4. multiply right to left
  5. fixed transform, precompute the maps

basics

~20 s

Each warp resamples the previous warp's output, so interpolation blur compounds and you pay three passes over the image. Compose the matrices into a single 3x3 and warp once, or precompute cv2.remap maps when the transform is fixed.

solid answer

~50 s

Every call to `cv2.resize`, `cv2.warpAffine` or `cv2.warpPerspective` resamples the image, and resampling is lossy: the second warp interpolates values that were themselves interpolated, so edges soften a little more at each step and the errors do not cancel. Three calls also mean three full passes over pixels and two throwaway intermediate buffers. Because all three are projective transforms, they compose: pad each 2x3 affine matrix with the row `[0, 0, 1]`, multiply the 3x3 matrices in the order the transforms apply (rightmost first), and run one `cv2.warpPerspective` with the product and the final `dsize`. One interpolation, one pass. If the transform is the same for every frame — a fixed lens correction or a fixed homography onto a ground plane — go further and precompute the sampling grid once with `cv2.initUndistortRectifyMap` or your own coordinate arrays, then call `cv2.remap` per frame.

code

python · 13 lines
python
import cv2, numpy as np

src = np.zeros((480, 640, 3), np.uint8)

S = np.array([[0.5, 0, 0], [0, 0.5, 0], [0, 0, 1]], np.float32)      # half size
R = np.vstack([cv2.getRotationMatrix2D((160.0, 120.0), 30.0, 1.0),
               [0, 0, 1]]).astype(np.float32)                         # then rotate

M = R @ S                                                             # rightmost applies first
out = cv2.warpPerspective(src, M, (320, 240),
                          flags=cv2.INTER_LINEAR,
                          borderMode=cv2.BORDER_CONSTANT)
print(M.shape, out.shape)

go deeper

for a junior

Know that warping resamples the image and that each extra warp costs quality as well as time, so prefer doing the geometry in one call when you can.

for a middle

Be able to build the composite: pad a 2x3 affine to 3x3, multiply right to left, choose dsize from the transformed corner bounding box, and set borderMode deliberately.

for a senior

Show the production reasoning — precompute cv2.remap maps for a fixed transform, keep an INTER_AREA downscale as its own step because warps cannot area-average, and know that WARP_INVERSE_MAP skips the internal inversion.

for a principal

Own the geometry contract across the system: one canonical transform per camera, computed once and versioned, so calibration, training crops and serving all warp identically instead of each stage inventing its own chain.

## Resampling is lossy and it accumulates A geometric warp does not move pixels; it builds an output grid and, for each destination pixel, works out where that pixel came from in the source and samples there. The source coordinate is almost never an integer, so the value is interpolated from neighbours — bilinear by default. Interpolation is a low-pass operation: it discards a little high-frequency detail every time. Run three warps in sequence and the third one is interpolating values that were already interpolated twice. Fine texture softens, thin lines fade, and small geometric errors from each stage add up rather than cancelling. The visible result is a pipeline whose output is noticeably blurrier than a single equivalent warp, with no error and no obvious culprit. ## Warps compose because they are matrices All three operations are projective (homographic) transforms of the plane: - A uniform `cv2.resize` by factors `fx, fy` is the 3x3 diagonal matrix `[[fx,0,0],[0,fy,0],[0,0,1]]`. - An affine 2x3 matrix `M` from `cv2.getRotationMatrix2D` or `cv2.getAffineTransform` becomes 3x3 by appending the row `[0, 0, 1]`. - A perspective matrix from `cv2.getPerspectiveTransform` is already 3x3. Matrix multiplication then expresses the whole chain. Order matters and it is right-to-left: if you scale, then rotate, then apply perspective, the composite is `M = P @ R @ S`. Feed the product to a single `cv2.warpPerspective(src, M, dsize)`. If the composite's bottom row is still `[0, 0, 1]` — that is, no genuine perspective component — you can drop back to `cv2.warpAffine` with the top two rows. ## Getting the output framing right The most common surprise when doing this is content falling outside the output. `dsize` is `(width, height)` and it is the size of the *destination canvas*, not something derived from the source. `cv2.getRotationMatrix2D(center, angle, scale)` rotates about a point but adds no translation, so rotating a rectangle inside its original canvas cuts the corners off. The fix is to transform the four source corners through your composite matrix, take the bounding box of the results, add a translation into `M[0,2]` and `M[1,2]` so the minimum corner lands at the origin, and set `dsize` to the bounding box size. Pixels of the destination with no valid source are filled per `borderMode`: `cv2.BORDER_CONSTANT` (the default, using `borderValue`, black unless set), `cv2.BORDER_REPLICATE`, `cv2.BORDER_REFLECT_101`, and so on. Composing warps has a quiet benefit here too — with three sequential warps, the black border introduced by stage one becomes real image content that stage two happily interpolates against, bleeding dark fringes inward. ## Direction of the mapping By default `cv2.warpAffine` and `cv2.warpPerspective` treat `M` as the forward, source-to-destination transform, and invert it internally so the sampling loop can run destination-driven (which is what avoids holes in the output). Pass `flags=cv2.WARP_INVERSE_MAP` when the matrix you hold is already the destination-to-source map and you want to skip the inversion. This is also why the maps handed to `cv2.remap(src, map1, map2, interpolation)` are inverse maps: for each destination pixel they give the *source* coordinate to sample, as `CV_32FC1` x and y arrays (or the fixed-point `CV_16SC2` pair). ## When to switch to remap If the transform is constant across frames, the per-frame matrix inversion and per-pixel coordinate arithmetic are pure repeated work. Building the two coordinate maps once and calling `cv2.remap` per frame moves that cost out of the loop and leaves only the sampling. Lens undistortion is the canonical case: `cv2.initUndistortRectifyMap` produces exactly these maps once from the camera intrinsics, and `cv2.remap` applies them per frame. `remap` also handles transforms that are not expressible as a matrix at all — barrel distortion, arbitrary mesh warps, fisheye unwrapping — which a composed homography cannot. ## What to keep separate Composition is not always the answer. If the net effect is a large reduction, one composed warp with bilinear sampling still aliases, because warping does not offer an area-averaging mode the way `cv2.resize` does with `cv2.INTER_AREA`. In that case do the antialiased downscale as its own `cv2.resize(..., interpolation=cv2.INTER_AREA)` step — ideally first, so the expensive warp then runs on far fewer pixels — and compose the rest. Two passes chosen deliberately beat one pass that aliases. That trade is the real senior content here: know that warps compose, know why fewer resampling steps is better, and know the one case where splitting is still correct.

  • Why does rotating with cv2.getRotationMatrix2D crop the corners, and how do you avoid it?
    The matrix rotates about a centre point but adds no translation, and warpAffine writes into a canvas of the dsize you asked for — usually the original size — so anything rotated outside it is clipped. Transform the four corners through M, take their bounding box, add the offset into M[0,2] and M[1,2], and pass the bounding-box size as dsize.
  • In what order do you multiply the matrices?
    Right to left, in application order: to scale, then rotate, then apply perspective, compute P @ R @ S. Each matrix acts on the coordinates produced by the one to its right. Reversing the order silently produces a valid-looking but wrong warp, so verify by pushing a couple of known points through the composite with cv2.perspectiveTransform.
  • When is cv2.remap the better tool than a composed matrix?
    When the transform is fixed across frames, so the coordinate maps can be built once and only the sampling runs per frame — lens undistortion via cv2.initUndistortRectifyMap is the classic case. Also when the transform is not a matrix at all: barrel distortion, fisheye unwrapping or an arbitrary mesh warp cannot be expressed as a homography.
  • Is composing always right if the pipeline includes a big downscale?
    No. Warping samples with bilinear or bicubic interpolation and has no area-averaging mode, so folding a 10x reduction into the warp aliases. Do that reduction as a separate cv2.resize with cv2.INTER_AREA, ideally first so the warp then processes far fewer pixels, and compose the remaining transforms into a single pass.

saying these in an interview costs you the question

  • Assumes repeated warps are lossless because they are just geometry
  • Multiplies the matrices in the wrong order
  • Passes the source size as dsize and wonders where the corners went
  • Folds a large downscale into the warp and gets moire
  • Rebuilds remap maps every frame for a fixed transform

context