skip to content

How do you use cv2.findHomography with RANSAC to reject outlier matches?

level: seniorimportance: must knowfreq 55%

answer

  1. four points minimum, eight degrees of freedom
  2. two returns, not one
  3. threshold is in pixels
  4. count inliers, not matches
  5. planar scene or pure rotation

basics

~20 s

Pass matched point pairs as (N,1,2) float32 arrays with method cv2.RANSAC and a reprojection threshold in pixels. The call returns the 3x3 homography plus an inlier mask; the inlier count from that mask, not the raw match count, is the signal that the fit is trustworthy.

solid answer

~50 s

Build `src_pts` and `dst_pts` from the surviving matches — `kp1[m.queryIdx].pt` and `kp2[m.trainIdx].pt`, reshaped to `(-1, 1, 2)` float32 — then call `H, mask = cv2.findHomography(src_pts, dst_pts, cv2.RANSAC, 5.0)`. Four correspondences are the mathematical minimum, and RANSAC repeatedly samples four, fits a candidate homography, and counts how many of the remaining pairs it maps within `ransacReprojThreshold` pixels; the model with the most inliers wins and is refined on them. The returned `mask` is an `(N,1)` uint8 array with 1 for inliers, so `int(mask.sum())` is your quality metric — a plausible-looking `H` backed by six inliers out of four hundred matches is a failed fit, not a good one. `maxIters` and `confidence` bound the sampling. Also check that `H` is not `None`, and remember a homography is only valid for a planar scene or a purely rotating camera.

code

python · 12 lines
python
import cv2
import numpy as np

src = np.float32([[0, 0], [100, 0], [100, 100], [0, 100],
                  [50, 20], [20, 50], [80, 80]]).reshape(-1, 1, 2)
H_true = np.float32([[1.0, 0.2, 10.0], [0.1, 1.0, 5.0], [0.0, 0.0, 1.0]])
dst = cv2.perspectiveTransform(src, H_true)
dst[6] = [[900.0, 900.0]]                       # one gross outlier

H, mask = cv2.findHomography(src, dst, cv2.RANSAC, 3.0)
print(mask.ravel())                             # [1 1 1 1 1 1 0]
print(int(mask.sum()), "inliers of", len(src))

go deeper

for a junior

Know the call shape: two (N,1,2) float32 point arrays, cv2.RANSAC, a pixel threshold, and that it returns both the 3x3 matrix and an inlier mask.

for a middle

Explain the sample-fit-count-refine loop, why four correspondences are the minimum for eight degrees of freedom, and what ransacReprojThreshold, maxIters and confidence each control.

for a senior

Demonstrate validation habits: check H is not None, read the inlier count and ratio, guard against degenerate collinear inliers by projecting corners, and recognise when a homography is simply the wrong model for the scene.

for a principal

Own the robustness budget end to end — how much outlier rejection belongs in the ratio test versus the estimator, when to move to USAC or LMEDS, and what inlier thresholds gate an automated pipeline from accepting a bad registration.

## What a homography is and when it applies A homography is a 3x3 matrix mapping points from one image plane to another in homogeneous coordinates, up to scale — eight degrees of freedom, which is why four point correspondences (two equations each) are the minimum. It exactly describes the relationship between two views in only two situations: the scene points all lie on a **plane**, or the camera **rotates about its optical centre without translating**. Panorama stitching relies on the second; document scanning, marker tracking and aerial-to-map registration rely on the first. If neither holds — a camera that translates through a scene with depth — no homography fits, and forcing one produces a matrix that satisfies the dominant plane and mangles everything else. Two-view geometry with depth needs the fundamental or essential matrix instead (`cv2.findFundamentalMat`), not this function. ## Preparing the inputs OpenCV wants both point sets as float32 arrays of shape `(N, 1, 2)`, in corresponding order: row i of `src` maps to row i of `dst`. The standard construction from matches is: src = np.float32([kp1[m.queryIdx].pt for m in good]).reshape(-1, 1, 2) dst = np.float32([kp2[m.trainIdx].pt for m in good]).reshape(-1, 1, 2) Getting `queryIdx` and `trainIdx` the wrong way round produces the inverse mapping — the warp goes the wrong direction and the result looks confidently broken. ## What RANSAC does here With `method=cv2.RANSAC`, the estimator does not fit all the points. It: 1. Samples a random minimal subset of 4 correspondences. 2. Fits the exact homography through them. 3. Maps every `src` point with that candidate and counts how many land within `ransacReprojThreshold` pixels of their `dst` partner. Those are the inliers. 4. Repeats up to `maxIters` times (default 2000), keeping the model with the largest inlier set, stopping early once `confidence` (default 0.995) is reached. 5. Re-estimates the final homography using **all** inliers, so the answer is not merely the best random quadruple. This is why the estimator survives a match list that is half wrong: outliers never form a consistent large inlier set, so they cannot win the vote. ## The parameters that matter - **ransacReprojThreshold** is in **pixels of reprojection error**, defaulting to 3. It should reflect your keypoint localisation noise and image resolution: 1 to 3 for clean, high-resolution, well-localised features; 5 to 10 for noisy detections or heavily downscaled images. Too small and every model looks bad, so the inlier count collapses; too large and outliers get counted as inliers, dragging the refined fit off. - **maxIters** caps the sampling. If the true inlier ratio is low, the probability of drawing four clean points at random falls steeply, and 2000 iterations may not be enough. - **confidence** is the desired probability of finding a good model and governs early termination. ## Alternatives to the method argument Passing `0` fits least squares over all points, which any single outlier ruins — only use it when the correspondences are already known-clean. `cv2.LMEDS` needs no threshold but assumes more than half the matches are inliers. `cv2.RHO` is a faster variant that does well at low inlier ratios. OpenCV 4.5 and later, including OpenCV 5, also expose the USAC family such as `cv2.USAC_MAGSAC`, which is generally more accurate and less threshold-sensitive than plain RANSAC and is worth benchmarking on your data. ## Reading the mask, and validating the result The second return value is an `(N, 1)` uint8 mask. It is the diagnostic that separates a real fit from a lucky one: - `H is None` — estimation failed outright, usually too few points or a degenerate configuration. - Low absolute inlier count (single digits) — the model was fit on almost nothing. Do not trust it, whatever it looks like. - Low inlier *ratio* — matching is producing noise; look upstream at the descriptor norm, the ratio test, or whether the images overlap at all. Even a healthy inlier count can hide a degenerate case: if all inliers are collinear or clustered in a small image region, the homography is unconstrained elsewhere and extrapolates absurdly. A cheap, effective sanity check is to push the source image corners through with `cv2.perspectiveTransform(corners, H)` and verify the result is still a convex, sensibly sized quadrilateral rather than a self-intersecting bowtie or something a thousand pixels off-frame. Checking that the determinant of the top-left 2x2 block is positive catches mirror flips. Once `H` is trusted, `cv2.warpPerspective` applies it, and `cv2.drawMatches` with `matchesMask=mask.ravel().tolist()` visualises exactly which correspondences the estimator kept — the single most useful debugging image in this whole pipeline.

  • Four hundred matches produced a homography with nine inliers. What do you conclude?
    That the fit is not trustworthy, regardless of how the matrix prints. Nine inliers is barely above the four-point minimum, so RANSAC found a random quadruple that a few noisy points happened to agree with. The real problem is upstream — wrong matcher norm, a too-permissive ratio test, or images that do not actually overlap.
  • How do you pick ransacReprojThreshold?
    From keypoint localisation noise at your working resolution. Well-localised features in a full-resolution image justify 1 to 3 pixels; noisy detections or heavily downscaled input need 5 to 10. Too tight and no model gathers enough inliers; too loose and genuine outliers get voted in, biasing the final refinement fit.
  • When is a homography the wrong model entirely?
    When the camera translates through a scene with real depth. A homography has only eight degrees of freedom and cannot express parallax, so it will lock onto the dominant plane and misplace everything at other depths. Two-view geometry with depth calls for the fundamental or essential matrix instead.
  • Why can a homography be wrong even with a high inlier count?
    Because of degenerate configurations. If the inliers are collinear or crammed into one small image region, the transform is unconstrained everywhere else and extrapolates wildly. Project the image corners with cv2.perspectiveTransform and check the result is still a sensible convex quadrilateral before trusting it.

RANSAC is a vote by random juries: draw four correspondences, ask the rest of the crowd whether they agree with the story those four tell, and keep whichever story wins the most agreement.

saying these in an interview costs you the question

  • Judging the fit by match count instead of inlier count
  • Swapping queryIdx and trainIdx and inverting the mapping
  • Treating ransacReprojThreshold as a normalized value
  • Fitting a homography to a scene with real depth
  • Ignoring that findHomography can return None

context