skip to content

What do the two threshold arguments of cv2.Canny control?

level: middleimportance: must knowfreq 70%

answer

  1. not one cut but two
  2. strong, weak, and rejected
  3. weak survives if connected
  4. recommended ratio 2:1 to 3:1
  5. blur it yourself first

basics

~20 s

They drive hysteresis. Pixels with gradient magnitude above the upper threshold are accepted as strong edges; below the lower they are discarded; in between they are kept only if they connect to a strong edge. A common ratio is 1:2 to 1:3.

solid answer

~40 s

`cv2.Canny(image, threshold1, threshold2)` uses the pair for **hysteresis thresholding**, the final step of the algorithm. After computing the gradient and thinning it with non-maximum suppression, every candidate pixel is classified: above the higher threshold it is a strong edge and is kept unconditionally; below the lower threshold it is dropped; between the two it is *weak* and survives only if it is connected, through a chain of weak pixels, to a strong one. That is what lets a real contour dim in the middle without breaking into fragments, while isolated noise responses at the same magnitude get rejected. Canny's own recommendation is a high:low ratio between 2:1 and 3:1. Two other arguments matter: `apertureSize` is the Sobel kernel used for the gradient, and `L2gradient=True` switches magnitude from `|dx|+|dy|` to `sqrt(dx^2+dy^2)`.

code

python · 10 lines
python
import cv2, numpy as np

img = np.zeros((200, 200), np.uint8)
img[50:150, 50:150] = 255
img += np.random.randint(0, 40, img.shape, dtype=np.uint8)

blur  = cv2.GaussianBlur(img, (5, 5), 1.4)      # Canny does not do this for you
edges = cv2.Canny(blur, 50, 150, apertureSize=3, L2gradient=True)

print(edges.dtype, sorted(np.unique(edges)))

go deeper

for a junior

Know that Canny takes a low and a high threshold, that the high one decides what is definitely an edge, and that you convert to grayscale and blur before calling it.

for a middle

Explain hysteresis precisely — strong, weak-but-connected, rejected — and know that apertureSize and L2gradient change the magnitude scale, so tuned thresholds do not transfer.

for a senior

Talk about stability across real data: normalise or equalise input rather than retuning per scene, derive thresholds from image statistics, and know that gaps in the edge map are a hysteresis consequence downstream stages must handle.

for a principal

Weigh whether an edge detector is the right primitive at all given lighting variance and the downstream consumer, and decide where robustness belongs — capture settings, preprocessing, or a learned detector.

## The pipeline behind the two numbers The Canny detector is a sequence of steps, and the thresholds only enter at the end: 1. **Smooth** the image, because differentiation amplifies noise. 2. **Compute the gradient**, magnitude and direction, with a Sobel operator. 3. **Non-maximum suppression**: walk along the gradient direction and keep a pixel only if its magnitude is a local maximum across the edge. This thins fat gradient ridges down to one-pixel lines. 4. **Hysteresis thresholding**: the two-threshold rule. An important practical detail: OpenCV's `cv2.Canny()` performs steps 2 to 4 for you but **does not smooth the input**. The Gaussian in step 1 is your job — call `cv2.GaussianBlur` first. Skipping it is the most common cause of a Canny output that looks like static, and turning the thresholds up to hide the static is how people lose their real edges too. ## Why one threshold is not enough With a single threshold you must choose between two failure modes. Set it low and every noise fluctuation with a modest gradient becomes an edge pixel, so the output is speckled. Set it high and real contours break up wherever they run through low-contrast regions — a shadowed section of an object outline drops below the cut and the edge becomes a dashed line, which then defeats any downstream contour tracing. Hysteresis resolves this by using connectivity as evidence. A moderate-magnitude pixel that is part of a continuous structure reaching up to a confidently strong pixel is probably a real edge. A moderate-magnitude pixel with no strong neighbour anywhere along its chain is probably noise. So `threshold2` (the higher one) controls *how confident a pixel must be to start an edge*, and `threshold1` (the lower one) controls *how faint an edge may become and still be followed*. Note that OpenCV does not require you to pass them in a particular order — it uses the larger of the two as the upper threshold — but every readable codebase passes low first, matching the parameter names. ## Choosing values There is no universal pair, because gradient magnitudes depend on contrast, on the Sobel aperture and on whether the input was blurred. Practical approaches: - Start from Canny's recommended ratio: upper between 2x and 3x the lower. - Derive them from image statistics — a widely used heuristic sets the pair around the median intensity, e.g. lower = 0.66 * median, upper = 1.33 * median. - Tune interactively on a representative sample, then check the extremes of your data (dark frames, glare) rather than the pretty one. - If a lighting change breaks the tuning, normalise the input first (histogram equalisation, or per-frame contrast normalisation) instead of chasing the thresholds per scene. ## The other arguments - `apertureSize` (default 3) is the Sobel kernel size for the gradient step. Larger apertures smooth more, respond to broader edges, and change the magnitude scale — so the thresholds you tuned at 3 will be wrong at 5. - `L2gradient` (default `False`) selects the magnitude formula. The default `|dx| + |dy|` is an L1 approximation that is cheaper; `True` uses the exact Euclidean `sqrt(dx^2 + dy^2)`. The L1 form overestimates diagonal gradients by up to about 40%, so switching this flag shifts the effective sensitivity and again invalidates tuned thresholds. - Input should be an 8-bit image; the usual call passes a single-channel grayscale image obtained with `cv2.cvtColor`. ## What comes out The result is an 8-bit single-channel map with 255 on edge pixels and 0 elsewhere — already binary, already thinned to roughly one pixel wide. That is why it feeds naturally into contour or line-finding stages. It is also why Canny is not a segmentation: the edges are not guaranteed closed, and gaps left by hysteresis are common at low-contrast junctions. The answer an interviewer is listening for is the connectivity idea. Anyone can say "two thresholds"; the signal is knowing that the middle band is decided by whether a pixel touches a strong edge, and being able to say what each threshold buys you when you move it.

  • What happens if you raise only the lower threshold?
    Fewer weak pixels qualify to continue an edge, so contours that dim in the middle break into fragments while the strong, high-contrast edges are untouched. Raising only the upper threshold does the opposite: fewer chains ever get started, so whole faint objects disappear even though their pixels would have been followable.
  • Does cv2.Canny blur the image for you?
    No. OpenCV's implementation computes the Sobel gradient, applies non-maximum suppression and runs hysteresis, but the noise-reduction Gaussian described in step one of the algorithm is left to the caller. Apply cv2.GaussianBlur yourself, otherwise sensor noise produces a speckled edge map that no threshold pair cleans up without also erasing real edges.
  • Why might edges detected at apertureSize=5 need different thresholds than at 3?
    The Sobel kernel is not normalised, so a larger aperture both smooths more and scales the gradient magnitudes up. Threshold values are absolute magnitudes, not percentiles, so the pair tuned for a 3x3 aperture becomes far too permissive at 5. The same applies when switching L2gradient, which changes the magnitude formula.
  • Is the Canny output guaranteed to give closed contours?
    No. Hysteresis drops any weak chain that never reaches a strong pixel, so low-contrast stretches of a real boundary can be missing entirely and the outline is left open. If a downstream step needs closed regions, either close the gaps morphologically or use a region-based segmentation rather than an edge detector.

It is like vetting a rumour: a loud, well-sourced claim is accepted outright, a whisper on its own is ignored, and a whisper is believed only when you can trace it back to the loud source.

saying these in an interview costs you the question

  • Says the two thresholds are just min and max intensity
  • Thinks the pixels between the thresholds are always kept
  • Feeds a raw noisy frame to Canny without blurring
  • Retunes thresholds per image instead of normalising input
  • Assumes Canny returns closed object outlines

context