skip to content

In OpenCV, what changes when you add cv2.THRESH_OTSU to cv2.threshold?

level: middleimportance: must knowfreq 62%

answer

  1. the thresh argument stops mattering
  2. the histogram decides
  3. the return value now tells you something
  4. assumes two humps
  5. still one number for the whole frame

basics

~20 s

OpenCV then ignores the threshold value you passed and computes one from the image histogram by maximizing between-class variance. The chosen value comes back as the first return value of cv2.threshold, which is why the call is written with 0 as the threshold argument.

solid answer

~40 s

`cv2.threshold` normally returns `(retval, dst)` where `retval` is just the threshold you supplied. Add the `cv2.THRESH_OTSU` flag - `cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)` - and the `thresh` argument is ignored; OpenCV scans the grayscale histogram and picks the cut that maximizes the variance between the two resulting pixel classes, then applies it with the base mode you combined it with. `retval` now carries that computed value, which is worth logging as a drift signal in production. The input must be single-channel grayscale. Otsu assumes a clearly **bimodal** histogram - foreground and background as two separated humps. It is also **global**: one threshold for the whole frame, so uneven illumination or a vignette defeats it and you want adaptive thresholding instead. `cv2.THRESH_TRIANGLE` is the alternative when one class dominates.

code

python · 11 lines
python
import cv2
import numpy as np

gray = np.full((100, 100), 30, np.uint8)
gray[30:70, 30:70] = 200          # a clearly bimodal histogram

t, binary = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
print(t)                          # the value Otsu chose; the 0 was ignored

plain_t, _ = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY)
print(plain_t)                    # 127.0 - just an echo of the argument

go deeper

for a junior

Know the idiom cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU) and be able to say that OpenCV picks the threshold for you from the image instead of you hard-coding 127.

for a middle

Explain that the thresh argument is ignored, that retval returns the computed value, and describe the between-class variance criterion and the bimodal-histogram assumption behind it.

for a senior

Demonstrate the operational angle: log retval as a drift signal, blur before thresholding to stabilise it, and recognise the half-black-frame symptom as a global-versus-local problem rather than a tuning problem.

for a principal

Own the decision of where automatic thresholding belongs at all - controlled enclosure and fixed exposure versus per-frame estimation - and what the failure budget is when a silently wrong mask propagates into counted or measured output.

## What plain thresholding does first `cv2.threshold(src, thresh, maxval, type)` compares every pixel against a fixed number. With `cv2.THRESH_BINARY`, pixels above `thresh` become `maxval` and the rest become 0; `cv2.THRESH_BINARY_INV` flips that; `THRESH_TRUNC`, `THRESH_TOZERO` and `THRESH_TOZERO_INV` clamp rather than binarize. The return is a tuple `(retval, dst)`, and in the plain case `retval` is simply the `thresh` you passed - uninteresting, which is why so much code writes `_, binary = cv2.threshold(...)`. The problem with a hard-coded number is that it encodes one camera, one lamp and one day. Change the exposure and 127 is wrong. ## What the Otsu flag adds `cv2.THRESH_OTSU` is not a threshold type of its own; it is a flag you **combine** with a base type using `+` or `|`: ``` t, binary = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU) ``` Two consequences follow. First, the `thresh` argument is ignored - passing 0 is pure convention, a signal to the reader that the value is not used. Second, `retval` stops being an echo and becomes real information: it is the threshold Otsu selected for this particular image. The algorithm builds the 256-bin histogram of the grayscale image and tries every possible cut. For each cut it splits pixels into two classes and computes the variance *between* the class means, weighted by class size. The cut that maximizes that between-class variance - equivalently, minimizes the intensity variance *within* the two classes - wins. Intuitively it finds the valley between two humps in the histogram. ## The assumptions you must be able to state **Single-channel input.** Otsu is defined over a one-dimensional intensity histogram, so you convert to grayscale first. A colour image is not accepted for the Otsu path. **Bimodality.** The method presumes two populations with a genuine valley between them - ink and paper, part and conveyor. If the histogram is unimodal (an almost blank scan, a frame where the object is absent, a scene with three distinct materials) Otsu still returns *a* number, confidently and wrongly, usually slicing a single mode in half and producing a mask full of texture noise. This is a silent failure with no exception raised. `cv2.THRESH_TRIANGLE` is the standard fallback for a heavily skewed, single-dominant-mode histogram. **Globality.** One threshold is applied to every pixel of the frame. A gradient across the field of view - a lamp on one side, vignetting, a curved page - guarantees that any single value is too high at one end and too low at the other. The symptom is the classic half-black scan. That is the boundary where you switch to a locally computed threshold. **Noise sensitivity.** Otsu reads the histogram, and per-pixel noise smears the two humps together. Smoothing the grayscale image before the call, typically with a small Gaussian blur, sharpens the valley and stabilizes the chosen value considerably. This is one of the most reliable cheap improvements to an Otsu pipeline. ## Using retval as a production signal Because `retval` is recomputed per frame, it is a free health metric. Log it. On a stable line it should sit in a narrow band; a slow drift means the lamp is aging or the lens is dirtying, and a sudden jump usually means an empty or occluded frame. Alarming on a threshold that leaves its expected range catches illumination faults before the downstream blob count goes wrong - and blob counts going quietly wrong is the expensive failure mode in industrial vision. You can also clamp: compute Otsu, and if the value falls outside a sane range fall back to a configured default rather than binarizing garbage. ## Where it sits in the pipeline The standard shape-analysis recipe is: grayscale, light smoothing, Otsu (or adaptive) threshold with `THRESH_BINARY_INV` if the objects are dark, a morphological opening to kill speckle, then `cv2.findContours`. Otsu's job is only to produce the binary mask; everything after it assumes that mask is right, which is why a bad threshold is never a visible error but always a wrong number at the end. ## Otsu versus the alternatives, briefly A hard-coded value is fine only in a fully controlled enclosure with fixed exposure. Otsu is the right default when lighting varies globally but is uniform across the frame. A locally adaptive threshold is the answer when the illumination varies *within* the frame. And if the two classes genuinely overlap in intensity - a grey part on a grey belt - no intensity threshold will save you; the fix is colour, lighting or a different feature entirely.

  • Why do people pass 0 as the thresh argument when using THRESH_OTSU?
    Purely as a readability convention. With the Otsu flag set the value is ignored entirely, so any number would behave identically; writing 0 tells the next reader that the argument is dead and the threshold is computed. The value actually used comes back as the first return value.
  • A scan is lit brightly on the left and dimly on the right, and Otsu blacks out half the page. What do you do?
    Switch to a locally computed threshold - cv2.adaptiveThreshold - which derives a value per pixel from its own neighbourhood, so a global illumination gradient cancels out. Alternatively estimate and divide out the illumination field first, for example with a morphological tophat, and then apply Otsu to the flattened image.
  • Why does blurring before Otsu often improve the result?
    Otsu decides from the histogram alone. Pixel noise widens and merges the foreground and background humps, blurring the valley the method is looking for, so the chosen cut moves around frame to frame. A light smoothing pass tightens both modes, sharpens the separation, and makes the selected threshold far more stable.
  • What is cv2.THRESH_TRIANGLE for?
    It is the other automatic threshold OpenCV ships, and it suits histograms with one dominant peak and a long tail rather than two comparable modes - a mostly-empty background with a small bright object. Where Otsu would cut the dominant mode in half, the triangle method places the cut out along the tail.

saying these in an interview costs you the question

  • Thinks the thresh argument is still applied
  • Ignores the returned computed threshold value
  • Claims Otsu adapts to local illumination
  • Believes it works on a color image directly
  • Trusts it on a flat unimodal histogram

context