How does cv2.createBackgroundSubtractorMOG2 label pixels, and what does detectShadows change?
answer
- mask is 8-bit, not boolean
- three labels, not two
- shadows are dark but same colour
- learningRate freezes or rebuilds the model
- getBackgroundImage shows what it learned
basics
~20 sMOG2 models each pixel's recent history as a mixture of Gaussians and returns an 8-bit mask: 0 for background, 255 for foreground. With detectShadows=True (the default) it adds a third label, 127, for pixels it judges to be shadow — so a naive nonzero test counts shadows as objects.
solid answer
~40 s`cv2.createBackgroundSubtractorMOG2(history, varThreshold, detectShadows)` keeps a per-pixel mixture of Gaussians describing what that pixel usually looks like. Each call to `apply(frame)` classifies pixels against that model and then updates it, returning a `CV_8U` mask. The values are three, not two: 0 background, 255 foreground, and 127 shadow when `detectShadows=True`, which is the default. That third value is the trap — `mask > 0` or `cv2.findContours` on the raw mask treats every shadow as a detection, so you either threshold at something like 200 to keep only hard foreground, or pass `detectShadows=False` (which is also measurably faster). `apply()` also takes a `learningRate`: -1 lets the subtractor pick from `history`, 0 freezes the model, and 1 rebuilds it from the current frame. `cv2.createBackgroundSubtractorKNN` is the non-parametric sibling with the same mask convention.
code
python · 16 linesimport cv2
cap = cv2.VideoCapture("input.mp4")
sub = cv2.createBackgroundSubtractorMOG2(
history=500, varThreshold=16, detectShadows=True
)
while True:
ok, frame = cap.read()
if not ok:
break
mask = sub.apply(frame) # 0 bg, 127 shadow, 255 fg
hard = cv2.threshold(mask, 200, 255, cv2.THRESH_BINARY)[1]
bg = sub.getBackgroundImage() # what the model believes
cap.release()go deeper
Remember the three mask values — 0, 127, 255 — and that 127 only appears when shadow detection is on. Know that apply() both classifies and updates on every call.
Explain the mixture-of-Gaussians model per pixel, what varThreshold and history control, and how learningRate values of -1, 0 and 1 differ. Name the KNN alternative and its identical mask convention.
Show the diagnosis habits: inspecting getBackgroundImage() for ghosts, recognising absorbed stationary objects, and knowing that a mask is a pixel classification that still needs cleanup and grouping before anything downstream calls it an object.
Own the choice of technique. Argue when classical subtraction is the right tool at all — fixed camera, tight compute, no labelled data — versus a learned detector, and price the ongoing cost of tuning thresholds per camera across a fleet.
## The idea Background subtraction turns a video into a per-pixel foreground mask by learning what each pixel normally looks like and flagging deviations. OpenCV ships two of these in the main `video` module: `cv2.createBackgroundSubtractorMOG2()` and `cv2.createBackgroundSubtractorKNN()`. Both are adaptive — the model updates as the video plays, so gradual lighting change, a parked car that becomes scenery, and slow camera auto-exposure drift are absorbed over time rather than reported forever as motion. ## What MOG2 models MOG2 keeps, for every pixel, a small mixture of Gaussian distributions over that pixel's observed values, together with weights. The number of components adapts per pixel, which is the point of the "2": a pixel showing a swaying leaf against sky needs two or three modes, while a pixel on a blank wall needs one. A new observation is background if it falls within `varThreshold` (in Mahalanobis terms) of a well-supported component; otherwise it is foreground. The constructor arguments you actually tune are `history` (roughly how many frames of context the model holds, default 500), `varThreshold` (default 16 — raise it to suppress noise, lower it to catch subtle objects), and `detectShadows` (default True). ## The three mask values The return of `apply()` is a single-channel 8-bit image, and this is where most bugs live. Background is 0 and foreground is 255, as expected. But with shadow detection on, pixels classified as shadow get the shadow value, which defaults to 127 and is readable and settable via `getShadowValue()` / `setShadowValue()`. A shadow is detected by the classic chromaticity heuristic: the pixel is darker than its background model but keeps roughly the same colour ratios, controlled by `setShadowThreshold()`. The consequence is concrete. If you write `mask > 0`, or feed the raw mask to contour finding, or `cv2.bitwise_and` the frame with it, every cast shadow becomes part of the object — a walking person acquires a long limb across the floor, and blob areas used for counting or classification are wrong. The two fixes are `cv2.threshold(mask, 200, 255, cv2.THRESH_BINARY)` to keep only hard foreground, or `detectShadows=False` at construction if you never want the label. Turning it off also costs less per frame, which matters in a real-time loop. ## learningRate `apply(frame, learningRate=...)` controls how fast the model absorbs the current frame. The default -1 means "choose automatically from `history`". A value of 0 freezes the model — nothing about this frame is learned, which is what you want if you have to hold a clean background across a period where a large object is stationary in view. A value of 1 reinitializes the model from this frame alone. Small values like 0.001 adapt slowly, larger ones quickly. Getting this wrong is the second classic bug: with too high a learning rate, an object that stops moving is absorbed into the background within a second and vanishes from the mask; with too low a rate, a genuine lighting change leaves the whole scene lit up as foreground for a long time. ## MOG2 versus KNN `createBackgroundSubtractorKNN(history, dist2Threshold, detectShadows)` uses a non-parametric K-nearest-neighbours decision over recent samples instead of fitting Gaussians. It shares the mask convention, including the 127 shadow label, and the same `apply()`/`getBackgroundImage()` interface. In practice KNN often handles small, sparse foreground better and copes with multi-modal backgrounds without needing enough samples to fit a distribution, while MOG2 tends to be faster on busy scenes. There is no universal winner — the honest interview answer is that you try both on your footage and compare mask quality against your downstream stage. Additional variants (MOG, GMG, CNT, LSBP) live in the contrib `cv2.bgsegm` module rather than in the main build. ## What background subtraction cannot do All of this assumes a static camera. On a pan-tilt or handheld camera every pixel changes and the mask is garbage; you need motion compensation or a different technique entirely. It also assumes the object differs from the background — a person in a grey coat against a grey wall produces almost nothing. And the mask is a per-pixel classification with no notion of objects: it is noisy at the edges, splits one person into several blobs when part of them matches the background, and merges two people who touch. Turning the mask into object hypotheses is a separate stage, and in a shipped pipeline the subtractor is nearly always followed by cleanup and then a grouping step before anything downstream sees a "detection". ## Inspecting the model `getBackgroundImage()` returns the subtractor's current idea of the background as an image. It is the single best debugging tool here: if the returned image contains a ghost of a person, your learning rate absorbed a stationary object; if it is a blur of the whole scene, the model never converged.
- A person stops moving and disappears from the MOG2 mask after a few seconds. Why, and what would you do?The subtractor is adaptive, so a stationary object is gradually absorbed into the background model at whatever rate `history`/`learningRate` imply. Either slow adaptation down with a small explicit `learningRate`, or freeze the model with `learningRate=0` during the period you need the object held, or accept it and keep object identity in a tracker rather than re-deriving it from the mask every frame.
- How would you choose between MOG2 and KNN for a given camera?Empirically. Both share the same interface and the same mask convention, so swapping the factory call is a one-line change. Run both over representative footage — including the hard lighting and the busiest hour — and compare mask quality and per-frame cost against the stage that consumes the mask. KNN often wins on sparse foreground and multi-modal backgrounds; MOG2 often wins on throughput.
- What happens to background subtraction on a camera that pans?It breaks. The technique assumes a fixed viewpoint so that a pixel's history describes one piece of the scene. Under pan or handheld motion every pixel changes and nearly everything is flagged foreground. You either stabilize/register frames to a reference first, or switch to a technique that models motion rather than appearance.
saying these in an interview costs you the question
- Treating the mask as binary and testing mask > 0
- Believing detectShadows removes shadows rather than labelling them
- Assuming the background model is fixed after the first frames
- Using background subtraction on a moving camera
- Calling raw mask blobs 'detections' with no cleanup or grouping