skip to content

Video Processing

Working frame by frame: the capture and write loop, background subtraction to isolate motion, optical flow to estimate it, and the built-in trackers. Interviewers ask about the accuracy-versus-latency trade-off inside a real-time loop.

on this pageshow

questions

6

When do you choose Lucas-Kanade over Farneback optical flow in OpenCV?

level: middleimportance: must knowfreq 55%

answer

  1. sparse points versus a full field
  2. status array is not optional
  3. both want grayscale input
  4. brightness constancy is the shared assumption
  5. pyramids buy you larger motion

basics

~20 s

cv2.calcOpticalFlowPyrLK is sparse: you give it specific points and it returns where each moved, plus a status array. cv2.calcOpticalFlowFarneback is dense: it returns a displacement vector for every pixel. Choose sparse when you track a few features cheaply, dense when you need motion everywhere.

solid answer

~50 s

They answer different questions. `cv2.calcOpticalFlowPyrLK(prevGray, nextGray, prevPts, None, ...)` is the sparse pyramidal Lucas-Kanade tracker: you hand it an Nx1x2 float32 array of points — usually from a corner detector — and it returns `(nextPts, status, err)`, where `status` marks which points it managed to follow. Cost scales with the number of points, so a few hundred features tracked per frame is nearly free, which is why it underpins visual odometry and feature-based stabilization. `cv2.calcOpticalFlowFarneback(prevGray, nextGray, None, 0.5, 3, 15, 3, 5, 1.2, 0)` instead returns an HxWx2 float32 field giving a displacement for every pixel — useful for segmenting motion regions, warping frames, or building a motion energy image, but orders of magnitude more expensive. Both take single-channel images, both assume brightness constancy and small displacement between frames, and both use an image pyramid so that motion larger than the window is still recoverable at a coarser level.

code

python · 21 lines
python
import cv2

cap = cv2.VideoCapture("input.mp4")
ok, first = cap.read()
prev = cv2.cvtColor(first, cv2.COLOR_BGR2GRAY)
pts = cv2.goodFeaturesToTrack(prev, maxCorners=200, qualityLevel=0.01, minDistance=10)

while True:
    ok, frame = cap.read()
    if not ok:
        break
    gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
    nxt, status, err = cv2.calcOpticalFlowPyrLK(
        prev, gray, pts, None, winSize=(21, 21), maxLevel=3
    )
    pts = nxt[status.flatten() == 1].reshape(-1, 1, 2)
    if len(pts) < 50:
        pts = cv2.goodFeaturesToTrack(gray, 200, 0.01, 10)
    prev = gray

cap.release()

go deeper

for a junior

Know that one is sparse and one is dense, what each returns, and that both need grayscale input. Be able to name calcOpticalFlowPyrLK and calcOpticalFlowFarneback without hesitating.

for a middle

Explain the return shapes precisely — (nextPts, status, err) versus an HxWx2 float32 field — why status must be filtered, and how the pyramid lets either method survive motion larger than its window.

for a senior

Bring the failure modes from real footage: exposure changes faking global motion, point attrition needing re-seeding, the aperture problem in textureless regions, and downscaling as the standard way to afford dense flow.

for a principal

Own the build-versus-buy call on motion estimation: when classical flow is enough, when a learned flow model or a tracker-by-detection design is worth the dependency, and what the compute budget per camera implies for the fleet's hardware.

## What optical flow gives you Optical flow is the apparent displacement of brightness patterns between two consecutive frames. OpenCV exposes it in two shapes, and knowing which shape you want is most of the interview answer. ## Sparse: calcOpticalFlowPyrLK The signature is `cv2.calcOpticalFlowPyrLK(prevImg, nextImg, prevPts, nextPts, winSize=(21,21), maxLevel=3, criteria=...)` and it returns a three-tuple `(nextPts, status, err)`. - `prevImg` / `nextImg` are consecutive **grayscale** frames (8-bit single channel is the normal case). - `prevPts` is an Nx1x2 array of float32 coordinates. You do not get these from LK itself — they come from a corner detector such as `cv2.goodFeaturesToTrack`, or from wherever your pipeline already has points of interest. - `nextPts` is the estimated new location of each point. - `status` is an Nx1 uint8 array: 1 where the point was found, 0 where tracking failed. **Filtering by status is mandatory.** Points that leave the frame, get occluded, or land in a textureless region come back with garbage coordinates and status 0, and a loop that ignores status accumulates nonsense over a few hundred frames. - `err` holds a per-point error measure you can threshold as a second quality filter. Cost is proportional to the number of points, not the image area. Tracking 300 corners at 1080p is cheap enough for real time on a laptop CPU. ## Dense: calcOpticalFlowFarneback `cv2.calcOpticalFlowFarneback(prev, next, flow, pyr_scale, levels, winsize, iterations, poly_n, poly_sigma, flags)` returns a single array of shape (H, W, 2) and dtype float32: `flow[y, x]` is the (dx, dy) displacement of that pixel. Typical arguments are `0.5, 3, 15, 3, 5, 1.2, 0`. The algorithm fits local polynomial expansions of the image neighbourhood and solves for the displacement that maps one to the other, over an image pyramid. To visualize or threshold it you normally convert to polar form: `mag, ang = cv2.cartToPolar(flow[..., 0], flow[..., 1])`, then map angle to hue and magnitude to value in an HSV image. Cost scales with pixel count, and at full resolution it is far too slow for a 30 fps loop on modest hardware — downscaling the input before computing flow and scaling the vectors back is the standard mitigation. `cv2.DISOpticalFlow` is the faster dense alternative in the main build when Farneback cannot keep up. ## The assumptions both share **Brightness constancy.** Flow assumes a physical point keeps the same intensity between frames. A cloud crossing the sun, an auto-exposure adjustment, a fluorescent flicker, or an IR illuminator switching on violate this and produce large spurious flow across the whole frame — a classic production surprise on outdoor cameras. **Small motion.** The local solution is only valid for displacements within the analysis window. Both functions mitigate this with a pyramid: motion that is 40 pixels at full resolution is 5 pixels three levels up, where the solver is happy, and the estimate is refined downwards. `maxLevel` in LK and `levels`/`pyr_scale` in Farneback control this. Fast motion still defeats a shallow pyramid, and the symptom is flow that snaps to zero or to a neighbouring texture. **Spatial coherence / the aperture problem.** Inside a local window, only the motion component perpendicular to an edge is observable — a featureless edge sliding along itself looks static. This is why LK wants corners, not edges, and why textureless regions in a dense field are unreliable no matter how confident the numbers look. ## Choosing Ask what the consumer of the flow needs. If it needs *the motion of specific things* — track these corners for stabilization, estimate how the camera moved, follow these points on an object — sparse LK is right, and the sparse output is directly consumable. If it needs *motion everywhere* — which regions of the frame are moving, warp this frame to align with the last, produce a motion field for a downstream model — dense flow is the only option, and you pay for it. A practical middle ground worth mentioning: run dense flow on a downscaled frame for region-level decisions, and sparse LK at full resolution for the handful of points where precision matters. And note the maintenance detail for the sparse path — tracked points die off over time, so a long-running loop must periodically re-seed with a fresh corner detection rather than tracking the original set forever. ## Practical shape and dtype notes Most debugging here is shape and dtype. LK wants float32 points with shape (N, 1, 2) and will complain or misbehave with a plain (N, 2) int array. Both functions want single-channel input; passing a BGR frame is the most common error. And after filtering by status you must reshape back to (-1, 1, 2) before the next call, or the following iteration fails.

  • How do you keep a Lucas-Kanade tracker alive over thousands of frames?
    Points attrit — they leave the frame, get occluded, or drift into textureless regions — so you filter by the returned status and error every frame and periodically re-seed. A common pattern is to re-run corner detection whenever the surviving point count falls below a threshold, masking out regions where points already exist so the new ones spread out. A forward-backward consistency check, tracking to the next frame and back, discards drifting points early.
  • An outdoor camera produces huge dense flow across the whole frame for one frame, with nothing actually moving. What happened?
    A global brightness change broke the brightness-constancy assumption — auto-exposure or auto-gain stepping, a cloud crossing the sun, or an IR cut filter switching. Flow interprets the intensity change as motion. Mitigations are locking camera exposure and gain where the hardware allows, normalizing frame intensity before computing flow, and rejecting frames whose global flow magnitude spikes coherently.
  • Farneback at 1080p cannot hit 30 fps. What are your options before giving up on dense flow?
    Downscale the input, compute flow on the smaller pair, and scale the vectors back — flow cost scales with pixel count and region-level decisions rarely need full resolution. Reduce `levels`, `winsize` and `iterations`. Compute flow on every Nth frame pair rather than every frame. Or switch to `cv2.DISOpticalFlow`, which is designed for the speed-accuracy point Farneback misses.

saying these in an interview costs you the question

  • Ignoring the status array and using every returned point
  • Passing BGR frames instead of grayscale
  • Expecting LK to find points for you rather than tracking given ones
  • Assuming dense flow at full resolution can run in a real-time loop
  • Believing flow works under changing exposure or lighting

context

open as a page

How does cv2.createBackgroundSubtractorMOG2 label pixels, and what does detectShadows change?

level: middleimportance: must knowfreq 58%

basics

~20 s

MOG2 models each pixel's recent history as a mixture of Gaussians and returns an 8-bit mask: 0 for background, 255 for foreground. With detectShadows=True (the default) it adds a third label, 127, for pixels it judges to be shadow — so a naive nonzero test counts shadows as objects.

open as a page

Why does cv2.VideoWriter produce an empty video file without ever raising an error?

level: middleimportance: must knowfreq 62%

basics

~20 s

VideoWriter.write() silently ignores any frame whose size or channel count differs from the frameSize and isColor passed to the constructor, and the writer never opens at all if the fourcc codec is unavailable. Check isOpened() and pass (width, height).

open as a page

How do OpenCV's CSRT and KCF trackers differ, and how do you drive one per frame?

level: middleimportance: should knowfreq 52%

basics

~20 s

Both are single-object trackers seeded once with a box: create with cv2.TrackerCSRT.create() or cv2.TrackerKCF.create(), call init(frame, bbox), then update(frame) each frame for (success, bbox). CSRT is more accurate and handles scale change better; KCF is considerably faster.

open as a page

An OpenCV loop on a live RTSP camera keeps showing frames seconds old — why, and how do you fix it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Processing is slower than the camera's frame rate, so decoded frames queue up behind the loop and VideoCapture.read() keeps handing back the oldest one. Latency grows without bound. Fix it by decoupling capture from processing — read in a thread that keeps only the newest frame, or drop frames with grab().

open as a page

How do you split a real-time video pipeline's frame budget between detection and tracking?

level: principalimportance: should knowfreq 36%

basics

~20 s

Run the expensive detector every N frames and carry boxes between detections with cheap trackers, sizing N so the average per-frame cost fits the frame interval. The cost of a larger N is tracker drift and delayed discovery of new objects, so N is measured, not guessed.

open as a page