How do OpenCV's CSRT and KCF trackers differ, and how do you drive one per frame?
answer
- create, init, update
- one tracker, one object
- accuracy against speed
- update returns a flag you must check
- no re-detection after loss
basics
~20 sBoth are single-object trackers seeded once with a box: create with cv2.TrackerCSRT.create() or cv2.TrackerKCF.create(), call init(frame, bbox), then update(frame) each frame for (success, bbox). CSRT is more accurate and handles scale change better; KCF is considerably faster.
solid answer
~50 sOpenCV's tracker classes share one interface. You construct — `cv2.TrackerCSRT.create()` or `cv2.TrackerKCF.create()`, also spelled `cv2.TrackerCSRT_create()` — then `tracker.init(frame, bbox)` with `bbox` as `(x, y, w, h)`, then call `ok, bbox = tracker.update(frame)` on every subsequent frame. `update` returns a success flag and the new box; when the flag is False the tracker has lost the target and there is no recovery from within the tracker. KCF uses kernelized correlation filters and is fast enough to run several instances in a real-time loop, but it is weak on scale change and cannot recover after full occlusion. CSRT (channel and spatial reliability) is noticeably more accurate, copes better with deformation and scale, and costs several times more per frame. Neither knows what it is tracking: they follow appearance from the seed box, so both drift, and a shipped pipeline pairs them with a detector that re-seeds periodically and on failure.
code
python · 20 linesimport cv2
cap = cv2.VideoCapture("input.mp4")
ok, frame = cap.read()
bbox = (100, 80, 60, 120) # x, y, w, h from a detector
tracker = cv2.TrackerCSRT.create() # same as cv2.TrackerCSRT_create()
tracker.init(frame, bbox)
while True:
ok, frame = cap.read()
if not ok:
break
found, bbox = tracker.update(frame)
if not found:
break # re-detect and construct a new tracker
x, y, w, h = (int(v) for v in bbox)
cv2.rectangle(frame, (x, y), (x + w, y + h), (0, 255, 0), 2)
cap.release()go deeper
Memorize the three-call shape — create, init with an (x, y, w, h) box, update per frame — and that update returns a success flag alongside the new box.
Contrast CSRT and KCF concretely on accuracy, scale handling and cost, explain that neither re-detects after loss, and know that meanShift and CamShift consume a back-projection image rather than a frame.
Demonstrate that you plan around drift: external validation against detections, dropping trackers that fail, re-seeding, and the fact that a True from update() is not evidence the box is correct.
Own the architecture: whether identity should come from per-object appearance trackers at all versus association across detections, what per-camera compute that implies, and how the choice degrades as object count and occlusion rate rise.
## The shared interface Every tracker in OpenCV implements the same three-call contract: 1. Construct: `cv2.TrackerCSRT.create()`, `cv2.TrackerKCF.create()`, `cv2.TrackerMIL.create()`. The flat aliases (`cv2.TrackerCSRT_create()`) are the same function. 2. `tracker.init(frame, bbox)` — seed it with the frame and a bounding box given as `(x, y, w, h)` in pixels, top-left origin. `cv2.selectROI` is the interactive way to get one for a demo; in a real system it comes from a detector. 3. `ok, bbox = tracker.update(frame)` — once per subsequent frame. It returns whether the target was found and the new box. One tracker instance follows one object. Multiple objects means multiple instances that you own and index yourself. A tracker cannot be re-pointed at a new target — to change targets you construct a fresh one and `init` it again. ## KCF KCF is built on a kernelized correlation filter: it learns a filter that responds strongly on the target's appearance and evaluates it densely over a search region using the frequency domain, which is why it is fast. Its weaknesses are structural. The filter is learned for a fixed template size, so scale change — an object walking towards the camera — degrades it. Full occlusion is fatal: the filter updates on whatever is in the box, so an occluder is learned as the target, and there is no re-detection to recover. Fast motion that carries the object outside the search region loses it too. ## CSRT CSRT extends discriminative correlation filtering with channel reliability weighting and a spatial reliability map, which lets it use a non-rectangular support region within the box and weight feature channels by how informative they are. The practical result is better handling of deformation, partial occlusion, rotation and scale, at several times KCF's cost per frame. On typical CPU hardware CSRT is the choice when accuracy matters and you have one or a few objects; KCF when you have many objects or a tight budget. ## Other options in the box `cv2.TrackerMIL` is the third classical tracker in the main module — robust to partial occlusion but slower and less precise on the box. `cv2.TrackerGOTURN` needs a downloaded model. Recent OpenCV also ships lightweight neural trackers (`cv2.TrackerNano`, `cv2.TrackerVit`, `cv2.TrackerDaSiamRPN`) which need their model files supplied. Older names — MOSSE, Boosting, TLD, MedianFlow — moved to the `cv2.legacy` namespace in the contrib build. ## meanShift and CamShift `cv2.meanShift` and `cv2.CamShift` are the older, lighter tracking pair and they are the ones people get wrong at the interface. Neither takes a frame. Their first argument is a **probability image** — typically a hue histogram of the target computed with `cv2.calcHist` and then back-projected onto each new frame with `cv2.calcBackProject`. `cv2.meanShift(probImage, window, criteria)` returns `(iterations, window)` and shifts a fixed-size window to the local density peak. `cv2.CamShift` (continuously adaptive mean shift) additionally adapts the window size and returns a rotated rectangle, which is why it survives an object approaching the camera. Both are cheap and both fail when a background region shares the target's colour distribution — they track colour, nothing more. ## Why every tracker drifts All of these follow appearance from a seed. Every update either adapts the model — and slowly absorbs background, drifting off the target — or does not adapt, and fails as soon as the object changes appearance. There is no configuration that escapes the tradeoff. The consequences you must plan for: - **No identity.** A tracker will happily follow the wrong object after a crossing; it has no notion of who it is following. - **No recovery.** A False from `update` is terminal for that instance. A returned True is not a guarantee either — a drifting tracker reports success while sitting on background. - **Unbounded error growth.** Without external correction the box degrades monotonically over hundreds of frames. The production shape is therefore always tracking-by-detection: a detector supplies boxes at some interval, trackers carry them between detections, associations are checked, failed trackers are dropped and new detections seed new ones. Validating tracker output against the next detection — an IoU check, say — catches drift the tracker itself never reports. ## Practical details that bite `update` returns box coordinates as floats; cast before slicing or drawing. Boxes that run outside the frame after a resize or crop raise or misbehave, so clamp. And seeding matters more than the tracker choice: a loose box that includes a lot of background gives either tracker a model contaminated from frame one, and no amount of algorithm quality recovers from a bad seed.
- tracker.update() keeps returning True but the box has slid onto the background. How do you catch that?The tracker cannot tell you — success only means its filter found a peak. You need external validation: periodically run a detector and check IoU or centre distance between the tracked box and the nearest detection, dropping or re-seeding the tracker when they disagree. Cheaper heuristics help too, such as rejecting implausible box velocity or size change between frames.
- You need to track twenty objects at once. What changes?Nothing about the interface — you keep a list of tracker instances and call update on each — but the cost is linear, so twenty CSRT instances will not fit a 30 fps budget on a modest CPU. That is where you switch to KCF or a lighter tracker, drop the tracking resolution, or move to a detect-and-associate design where identity comes from matching detections across frames rather than from a per-object appearance filter.
- What does cv2.CamShift give you that cv2.meanShift does not?An adapting window. meanShift shifts a fixed-size window to the local peak of the probability image, so an object growing as it approaches the camera outgrows the window. CamShift resizes and orients the window each iteration and returns a rotated rectangle, which handles scale and in-plane rotation. Both still track a colour distribution and both fail on background that shares the target's colours.
saying these in an interview costs you the question
- Expecting a tracker to re-find the object after occlusion
- Ignoring the boolean returned by update()
- Believing a True result means the box is still on the target
- Passing a raw frame to cv2.meanShift instead of a back-projection
- Assuming one tracker instance can follow several objects