How do you split a real-time video pipeline's frame budget between detection and tracking?
answer
- mean cost under the frame interval
- detect every N, track between
- N is bounded by drift, not by math
- async detections arrive stale
- measure IoU at correction
basics
~20 sRun the expensive detector every N frames and carry boxes between detections with cheap trackers, sizing N so the average per-frame cost fits the frame interval. The cost of a larger N is tracker drift and delayed discovery of new objects, so N is measured, not guessed.
solid answer
~60 sThe constraint is a hard one: mean per-frame cost must sit under 1/fps with headroom for jitter, or latency grows without bound. A detector at, say, 80 ms cannot run every frame at 30 fps, so the standard shape is detect every N frames and track in between — trackers such as KCF cost a few milliseconds each, so the amortized budget is `(detector + N x trackers) / N`. Choosing N is an accuracy decision, not an arithmetic one: a larger N means boxes drift further before correction, and a new object entering the scene goes unseen for up to N frames, which is the number that matters for anything safety- or event-driven. Measure both — IoU of tracked boxes against the next detection, and time-to-first-detection — on real footage, and pick the largest N that keeps them acceptable. The other lever is moving detection off the critical path onto a worker, which raises throughput but delivers detections that are already stale, so boxes need forward-propagation before they are used for association.
code
python · 35 linesimport cv2
cap = cv2.VideoCapture("input.mp4")
N = 10
trackers = []
frame_idx = 0
def detect(frame):
# returns a list of (x, y, w, h) boxes from whatever detector you run
return []
while True:
ok, frame = cap.read()
if not ok:
break
if frame_idx % N == 0:
trackers = []
for box in detect(frame):
t = cv2.TrackerKCF.create()
t.init(frame, box)
trackers.append(t)
else:
alive = []
for t in trackers:
found, box = t.update(frame)
if found:
alive.append(t)
trackers = alive
frame_idx += 1
cap.release()go deeper
Understand the basic shape: an expensive detector cannot run every frame, so cheap trackers carry boxes in between. Know that the average per-frame time must fit inside the frame interval.
Compute the amortized budget explicitly and explain what raising the detection interval costs — drift between corrections and a delay before new objects are seen at all.
Show measurement-driven tuning: IoU at correction versus frames-since-detection, time-to-first-detection, p99 frame time, and the staleness handling an asynchronous detector demands.
Frame it as the product tradeoff it is — recency versus completeness versus identity stability — defend a fixed or adaptive interval with telemetry, and connect the choice to fleet capacity: how many cameras one accelerator serves follows directly from it.
## The budget is arithmetic; the value of N is not A real-time loop at F frames per second has 1000/F milliseconds per frame. Everything — decode, preprocess, detect, track, draw, encode, publish — comes out of that. Exceed it on average and you are in the growing-backlog regime where displayed frames fall ever further behind reality. With detection every N frames the amortized cost is `(C_detect + N x C_track x K) / N` where K is the number of tracked objects, plus fixed per-frame costs. That gives you a minimum feasible N immediately. The interesting question is the maximum acceptable N, and that comes only from measurement on representative footage. ## What raising N actually costs Three distinct penalties, and they matter differently per product: 1. **Drift.** Tracked boxes degrade between corrections. The measurable version is IoU between each tracked box and the matching detection when detection next runs. Plot mean IoU against frames-since-detection; the curve tells you where quality falls off a cliff for your scene, and it is scene-dependent — a crowded platform drifts far faster than an empty corridor. 2. **Detection latency for new objects.** A new object is invisible until the next detection frame, so worst-case discovery delay is N/F seconds. For a counting application 300 ms is irrelevant; for anything that triggers an action it may be the whole requirement. 3. **Identity errors.** Longer gaps mean more crossings and occlusions resolved by the tracker alone, and association at the next detection has more ambiguity to resolve. This grows superlinearly with object density, which is why an N validated on quiet footage collapses at rush hour. ## Cheaper levers before raising N Before trading accuracy, spend the obvious wins. Detect at a lower input resolution — most detectors lose little on medium-sized objects. Restrict detection to a region of interest where objects can actually appear. Use hardware decode so decoding is not eating the budget. Batch or quantize the detector. Choose a cheaper tracker: several KCF instances cost what one CSRT costs, and for many objects that swap buys more than any scheduling change. ## Asynchronous detection The next structural move is taking detection off the critical path: the main loop tracks every frame and hands frames to a detection worker that returns results some frames later. Throughput rises to whatever the tracking loop can do, and detection runs as often as its hardware allows rather than on a fixed schedule. The cost is that detections arrive **stale** — they describe frame t while the loop is on frame t+k. Merging them naively snaps boxes backwards. You either propagate the detection forward using the tracked motion between t and t+k before associating, or you keep a short ring buffer of frames and states and reconcile against the state at t. Getting this wrong produces a characteristic visual stutter that is worth naming in an interview, because it is the sign of an async pipeline nobody validated. ## Adaptive scheduling A fixed N is a compromise across conditions that vary. More sophisticated pipelines make N dynamic: detect more often when tracker confidence or IoU-at-correction is degrading, when object count is changing, or when scene motion energy is high; back off when the scene is static. This is strictly better in principle and worse in practice unless you can observe it, because a system whose compute demand varies with scene content is harder to capacity-plan. Whichever you choose, the scheduling policy must be visible in metrics. ## What to measure in production The telemetry that makes this decision auditable rather than folkloric: - p50 and p99 per-frame wall time against the frame interval — the p99 is what causes visible hitches. - frames dropped per minute, and by which stage. - mean IoU at correction, bucketed by frames-since-detection. - time-to-first-detection for newly appearing objects. - tracker failure rate and re-seed rate. A pipeline with those numbers can defend its N. One without is tuned by whoever last looked at the demo video. ## Where the decision really lives Finally, this is a product tradeoff dressed as an engineering one. Ask what the consumer of the output needs: recency, completeness, or identity stability. A monitoring dashboard tolerates dropped frames and moderate drift. An event trigger cannot tolerate discovery delay. An analytics count over an hour tolerates latency but not identity swaps, because those corrupt the count directly. Different answers put N in different places, and they may put the whole pipeline in a different place — some workloads are better served by processing a recorded segment offline at full quality than by squeezing a detector into a live loop at all. Fleet economics settle the rest: N interacts directly with how many cameras one accelerator serves, so the frame budget is also a capacity-planning decision.
- How do you actually choose N rather than guessing it?Measure two curves on representative footage: mean IoU between tracked boxes and the next detection as a function of frames-since-detection, and worst-case time-to-first-detection for newly entering objects. The first tells you where box quality collapses, the second is fixed by N/F. Pick the largest N that keeps both inside the product's tolerance, and re-measure on the busiest footage, since drift and association errors grow with object density.
- What breaks when detection runs asynchronously on a worker thread?Detections describe an older frame than the one the loop is on. Associating them directly against current tracker state snaps boxes backwards and corrupts identity assignment. The fix is to forward-propagate the detection through the tracked motion for the elapsed frames before matching, or to keep a short history of tracker states and reconcile against the state at the detection's own timestamp.
- When is a dynamic detection interval worth the complexity over a fixed N?When scene conditions vary widely and you have the telemetry to observe the adaptation — for example detecting more often as IoU-at-correction degrades or object count changes. The cost is capacity planning: compute demand now depends on scene content, so peak load is harder to bound across a fleet. Without metrics exposing why the rate changed, an adaptive policy is untunable in production.
- Which output requirement should drive this design before any of the numbers?What the consumer needs from the stream: recency, completeness, or identity stability. A live dashboard tolerates drops and drift; an event trigger cannot tolerate discovery delay; an hour-long count tolerates latency but is corrupted by identity swaps. Naming which of the three is primary settles both N and whether the workload belongs in a live loop at all.
saying these in an interview costs you the question
- Choosing the detection interval from arithmetic alone, with no accuracy measurement
- Assuming a tracker's output is as good as a detection
- Ignoring that a new object is invisible until the next detection frame
- Merging asynchronous detections without accounting for their staleness
- Tuning on quiet footage and shipping the same N for peak load