skip to content

When does uniform frame sampling fail, and how would you sample video adaptively instead?

level: seniorimportance: should knowfreq 48%

answer

  1. information is not spread evenly
  2. budget follows change, not the clock
  3. a floor rate plus bursts
  4. include seconds before and after a trigger
  5. motion gates are blind to stillness

basics

~20 s

Uniform sampling spends the same budget on empty footage as on the moment that matters, so short events fall between samples while idle aisles are described dozens of times. Adaptive sampling gates frames on change: dense while something moves, sparse while nothing does.

solid answer

~50 s

Uniform sampling assumes information is spread evenly through the footage, and in surveillance or long-form recordings it is not. On a warehouse shift, hours of an empty aisle consume the same tokens per minute as the ten seconds when a forklift reverses into a pallet, and a two-second event still falls between samples if the interval is five seconds. Adaptive sampling reallocates that budget: compute a cheap per-frame change signal on the client, such as frame differencing, a motion or scene-change score, or an existing detector, and sample densely while the signal is high, sparsely while it is flat. The practical shape is a floor rate so you never go blind, a burst rate around triggers, and a pre-roll and post-roll window so the model sees the approach and the aftermath rather than only the impact. Then check what the gate silently drops, because a motion gate misses stationary events.

code

python · 13 lines
python
def adaptive_frames(motion, floor_every=25, threshold=0.3, roll=10):
    picked = set()
    for i, m in enumerate(motion):
        if i % floor_every == 0:
            picked.add(i)
        if m >= threshold:
            lo = max(0, i - roll)
            hi = min(len(motion), i + roll + 1)
            picked.update(range(lo, hi))
    return sorted(picked)

motion = [0.0] * 60 + [0.8] * 3 + [0.0] * 60
print(len(adaptive_frames(motion)), "of", len(motion), "frames kept")

go deeper

for a junior

Understand that a fixed interval between frames means anything shorter than that interval can be missed entirely, and that most surveillance footage is mostly uneventful.

for a middle

Be able to describe a concrete adaptive scheme: a cheap change signal, a floor rate, a burst rate around triggers, and a window before and after the trigger, and explain why each part is there.

for a senior

Demonstrate that you treat the sampler as an evaluated component with its own recall against labelled incidents, and that you can name its blind spots, especially stationary and slow-developing events.

for a principal

Own the tradeoff across a fleet: where the gate belongs (camera, edge, or ingest), what recall target justifies its cost, and when a purpose-built detector should own detection with the model reserved for explaining flagged clips.

## The assumption uniform sampling makes Sampling every N seconds implicitly assumes that any interval is as informative as any other. That is roughly true for a densely edited film and badly false for the footage most systems actually process: security cameras, dashcams, machine monitoring, operating-room recordings, long meetings. In those, information arrives in bursts separated by long stretches where nothing changes. The consequence is a budget spent in the wrong place. Over an eight-hour shift, a uniform sample produces hundreds of near-identical frames of an empty aisle, each fully billed and each contributing nothing, while the ten seconds that the whole review exists for get the same two or three frames as any other ten seconds. ## The second failure: events smaller than the interval A sampling interval sets a hard floor on what is observable. If frames arrive every five seconds, an event that begins and ends inside three seconds may be represented by zero frames, or by one ambiguous frame at its very edge. There is no prompt that recovers information that was never sent. This is why lowering the rate to fit a budget is not a soft quality tradeoff: past a threshold set by the shortest event you care about, detection does not degrade gracefully, it goes to zero. ## The shape of an adaptive sampler Adaptive sampling separates the cheap job of deciding what is interesting from the expensive job of understanding it. **A cheap signal.** Frame differencing, a motion-energy or scene-change score, an encoder's own keyframe and motion-vector data, or a small purpose-built detector that already runs on the stream. All of these are orders of magnitude cheaper per frame than a model call, so you can afford to compute them on every decoded frame. **A policy over that signal.** Three parameters carry most of the value. A floor rate, so quiet periods are still sampled thinly and you never go completely blind. A burst rate, applied while the signal exceeds a threshold, dense enough to resolve the event. And pre-roll plus post-roll: when a trigger fires, include some seconds before and after it, because causality lives in the approach and the aftermath, not in the instant of the impact. **A budget cap.** Without one, a windy day of moving shadows triggers continuously and you have rebuilt uniform sampling at the burst rate. Cap frames per window and degrade by raising the trigger threshold when the cap binds. ## Keyframe selection as the other flavour A related technique picks a fixed number of frames spread by content rather than by clock: cluster frames by visual similarity and keep one representative per cluster, or keep the frames where the scene changes most. This is the right tool when you must hit an exact frame budget, for example to fit a context window, and it produces a storyboard rather than a timeline. Its weakness is the mirror image of motion gating: it under-samples long, slow, important activity because that activity is visually homogeneous. ## What adaptive sampling silently loses Every gate has a blind spot, and naming yours is what separates a senior answer from a clever one. A motion gate misses events defined by absence or stillness: a person who stops moving, a machine that should be running and is not, an item that was there an hour ago and is gone now. It also misses slow change, such as a gradual spill, because the per-frame difference never exceeds the threshold even though the hour-scale difference is enormous. It also breaks timing. If the model receives ten frames from one dense burst and two from an hour of quiet, and the frames arrive without explicit timestamps, it will tend to read them as an evenly spaced sequence and badly misjudge durations. Adaptive sampling therefore makes explicit per-frame timing more important, not less. Finally, the gate becomes a component you must evaluate. Its recall against a labelled set of known incidents is a number you should be able to quote, because the model can only be as good as the frames the gate let through, and a gate failure is invisible downstream: the model answers confidently about footage it never saw. ## Judgment to demonstrate Good answers tie the burst rate to the event duration, name the pre-roll and post-roll, admit the stillness blind spot, and treat the sampler as an evaluated component with its own recall metric rather than as a preprocessing detail.

  • Your motion gate is triggering constantly because a bay door lets in changing sunlight. What do you do?
    Make the signal robust before touching the threshold: mask the region the light hits, difference on a luminance-normalised image, or require the change to persist across several frames rather than firing on one. Then cap frames per window so a bad day degrades cost rather than exploding it. Blindly raising the threshold suppresses the sunlight and the forklift together.
  • How would you know your adaptive sampler is not dropping real incidents?
    Evaluate it as its own component. Take a labelled set of known incidents with their timestamps, run the sampler over the raw footage, and measure what fraction of incident windows produced at least one frame, and how many. That recall number caps everything downstream, because the model cannot reason about footage it never received, and its confident answers give you no warning.
  • When is plain uniform sampling still the right choice?
    When the footage is densely informative throughout, such as an edited lecture, a screen recording of a workflow, or a match where the ball is always in play, and when the question is about overall content rather than about rare events. Uniform sampling is also the honest default when you have no reliable cheap signal, since a badly tuned gate is worse than no gate at all.

saying these in an interview costs you the question

  • Assumes lowering the frame rate degrades detection gradually
  • Uses a motion gate for events defined by stillness
  • Sends only the trigger frame, with no pre-roll or post-roll
  • Runs the gate with no cap, so noisy footage explodes the budget
  • Never measures what fraction of known incidents the sampler let through

context