How does frame-sampling rate drive the token cost of sending video to a multimodal model?
answer
- cost is linear in something you choose
- frames, not seconds, are billed
- rate times duration times tokens per frame
- 1 fps over hours is millions of tokens
- shortest event sets the floor on rate
basics
~20 sEach sampled frame becomes a few hundred image tokens, so cost scales with duration times frames per second. At roughly 258 tokens a frame, one frame per second across an 8-hour shift is about 7.4 million tokens, far past any context window. Rate and per-frame resolution are the levers.
solid answer
~50 sA model never watches video; it receives sampled still frames that are tokenized like images, so the bill is duration x frames-per-second x tokens-per-frame. Native video APIs commonly default near 1 fps at roughly 258 tokens per frame, with a low-resolution mode around 66, plus about 32 tokens per second if the audio track is included. For 8 hours of dock-door CCTV (28,800 seconds), 1 fps at full resolution is about 7.4M image tokens, several times a 1M-token window. One frame per five seconds gives about 1.5M; combining that with low resolution gives about 380k, which fits. The tradeoff is real: at one frame per five seconds a three-second forklift near-miss can fall entirely between samples, and low resolution erases small or distant detail. Size the rate against the shortest event you must not miss, then check the budget.
code
python · 9 linesdef video_tokens(hours, fps, tokens_per_frame, audio_tokens_per_sec=0):
seconds = hours * 3600
frames = seconds * fps
return int(frames * tokens_per_frame + seconds * audio_tokens_per_sec)
print(video_tokens(8, 1.0, 258)) # 7430400
print(video_tokens(8, 0.2, 258)) # 1486080
print(video_tokens(8, 0.2, 66)) # 380160
print(video_tokens(8, 0.2, 66, 32)) # 1301760go deeper
Know that video is turned into sampled still frames and each frame is billed like an image. Be able to say that more frames per second means proportionally more tokens.
Be ready to write the formula on the whiteboard and run real numbers through it, including a low-resolution option and the audio track, and to explain what a lower rate costs you in missed short events.
Show that you set the sampling rate from a detection requirement and then engineer the budget around it, with clipping, tiering and resolution choices, rather than picking a round number and hoping.
Own the economics across a fleet: what a per-camera per-day cost implies at scale, where a cheap coarse pass plus selective re-analysis beats uniform treatment, and when the whole workload should be a classical detector with the model reserved for adjudication.
## Video is a token problem before it is a vision problem Models do not consume video as video. The pipeline, whether the provider runs it server-side or you run it yourself, decodes the file into a sequence of still frames, and each frame is tokenized exactly the way a photograph would be. An audio track, if included, is tokenized separately. Everything about cost, context pressure and latency follows from that. The governing arithmetic is simple and worth being able to do out loud: image tokens = duration_seconds x frames_per_second x tokens_per_frame ## Working the warehouse example Take a safety review of dock-door CCTV: 8 hours per door, so 28,800 seconds. Using representative native-video figures of about 258 tokens per frame at default resolution and about 66 in a low-resolution mode: - 1 fps, default resolution: 28,800 x 258 is roughly 7.43M tokens. That is several multiples of a 1M-token context window, so the request cannot even be made, let alone paid for. - 0.2 fps (one frame every five seconds), default resolution: 5,760 x 258 is roughly 1.49M tokens. Still over. - 0.2 fps, low resolution: 5,760 x 66 is roughly 380k tokens. This fits, with room for the prompt. - Audio, if you leave the track in at roughly 32 tokens per second, adds about 922k tokens on its own for the same 8 hours. On silent CCTV that is pure waste, which is why stripping or disabling audio is often the single biggest saving. The useful habit is doing this arithmetic before writing any code. It usually kills the naive design in one line. ## The three levers, and what each costs you **Sampling rate.** The dominant term. Halving it halves the bill. It also sets an event-duration floor: anything shorter than the sampling interval may fall entirely between two frames. A three-second near-miss is caught reliably at 1 fps, unreliably at 0.2 fps, and essentially never at 0.05 fps. Choose the rate from the shortest event you must detect, not from the budget. **Per-frame resolution.** Fewer tokens per frame means a coarser image. Coarse frames are fine for gross motion (is anyone in the aisle, is the door open) and useless for small or distant detail (a label, a hand position, a person at the far end of a bay). Lowering resolution is the cheapest lever when the question is about presence and movement rather than fine detail. **Segmentation.** Instead of one giant request, cut the footage into windows, analyse each independently, and combine the results. This is what makes hours of footage tractable at a rate that would otherwise not fit, at the price of losing cross-window context. ## Two common ways to get this wrong The first is confusing file size with token cost. Re-encoding a video at a lower bitrate makes the upload smaller and changes the token bill not at all, because tokens are a function of how many frames get sampled and at what pixel budget, not of how well the codec compressed them. The second is treating the context window as the only constraint. Even when a long request fits, you are paying for every frame on every call, and latency grows with the prompt. A design that fits but costs several dollars per review, repeated across hundreds of cameras and days, is not a viable design. ## What a good answer sounds like State the formula, put real numbers through it, name the rate and resolution you would choose, and then say what that choice gives up. The interviewer is testing whether you reason about the token bill as a first-class design constraint rather than discovering it after the first invoice, and whether you can tie a sampling decision back to the detection requirement it has to satisfy.
- You have 40 cameras and a fixed monthly budget. How do you decide which footage gets the expensive treatment?Tier it. Run every camera at a cheap rate and low resolution as a coarse detector, and re-run only the windows it flags at a high rate and full resolution. The cheap pass is a filter, not the answer, so it should be tuned for recall rather than precision: a false positive costs one expensive re-run, a false negative loses the incident entirely.
- Does re-encoding the video to a lower bitrate before upload reduce the token cost?No. Token cost depends on how many frames are sampled and the pixel budget each frame is given, not on the file size. Heavy re-encoding only shrinks the upload and can damage the very detail you need. The things that actually move the bill are sampling rate, per-frame resolution, clipping to relevant windows, and dropping the audio track when it carries nothing.
- Where does the audio track fit into this budget?It is billed per second of media rather than per frame, at roughly 32 tokens per second in current native video APIs, so an hour of audio is on the order of 115k tokens regardless of the frame rate. For silent or irrelevant-audio footage that is pure waste, so strip it. For footage where sound carries the signal, it is often cheaper than raising the frame rate to catch the same event visually.
saying these in an interview costs you the question
- Assumes the model streams video rather than sampling frames
- Thinks compressing the file lowers the token bill
- Picks a frame rate from the budget, ignoring event duration
- Leaves the audio track on silent CCTV in the request
- Treats fitting the context window as proof the design is affordable