What do you gain by using a native video input API instead of extracting frames yourself?
answer
- one path hands over the whole file
- the other hands over a list of images
- something is lost when frames arrive alone
- time and sound come free on one path
- control over selection comes free on the other
basics
~20 sA native video API takes the file plus a frame-rate hint and handles decoding, sampling, timestamps and the audio track for you, so the model can answer with times. Hand-extracted frames sent as an image list give you full control but arrive as anonymous stills.
solid answer
~50 sSending a video natively means uploading or referencing the file and letting the provider decode it, sample frames at a default or hinted rate, keep each frame's position in time, and tokenize any audio track into the same context. The practical payoff is timestamped answers: you can ask what happened and get a reply anchored at a point in the recording, and audio and picture are reasoned about together. Hand-extracting frames and sending them as an ordered list of images gives you control the API does not expose: motion-gated or keyframe sampling, cropping, per-frame preprocessing, and reuse of a cached prompt prefix. The cost is that you now own decoding, and the frames arrive as a bare sequence with no inherent timing, so you must state the timestamps in the prompt or burn them into the images. Audio also becomes a separate pipeline. Most teams start native and move to hand-extracted frames only when the sampling policy has to be smarter than a fixed rate.
go deeper
Be able to say that you can either upload the video and let the service sample it, or extract frames yourself and send them as images, and that the native path keeps timing and audio for you.
Explain the concrete tradeoff: control over sampling and preprocessing versus timestamps, audio in the same context, and no decoding pipeline, and describe how to restore timestamps on the DIY path.
Show that you can name the specific constraint that forces a migration off the native path, and that you have a plan for merging separately handled audio back in by timestamp.
Own the build-versus-inherit call: whether a frame pipeline is worth maintaining across teams, how provider default changes would affect your cost model, and what abstraction keeps both paths swappable.
## Two ways to put moving pictures in front of a model The first is native video input: you hand the provider the media, by upload, by inline bytes, or by a public URL the provider supports, optionally with a frame-rate hint and a resolution setting, and the service does the rest. Current native video APIs sample around one frame per second by default, bill roughly a couple of hundred tokens per frame, and tokenize an accompanying audio track at a per-second rate, all inside one context. The second is do-it-yourself: decode the file locally, choose frames by whatever policy you like, and send them as an ordered list of images in a single request, exactly as you would send a set of photographs. The model does the same thing in both cases. The difference is entirely about who owns the pipeline. ## What native input actually gives you **Time survives.** This is the big one. When the provider does the sampling, each frame carries its position in the recording, so the model can say that something happened near a given point and can be asked about a specific interval. A bare image list has no such anchor; the model knows the order and nothing else. **Audio and picture share a context.** A native request can include the sound, so questions that depend on both, an alarm sounding while a door opens, a crunch with no visible impact, are answerable in one pass. With hand-extracted frames the audio is a separate problem you must solve and then splice back together yourself. **No decoding pipeline to own.** Container formats, variable frame rates, rotation metadata, dropped frames and codec quirks are a genuine source of bugs, and they become the provider's problem. **Less code between the file and the answer.** For a prototype or a low-volume workload this alone usually decides it. ## What you give up **Sampling policy.** A frame-rate hint is a blunt instrument. If your footage needs dense frames around motion and almost nothing during idle stretches, or a fixed budget of the most visually distinct frames, you need to select the frames yourself. **Per-frame preprocessing.** Cropping to a region of interest, masking a privacy area, upscaling a small detail, or overlaying a timestamp are all things you can only do to frames you control. **Predictability of cost.** With a rate hint you are trusting the provider's defaults and limits, including caps on how many videos a single request may carry. With an explicit frame list you know exactly how many images you are sending. **Prompt caching.** Hand-built requests let you keep a stable prefix and vary the tail, which matters when you re-query the same footage many times. ## Making the DIY path work If you extract frames yourself, restore what you lost. State the timing explicitly, either as text alongside each image ("frame at 00:12:04") or by burning a timestamp overlay into the frame itself. Keep the ordering unambiguous. Decide what happens to the audio: transcribe it separately if it is speech, or treat non-speech sound as its own signal, and merge the results by timestamp rather than hoping the model interleaves them. ## A reasonable default Start native. It is fewer moving parts, it preserves time and audio for free, and for clips of minutes rather than hours the default rate is usually fine. Move to hand-extracted frames when one of three things is true: the footage is long enough that a fixed rate is unaffordable, the events are rare enough that a fixed rate misses them, or you need preprocessing the API cannot express. That is a migration driven by a specific constraint, not a matter of taste, and being able to name the constraint is what an interviewer is listening for.
- If you send hand-extracted frames, how do you let the model answer with timestamps?Carry the timing explicitly. Either interleave a short text label before each image giving its offset, or render a timestamp overlay into the frame itself so it is visible in the pixels. Both work; the overlay is more robust because it cannot be reordered or dropped by later context editing. Without one of them the model can only reason about order, and it will guess at durations.
- Why might you still prefer hand-extracted frames for a workload you re-query many times?Because you control the request layout. A stable set of frames plus a stable instruction block can sit in a cached prompt prefix, so repeated questions about the same footage pay for the frames once rather than on every call. You also get exact, predictable frame counts, which makes the cost per query a number you can quote rather than a provider default you inherit.
saying these in an interview costs you the question
- Thinks the model watches the video continuously when sent natively
- Sends bare frames and still expects timestamped answers
- Forgets the audio track disappears when you extract frames yourself
- Assumes a frame-rate hint gives motion-aware sampling
- Treats decoding and rotation metadata as trivial to handle