How do you cut Gemini's token cost for a long video without dropping content?
answer
- compression is not a lever here
- clip before you send
- frame rate is a knob you own
- resolution trades detail for tokens
- audio alone is far cheaper
basics
~20 sSend less of the video rather than a smaller file: clip with video_metadata start_offset and end_offset, lower the frame sampling rate, and set media_resolution to MEDIA_RESOLUTION_LOW. Re-encoding to a smaller file changes bytes, not tokens.
solid answer
~50 sBecause Gemini bills video by duration and sampled frames — roughly 263 tokens per second at default resolution — the levers are all about how much of the video the model actually sees. First, **clip it**: the SDK's `types.VideoMetadata` carries `start_offset` and `end_offset`, so you send the relevant 90 seconds of a 40-minute recording rather than the whole file. Second, **sample fewer frames**: VideoMetadata's `fps` field lets you drop below the default one frame per second for near-static footage such as a screen recording or a fixed-camera lecture. Third, **lower media resolution**: `GenerateContentConfig(media_resolution=types.MediaResolution.MEDIA_RESOLUTION_LOW)` reduces the tokens spent per frame, trading fine detail like small on-screen text for a large saving. Beyond that, segment long media into chunks and carry structured summaries forward instead of raw footage — and if the information is really in the speech, send the audio alone at about an eighth of the video rate.
code
python · 25 linesfrom google import genai
from google.genai import types
client = genai.Client(api_key="YOUR_KEY")
video = client.files.upload(file="keynote.mp4")
part = types.Part(
file_data=types.FileData(
file_uri=video.uri,
mime_type=video.mime_type,
),
video_metadata=types.VideoMetadata(
start_offset="600s",
end_offset="690s",
),
)
reply = client.models.generate_content(
model="gemini-2.5-flash",
contents=[part, "What is demoed in this segment?"],
config=types.GenerateContentConfig(
media_resolution=types.MediaResolution.MEDIA_RESOLUTION_LOW,
),
)
print(reply.text)go deeper
Know that a long video is expensive because cost follows its duration, and that sending only the relevant portion is the simplest way to spend less.
Explain the actual knobs — start and end offsets on the video Part, a custom frame rate, and the media_resolution setting — and why re-encoding the file changes none of them.
Show a staged pipeline: a cheap low-resolution or audio-only pass to locate moments, a precise high-resolution pass on clipped windows, and count_tokens verification at each stage.
Own the unit economics: which modality the product genuinely needs, what quality floor the use case demands, and how segmentation and summarisation keep cost and latency predictable at volume.
## Why the obvious levers do nothing Engineers reach first for compression: lower bitrate, smaller resolution, a more efficient codec. All of that reduces upload time and storage, and none of it reduces token cost. Gemini normalises frames before tokenising, and it counts *seconds*, so a 4K master and a 480p proxy of the same clip cost the same. The only things that move the number are how long the media is, how often it is sampled, and how much detail each sample is rendered at. ## Lever 1 — clip the time range The biggest saving is almost always "don't send the parts you don't need". The Part carrying a video accepts video metadata with a start and end offset, expressed as duration strings such as `"600s"`. If your retrieval layer already knows the interesting moment is around minute ten, sending 10:00 to 11:30 costs roughly 24,000 tokens instead of the ~630,000 for a full 40-minute recording. That is a 25x saving with zero loss on the question actually being asked. Combined with a cheap first pass — a transcript search, or a coarse low-resolution scan — clipping turns "analyse this video" into "analyse these three windows", which is both cheaper and more accurate, because the model is not diluted by irrelevant footage. ## Lever 2 — sample fewer frames The default sampling rate is one frame per second. That rate is tuned for general footage; it is wasteful for content that barely changes. A screen recording of someone typing, a fixed-camera talking head, a slideshow — all can be sampled well below 1 fps with no information loss, because consecutive frames are near-identical. `types.VideoMetadata` exposes an `fps` field for exactly this. The inverse also matters: fast-cut sports or a UI interaction you need to trace step by step may justify sampling *above* the default, spending more tokens deliberately because one frame per second genuinely misses events. The point is that frame rate is a knob you own, not a fixed property of the API. ## Lever 3 — media resolution `GenerateContentConfig` accepts a `media_resolution` setting (`types.MediaResolution`, with LOW, MEDIUM and HIGH variants) that controls how much token budget each frame or image is rendered into. LOW substantially reduces per-frame cost. The tradeoff is concrete rather than vague: low resolution is fine for "what is happening in this scene", "who is on screen", "does the shelf look empty", and poor for "read the error message in this screenshot" or "transcribe the slide". Small text is the first casualty. A common production pattern is a two-stage pipeline: LOW to find candidate moments across a long video, then a short HIGH-resolution clip on the moments that matter. ## Lever 4 — question whether you need video at all Audio runs at roughly 32 tokens per second against video's ~263. If the content is a meeting, a podcast or an interview, the visual channel may contribute nothing to the question you are asking. Sending audio alone is an eightfold saving before any other optimisation. Where a handful of visuals do matter — slides, a whiteboard — a hybrid works well: audio for the full duration plus a few extracted stills as image Parts. ## Lever 5 — segment and summarise Even with every lever pulled, hours of footage will not fit comfortably. The durable pattern is map-reduce: split the media into segments, analyse each in its own request, keep a compact structured summary (timestamps, entities, events) and reason over the summaries rather than the media. This bounds both cost and latency per request, parallelises cleanly, and keeps any single request well inside the context window. ## Measure, then choose Every one of these decisions should be validated with `client.models.count_tokens(...)` on the actual contents you intend to send — the same Parts, the same offsets, the same resolution setting — and reconciled afterwards against the response's `usage_metadata`. Estimating from documented rates is fine for design; shipping without measuring is how a pipeline discovers its unit economics from an invoice. ## The framing that lands in an interview Name the cost driver correctly (duration times sampling rate times per-frame budget), then present the levers in order of leverage: clip, sample less, lower resolution, drop the visual channel, segment. Make the quality tradeoff explicit for each, and say how you would verify the saving. Answering "compress the file" signals that the billing model was never understood.
- What quality do you actually lose at MEDIA_RESOLUTION_LOW?Fine detail first — small on-screen text, dense charts, subtle facial or product distinctions. Scene-level understanding (who is present, what is happening, whether a shelf is empty) survives well. The usual production shape is a low-resolution sweep to locate candidate moments, then a short high-resolution request on just those moments, which keeps the bill low without giving up detail where it matters.
- When would you deliberately sample video above one frame per second?When events are shorter than a second: fast-cut footage, sports, or a UI interaction where you must trace each click. At the default rate those events fall between samples and simply do not exist for the model. Raising fps is a conscious purchase of more tokens to buy temporal fidelity, and it is worth it only over a clipped window, never across an entire long recording.
- How would you architect analysis of a two-hour recording?Segment it. Split into windows, analyse each in its own request with clipping and low resolution, and emit a compact structured summary per window — timestamps, entities, events. Then reason over the summaries, re-requesting only the windows that need detail at higher fidelity. That bounds per-request cost and latency, parallelises cleanly, and keeps every call well inside the context window.
saying these in an interview costs you the question
- Suggests re-encoding at lower bitrate to save tokens
- Thinks downscaling the source video lowers the per-second rate
- Ignores clipping and sends whole recordings by default
- Treats one frame per second as unchangeable
- Uses low media resolution for tasks that require reading small text