How does Gemini turn a minute of video or audio into billable tokens?
answer
- duration drives it, not bytes
- one frame sampled per second
- audio is roughly thirty tokens a second
- video runs about eight times audio
- count_tokens before you commit
basics
~20 sGemini bills time-based media per second: roughly 32 tokens per second of audio (about 1,900 per minute) and about 263 tokens per second of video at default resolution (roughly 16,000 per minute), because video is sampled at one frame per second and its audio track counted too.
solid answer
~50 sNon-text modalities are converted to tokens by fixed documented rates rather than by file size. **Audio** costs about **32 tokens per second**, so a one-minute clip is roughly 1,900 tokens and an hour is over 115,000. **Video** costs about **263 tokens per second** at default media resolution — Gemini samples frames at one per second and also processes the audio track — so a minute is around 16,000 tokens and a ten-minute clip already consumes a sizeable slice of the context window. **Images** are cheaper and flat: an image with both dimensions at or below 384 px costs 258 tokens, and larger images are split into tiles that each cost the same 258. Bitrate, codec and file size do not enter the calculation — only duration and, for images, pixel dimensions. Before sending anything expensive, verify with `client.models.count_tokens(...)`, which accepts the same multimodal `contents` you would send to `generate_content`.
code
python · 10 linesfrom google import genai
client = genai.Client(api_key="YOUR_KEY")
clip = client.files.upload(file="interview.mp3")
counted = client.models.count_tokens(
model="gemini-2.5-flash",
contents=[clip, "Summarise the main points."],
)
print(counted.total_tokens)go deeper
Know that images, audio and video all become tokens you pay for, and that a minute of video costs far more than a minute of audio.
Quote the mechanism: roughly 32 tokens per audio second, roughly 263 per video second with frames sampled at one per second, and 258 tokens per image tile or PDF page.
Demonstrate budgeting: pre-flight with count_tokens, reconcile against usage_metadata, and recognise that long media threatens latency and context headroom, not only cost.
Own the economics of the pipeline — whether the workload needs video at all versus audio plus sparse stills, and how per-second rates set the unit cost of every feature built on top.
## Duration, not bytes The single most useful reframing here: for Gemini, media cost is a function of **time and dimensions**, never of file size. Re-encoding an hour of audio from 320 kbps to 64 kbps shrinks the upload by 80% and changes the token bill by exactly zero. A candidate who says "compress it to save tokens" has the wrong model of the pricing. ## The documented rates (Gemini 2.x families, mid-2026) - **Audio: ~32 tokens per second.** One minute is about 1,920 tokens; one hour about 115,000. - **Video: ~263 tokens per second** at default media resolution. One minute is about 15,800; ten minutes about 158,000. The figure reflects both the sampled frame and the accompanying audio. - **Images: 258 tokens** when both dimensions are at or below 384 px. Larger images are cropped into tiles, each resized to a fixed working size, and each tile costs 258 tokens — so a large screenshot costs a multiple of 258, not a smoothly scaling number. - **PDF pages** are treated like images: each page is documented at 258 tokens, with a cap of about 1,000 pages per document. A 40-page report is therefore roughly 10,000 tokens of page imagery before any output. These rates are stable enough to plan with, but they are model-family properties and have changed between generations — quote them with the family you are assuming, and check `count_tokens` rather than trusting memory in production. ## Why video is so much more expensive than audio Video is not "understood" as a stream. By default Gemini samples it at **one frame per second** and turns each sampled frame into image tokens, alongside the audio track at the audio rate. That is why the per-second video number is roughly eight times the audio one, and why a static talking-head recording costs exactly as much as fast-cut action footage of the same length — the sampler does not care that nothing moved. It is also why *duration*, not the resolution of the source, drives the bill: a 4K video and a 480p video of the same length cost the same, because frames are normalised before tokenisation. ## Context window pressure, not just money Token cost has two consequences and interviews usually probe the second one. At roughly 16,000 tokens per video minute, an hour of footage approaches a million tokens — it competes directly with everything else in the context window, and it slows the request down: prefill work scales with input tokens, so a long video means a long time to first token even before generation starts. This is why long-video workflows are usually segmented: analyse chunks, keep short structured summaries, and let the summaries rather than the raw media carry forward. ## Verifying instead of estimating `client.models.count_tokens(model=..., contents=[...])` accepts exactly the same multimodal contents as `generate_content`, including uploaded-file Parts, and returns `total_tokens`. Two practical uses: a pre-flight check so an over-budget request fails fast and locally rather than after a large upload, and a regression guard so a change in the documented rates shows up in your dashboards rather than in a surprise invoice. The response's `usage_metadata` closes the loop after the fact, reporting the prompt and candidate token counts the request was actually billed for. ## Estimating on the spot Useful round numbers to carry into an interview: **one minute of audio is about 2,000 tokens**, **one minute of video about 16,000 tokens**, **one image tile or PDF page 258 tokens**. From those, a 30-minute meeting recording is about 58,000 tokens of audio, but the same meeting as video is about 470,000 — an eightfold difference that usually settles the design question of whether you need the video at all. ## Where people go wrong Assuming file size drives the bill; assuming a lower-bitrate or lower-resolution source is cheaper; assuming silence or a static frame costs less; forgetting that video includes its audio track; and treating an image as a single flat 258 tokens regardless of size, when in fact anything much beyond 384 px on a side is billed as several tiles.
- Does re-encoding a video at a lower bitrate reduce its Gemini token cost?No. Token cost is driven by duration and the sampling rate of frames, not by file size or bitrate. Re-encoding shrinks the upload and speeds the transfer, but the model still sees roughly one frame per second plus the audio track, so the bill is unchanged. To spend fewer tokens you must shorten the media, lower the media resolution, or sample fewer frames.
- How are the pages of a PDF counted?Each page is handled like an image and documented at 258 tokens, with a limit of around 1,000 pages per document. So a 100-page report is roughly 25,800 tokens of page imagery before any output. That makes page count, not file size, the number to check when deciding whether a document fits your budget.
- Where do you see the tokens a multimodal request actually consumed?In the response's `usage_metadata`, which reports prompt and candidate token counts for the call that was billed. Pair it with a `count_tokens` pre-flight on the same contents: the pre-flight lets you reject an over-budget request before uploading, and the post-hoc metadata is what you aggregate for cost dashboards.
saying these in an interview costs you the question
- Thinks compressing the file lowers the token count
- Assumes silence or static frames cost fewer tokens
- Forgets video includes its audio track
- Treats any image as a flat 258 tokens regardless of size
- Believes source resolution changes the per-second video rate