When must a Gemini request use the Files API instead of inline media bytes?
answer
- the cap is on the request, not the file
- base64 inflates about a third
- two gigabytes, forty-eight hours
- free storage, per-project scope
- reuse is a reason too, not just size
basics
~20 sInline bytes only work while the whole generateContent request stays under about 20 MB, so anything larger — and effectively all video — must go through the Files API, which accepts up to 2 GB per file and keeps it for 48 hours.
solid answer
~50 sThe deciding number is the **total request size**, not the file size: a Gemini `generateContent` call carrying `inline_data` has to fit inside roughly 20 MB including the base64 expansion, your prompt text and system instruction. Under that, inline is simplest — one round trip, nothing to clean up. Above it you upload with `client.files.upload(...)` and pass the returned URI, which lifts the ceiling to 2 GB per file with about 20 GB of storage per project. The Files API is free, but files are deleted after **48 hours**, so a URI is a short-lived handle rather than durable storage. The second reason to upload even for a small file is reuse: an uploaded file is sent once and can be referenced by many requests, whereas inline bytes are re-uploaded on every call. Video and long audio essentially always take the Files API path.
code
python · 16 linesfrom google import genai
client = genai.Client(api_key="YOUR_KEY")
uploaded = client.files.upload(file="quarterly.pdf")
print(uploaded.name, uploaded.mime_type, uploaded.size_bytes)
print(uploaded.state.name, uploaded.expiration_time)
for question in ["Who signed it?", "What is the total?"]:
reply = client.models.generate_content(
model="gemini-2.5-flash",
contents=[uploaded, question],
)
print(reply.text)
client.files.delete(name=uploaded.name)go deeper
Remember the practical rule: small images can go inline, but video and big documents must be uploaded first, and uploaded files disappear after two days.
Explain that the roughly 20 MB ceiling applies to the whole request including base64 expansion and prompt text, and that the Files API raises it to 2 GB per file with 48-hour retention.
Show how you design around expiry and reuse: authoritative bytes in your own storage, the Gemini file as a disposable cache, and a re-upload-and-retry path when a URI has aged out.
Own the storage and quota strategy across services — who holds the project scope that uploads are bound to, how the per-project storage cap is managed at volume, and what the media ingestion contract looks like end to end.
## The ceiling is on the request, not the file The most common mistake is to compare the *file* against 20 MB. The documented limit applies to the **total request payload**: every inline Part plus the prompt text plus the system instruction plus the JSON envelope. Two 9 MB images plus a long system instruction will not fit even though neither image alone is close. There is a second, quieter multiplier: `inline_data` is transmitted as base64, which inflates binary by roughly 33%. A 16 MB JPEG is about 21 MB on the wire — over the line before you have written a word of prompt. ## What the Files API buys you `client.files.upload(file="lecture.mp4")` performs a resumable upload and returns a File resource with `name` (`files/abc123`), `uri`, `mime_type`, `size_bytes`, `state` and an expiration timestamp. You then reference it in `contents` — either as `types.Part.from_uri(file_uri=..., mime_type=...)` or by passing the File object straight into the list. The documented properties: - **Up to 2 GB per file**, with roughly **20 GB of storage per project**. - **Files are stored for 48 hours**, then deleted automatically. You can delete earlier with `client.files.delete(name=...)`. - **Storage is free** — you pay only for the tokens the media becomes when a request consumes it. - Files are **scoped to the project/API key** that uploaded them. Another key cannot read your URI, and you cannot download the bytes back; the API exposes metadata only. ## Reuse, not just size Even a 200 KB image is worth uploading if you are going to ask twenty questions about it. Inline bytes are re-sent, re-base64'd and re-parsed on every single call; an uploaded file is transferred once and referenced cheaply thereafter. In a chat UI where the user drops a PDF and then asks follow-up questions, uploading once and holding the URI for the session is the right shape — as long as the session cannot outlive the 48-hour window. Note that this is a *transport* optimisation, not a *token* optimisation: referencing a file by URI still re-tokenises the media on every request. Reducing repeated token cost is what context caching addresses, which is a separate mechanism. ## The 48-hour rule has architectural teeth Because URIs expire, a `file_uri` is not a durable identifier. Persisting one in your database as "the customer's uploaded contract" is a latent bug: two days later the reference stops resolving and the feature breaks. The durable copy of the bytes belongs in your own object storage; the Gemini file is a cache entry you can always recreate by re-uploading. Production code should treat "file missing or expired" as a normal, recoverable branch that re-uploads and retries. ## Choosing per modality - **Images** — usually inline. A photo or screenshot is typically well under a megabyte; upload only when reusing across many turns or when batching many images into one request. - **PDFs and documents** — depends on size. A 3-page invoice is fine inline; a 300-page manual is not. - **Audio** — short clips can be inline; a podcast episode cannot. A minute of decent-quality audio is already several megabytes. - **Video** — essentially always the Files API. Even a couple of minutes of 1080p footage blows past 20 MB. - **YouTube** — a documented special case: you can pass a YouTube URL as `file_data` without uploading anything, subject to per-tier quotas on how much YouTube video you may process. ## Operational details worth naming in an interview Uploads of larger files return a file in `PROCESSING` state, and a request that references it before it becomes `ACTIVE` fails. Uploads themselves can fail on network hiccups; because the file is content you already hold, retrying an upload is safe from your side — you simply get a new file name. And because storage is free while tokens are not, there is no cost argument for keeping files short-lived; the reason to delete early is hygiene and the per-project storage cap. ## The one-line answer Inline while the *whole request* fits in ~20 MB and the media is used once; Files API above that, for anything reused, and for video by default — remembering that the URI you get back lives for 48 hours and nothing longer.
- Does referencing an uploaded file instead of inline bytes reduce the token cost of a request?No. The upload changes how the bytes reach Google, not how the model consumes them — the media is re-tokenised on every request either way, and you are billed for those tokens each time. Files API storage itself is free. Cutting the repeated token cost is a job for context caching, not for the Files API.
- Your service stores a Gemini file_uri in Postgres against a customer document. What breaks?It breaks after 48 hours, when the Files API deletes the object and the URI stops resolving. Keep the authoritative bytes in your own storage and treat the Gemini file as a disposable cache: on a missing-file error, re-upload and retry transparently rather than surfacing the failure.
- Why can a 15 MB image still fail a Gemini request when the documented limit is 20 MB?Because base64 encoding inflates binary by about a third, so 15 MB becomes roughly 20 MB on the wire, and the limit covers the entire payload — prompt text, system instruction and JSON envelope included. Any other inline Part in the same request eats the remaining headroom too.
saying these in an interview costs you the question
- Compares only the file size against the 20 MB limit
- Thinks the Files API reduces per-request token cost
- Treats a file_uri as permanent storage
- Ignores base64 inflation when sizing inline payloads
- Assumes uploaded files are readable by any API key