How do you send an image to Gemini's generateContent as inline data or a file URI?
answer
- two doors into one parts list
- bytes now, or a URI later
- mime_type is never optional
- inline_data versus file_data
- Part.from_bytes versus Part.from_uri
basics
~20 sGemini takes media as a Part inside contents. Small files ride along as inline_data (raw bytes plus a mime_type); larger or reused files are uploaded to the Files API first and referenced as file_data with a file_uri and mime_type.
solid answer
~40 sA Gemini `generateContent` request carries a `contents` list; each entry has `parts`, and every Part holds exactly one thing — `text`, `inline_data`, or `file_data`. For a small image you build an inline Part: `types.Part.from_bytes(data=raw, mime_type="image/png")` in the google-genai SDK, which becomes `inline_data` with base64 bytes over REST. For anything large or reused, you upload once with `client.files.upload(file="chart.png")` and reference the returned URI with `types.Part.from_uri(file_uri=f.uri, mime_type=f.mime_type)`, producing `file_data`. The `mime_type` is required in both shapes — the model does not sniff it. Text and media Parts live in the same turn, so you put the image and "What does this chart show?" side by side in one `contents` list, and you can include several media Parts in a single request.
code
python · 14 linesfrom google import genai
from google.genai import types
client = genai.Client(api_key="YOUR_KEY")
raw = open("chart.png", "rb").read()
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[
types.Part.from_bytes(data=raw, mime_type="image/png"),
"What trend does this chart show?",
],
)
print(response.text)go deeper
Be able to name the two ways media enters a request — inline bytes or an uploaded file's URI — and say that every media Part needs a mime_type alongside the data.
Explain the Content/Part structure itself: one Part carries one payload, media Parts sit beside text Parts in the same turn, and inline bytes are base64 in the JSON body while file_data is a reference.
Show judgment about which door to use per workload: inline for small one-shot payloads, upload for anything large or reused, and be ready to explain why an arbitrary public URL is not an option.
Own the ingestion boundary: where in your pipeline bytes are fetched and validated, how MIME types are established rather than guessed, and how the choice between inline and uploaded references shapes retry and caching behaviour.
## Where media lives in a Gemini request The Gemini API's `generateContent` takes a `contents` argument: a list of Content objects, each with a `role` (`user` or `model`) and a `parts` array. A **Part** is a union type — it carries exactly one payload. For multimodal work the three that matter are `text`, `inline_data`, and `file_data`. There is no separate top-level `images` or `attachments` field: media is always a Part sitting next to your prompt text in the same conversational turn. The google-genai Python SDK lets you pass a bare string where a text Part is expected, which is why examples often look like `contents=[image_part, "Describe this"]`. Under the hood the SDK normalises that into one user Content with two Parts. ## inline_data — bytes travel in the request body `inline_data` is a Blob: over REST it is `{"inline_data": {"mime_type": "image/jpeg", "data": "<base64>"}}`. In the SDK you build it with `types.Part.from_bytes(data=raw_bytes, mime_type="image/jpeg")`; the SDK does the base64 encoding for you, and it will also accept a PIL image object or a local path as a convenience. Properties worth remembering: the bytes count toward the request's total size budget (roughly 20 MB for the whole payload), base64 inflates binary by about a third on the wire, and the bytes are re-sent on every request — nothing is stored server-side. ## file_data — a URI issued by the Files API The other shape is `{"file_data": {"mime_type": "video/mp4", "file_uri": "…/files/abc123"}}`. You obtain that URI by uploading first: `f = client.files.upload(file="lecture.mp4")` returns a File with `.name` (`files/abc123`), `.uri`, `.mime_type`, `.size_bytes` and `.state`. You then reference it with `types.Part.from_uri(file_uri=f.uri, mime_type=f.mime_type)`, or simply drop the File object itself into `contents` and let the SDK build the Part. The critical constraint is that `file_uri` is **not** a general-purpose URL fetcher. It must be a Files API URI belonging to the same project/API key (a YouTube video link is the one documented special case). Handing Gemini an arbitrary `https://example.com/cat.jpg` does not make it download the image — you fetch the bytes yourself and send them inline, or upload them. ## MIME types are mandatory, and the list is finite Both Part shapes require `mime_type`. Images are `image/png`, `image/jpeg`, `image/webp`, `image/heic`, `image/heif`. Audio includes `audio/wav`, `audio/mp3`, `audio/flac`, `audio/aac`, `audio/ogg`. Video includes `video/mp4`, `video/mpeg`, `video/webm`, `video/mov`, `video/3gpp`. Documents are `application/pdf` plus plain-text types. Vector formats such as SVG are not a supported image input; you rasterise first. A wrong or missing MIME type produces a 400 rather than a silent guess. ## Ordering, multiples, and mixing modalities You may include several media Parts in one turn — a handful of images, an image plus a PDF, audio plus a text instruction. Google's guidance for a single-image prompt is to place the image Part **before** the text Part, which tends to produce better grounding; for multi-image prompts it helps to label each image in the surrounding text ("Image 1: …", "Image 2: …") so the model can refer to them unambiguously. The two Part shapes mix freely inside the same request: a small logo inline next to a 400 MB video referenced by URI is perfectly legal. ## What goes wrong in practice - Passing a public URL as `file_uri` and wondering why the model hallucinates a description — nothing was fetched. - Wrapping base64 in a `data:image/png;base64,...` prefix. The `data` field takes bare base64; the data-URL prefix corrupts it. - Forgetting that inline bytes count against the request-size ceiling, so a 30 MB scan fails where a 300 KB thumbnail succeeds. - Sending an unsupported container (an SVG, or a video codec outside the supported list) and reading the resulting 400 as a model problem. - Assuming media must be a separate call from the prompt. It is one call; the model needs both Parts in the same turn to answer about the media. ## The mental model One list, one Part per payload, one MIME type per Part, and exactly two ways for bytes to arrive: in the body now, or by reference after an upload. Everything else about Gemini multimodality — token cost, size ceilings, file lifecycle — hangs off which of those two doors you used.
- Can you point file_uri at any public image URL on the internet?No. `file_uri` must be a URI issued by the Files API for your own project, with the documented exception of a YouTube video URL. Gemini does not act as a general fetcher, so for a public image you download the bytes yourself and send them as `inline_data`, or upload them first and use the returned URI.
- Does the order of the image Part and the text Part matter?It is not a hard requirement, but Google's prompting guidance is to put a single image before the instruction text, which tends to ground the answer better. With several images, label them in the prose ("Image 1 shows…") so the model can distinguish them, since the Parts themselves carry no names.
- How do you send several images plus a PDF in one request?Append one Part per file to the same `parts` list, each with its own `mime_type`, then the text instruction. They can mix shapes freely — small images inline, the PDF by `file_uri`. The only ceiling is the total request size for the inline ones and the model's context window for the combined token count.
saying these in an interview costs you the question
- Thinks file_uri can be any public https image URL
- Omits mime_type and expects Gemini to sniff the format
- Sends the image in a separate call from the prompt
- Wraps base64 bytes in a data: URL prefix
- Believes only one media item is allowed per request