How do you send an image to Claude's Messages API inside a user message?
answer
- content becomes an array, not a string
- one block type for image, one for text
- source object picks how bytes arrive
- three source types: base64, url, file
- media_type only on the base64 form
basics
~20 sSet the user message content to an array of blocks and add an image block next to your text block. The image block carries a source object that is either base64 (with media_type and data) or a url. Image blocks belong in user turns.
solid answer
~40 sAn image travels as a **content block**, not as a separate parameter. Set the user message's `content` to an array and include an `image` block alongside a `text` block. The `source` object inside it takes one of three shapes: type `base64` with `media_type` (for example `image/png`) plus a `data` string, type `url` with a `url` field so Anthropic's servers fetch the bytes, or type `file` with a `file_id` for an image already uploaded through the Files API. The `url` form carries no `media_type` — the server infers it. Image blocks are valid only in `user`-role turns; the assistant never emits one. Anthropic's guidance is to put the image block before the text that asks about it. The pixels are billed as ordinary input tokens, not as a separate image charge.
code
python · 27 linesimport base64
from anthropic import Anthropic
client = Anthropic()
with open("chart.png", "rb") as f:
encoded = base64.standard_b64encode(f.read()).decode("utf-8")
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=[{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/png",
"data": encoded,
},
},
{"type": "text", "text": "What trend does this chart show?"},
],
}],
)
print(response.content[0].text)go deeper
Be ready to write the request from memory: content becomes an array, the image block holds a source, and base64 needs media_type plus data. Say plainly that images go in user messages.
Explain why there are three source shapes and what each costs you operationally — inline bytes, a server-side fetch, or a one-time upload referenced by file_id — and that block order matters to the model.
Show that you validate media types against real bytes on ingest, handle oversized or unsupported uploads before the API call, and know that image pixels land in usage.input_tokens rather than a separate charge.
Own the decision of where image bytes live in your architecture: inlined per request, fetched from a public URL, or uploaded once and referenced, and what each choice implies for egress, retention and portability across hosting platforms.
## The content-block model Anthropic's Messages API has no `image` parameter and no separate vision endpoint. Everything goes through `POST /v1/messages`, and multimodality is expressed by widening one field: a message's `content` may be a plain string, or an **array of typed content blocks**. Text is one block type, images are another. So the move from a text-only call to a vision call is not a different API — it is replacing a string `content` with an array holding an `image` block and a `text` block. Every other request field (`model`, the required `max_tokens`, `system`, `tools`, `stream`) is unchanged, and the response comes back in exactly the same shape as a text-only call. ## The three source shapes The `image` block has one required child, `source`, and `source.type` selects how the bytes reach Anthropic. **base64** — you read the file yourself, base64-encode it, and inline it. You must also supply `media_type`, and it must match the real bytes: labelling a PNG as `image/jpeg` is a request error rather than something the server silently corrects. The encoded string must not contain newlines. **url** — you supply an `https` URL and Anthropic's servers download the image during the request. There is no `media_type` field here; the fetch determines it. The URL has to be reachable from Anthropic's network, so private object-storage paths and links behind your own auth do not work unless you pre-sign them. **file** — you upload the image once to the Files API and pass the returned identifier as `file_id`. The block type must match the file's MIME type: images go in an `image` block, PDFs and text go in a `document` block. The Files API is a beta surface, so the beta flag `files-api-2025-04-14` has to be present on both the upload and the message request. ## Where image blocks are legal Image blocks are an input-side construct. They may appear in `user`-role messages, including inside a `tool_result` block that returns a screenshot to the model. They may not appear in an assistant message you construct — Claude produces text, `thinking`, and `tool_use` blocks, never an `image` block. If you want the model to produce a picture, you need an image-generation tool or a different product entirely; the Messages API only reads images. ## Ordering inside the array Block order is meaningful because it is the order the model reads. Anthropic's documented recommendation is to place the image before the text that refers to it, which mirrors how the question makes sense to a human reader: look, then ask. With several images, label them in the text ("Image 1 is the before state, Image 2 is the after state") so that references in the answer are unambiguous. ## What it costs and what comes back Image content is converted into input tokens and billed at the model's normal input rate. There is no separate per-image line item; the count simply shows up inside `usage.input_tokens` on the response, folded together with your text. That means anything that reduces the number or the resolution of images reduces spend directly, and it also means the token-counting endpoint (`POST /v1/messages/count_tokens`) can price a vision request before you send it. ## Common mistakes The most frequent one is reaching for another vendor's shape. There is no `image_url` block type here and no nested object with a `url` key under `image_url`; that is a different provider's schema. The second is forgetting that `content` must become an array — leaving it as a string and trying to attach the image elsewhere in the request body produces a validation error. The third is sending a data URI (`data:image/png;base64,...`) as the `data` value; the field wants the bare base64 payload, not the URI prefix. The fourth is omitting `max_tokens`, which the Messages API requires on every call, vision or not.
- Can Claude return an image block in its response the way it returns text?No. The assistant side of the Messages API emits text, thinking, and tool_use blocks only — there is no assistant image block. Vision here is strictly an input capability: Claude reads pixels you send and answers in text. Producing pictures requires an image-generation model or a tool you expose, not this endpoint.
- What happens if the media_type you declare does not match the actual bytes?The request fails validation rather than being quietly corrected. The declared media_type is how the API decodes the payload, so a PNG labelled image/jpeg is rejected. In pipelines that accept user uploads, sniff the real type from the file header instead of trusting the filename extension, and normalise unsupported types before building the block.
- Where do images fit when you are returning a screenshot from a tool call?A tool_result block's content can itself be an array of blocks, so you nest the image block inside the tool_result rather than adding a bare image alongside it. Because tool_result blocks live in user-role messages, the rule that images only appear on the user side still holds.
Think of the content array as an email body: the attachment and the sentence asking about it travel together in one message, in the order the reader will meet them.
saying these in an interview costs you the question
- Looking for an images endpoint separate from /v1/messages
- Using an image_url block, which is another vendor's schema
- Putting an image block in an assistant message
- Passing a full data: URI as the base64 data value
- Leaving content as a string and attaching the image elsewhere