How do you combine several images with text in one Claude Messages API turn?
answer
- one array, many image blocks
- order is what the model reads
- name each image in adjacent text
- token counts add up, they do not merge
- independent images belong in separate requests
basics
~20 sPut every image as its own block in the same user message content array, place them before the text that asks about them, and label them in the text so answers can refer to them unambiguously. Each image is tokenised and billed separately.
solid answer
~50 sOne user message's `content` array can hold many `image` blocks interleaved with `text` blocks, and order is what the model reads. Anthropic's documented guidance is to put the images ahead of the question and to label them in the prose — for example a short `text` block reading "Image 1: before" in front of each one — so that the model's answer can reference them without ambiguity. Every image is tokenised independently and the counts sum, so a four-screenshot turn costs roughly four times a one-screenshot turn; the per-request ceiling of 100 images is generous but the token budget binds long before it. Comparison tasks belong in a single turn, because the model can only contrast images it sees together. Conversely, unrelated images should be separate requests: they add cost to every subsequent turn of the conversation without adding signal.
code
json · 12 lines{
"role": "user",
"content": [
{ "type": "text", "text": "Image 1 - before the deploy:" },
{ "type": "image",
"source": { "type": "url", "url": "https://example.com/before.png" } },
{ "type": "text", "text": "Image 2 - after the deploy:" },
{ "type": "image",
"source": { "type": "url", "url": "https://example.com/after.png" } },
{ "type": "text", "text": "Which UI elements changed between them?" }
]
}go deeper
Know that several image blocks can live in one user message alongside text, and that you label them in the text so the answer can refer to them clearly.
Explain that block order is meaningful, that images should precede the question, and that per-image token counts simply add up with no volume discount.
Show judgment about batching: comparisons in one turn, independent items as separate requests, and pruning consumed screenshots out of history so a long agent loop does not re-pay for them every turn.
Own the context strategy for a vision-heavy product — how many images a turn may carry, when images are replaced by text summaries, and how that budget is enforced so no single team's prompt destabilises shared cost and latency targets.
## Several blocks, one array The `content` array of a user message is an ordered sequence, and nothing restricts it to one image. A comparison turn typically looks like: text block naming the first image, image block, text block naming the second image, image block, then a final text block with the actual question. That structure is not required by the schema — a bare run of image blocks followed by one question works — but it is what makes multi-image answers reliable. ## Why labelling matters Without labels, the model has to refer to the images positionally, and so do you. Answers drift into "the first picture" and "the other one", which is fine for a human reading the transcript and terrible for a program parsing it. Naming the images in adjacent text blocks gives both sides a shared vocabulary: you can ask "which elements changed between Screenshot A and Screenshot B" and get an answer that uses those names. This is the documented recommendation and it costs a handful of tokens. The related recommendation is to put images before the text that asks about them. The model reads the array in order, so presenting the evidence before the question mirrors how a person would be briefed. ## Cost is additive, and it repeats Each image is tokenised on its own and the counts sum into `usage.input_tokens`. There is no discount for volume and no deduplication of identical images. At the high-resolution tier a single image can approach roughly 4784 tokens, so a handful of full-resolution screenshots is already a substantial prompt before any text. Because the Messages API is stateless, every image in the history is re-sent — and re-billed — on each subsequent request in that conversation. A multi-image turn early in a long agent loop is therefore not a one-time cost. Once the model has extracted what it needs, replacing those image blocks with a short text summary of the findings is one of the largest savings available in a vision-heavy agent, and it usually improves focus as well. ## Choosing a batching strategy Put images in the same turn when the task genuinely requires seeing them together: before/after diffs, picking the best of several candidates, reading a document split across page images, tracking a UI flow across steps. Split them into separate requests when the images are independent — classifying a thousand product photos is a thousand requests (or a batch job), not one request with a thousand images, because separate requests parallelise, fail independently, retry cheaply, and never carry each other's tokens. A useful middle ground for very long document flows is chunking: send pages in groups, summarise each group into text, and carry only the summaries forward. That keeps any single request within a sane token budget while preserving continuity. ## Limits and legality Up to 100 images may appear in one request. Image blocks remain user-side only, including when they arrive nested inside a `tool_result` block — for example an agent tool that returns a screenshot. The assistant never produces image blocks, so a multi-image conversation is always you supplying pictures and Claude answering in text. ## Ordering pitfalls Two failure patterns recur. The first is asking the question first and appending the images afterwards, which weakens grounding for exactly the reason the guidance exists. The second is interleaving so heavily that the question is buried — a long alternation of images and commentary with the actual instruction lost in the middle. Keep the structure simple: label, image, label, image, then one clear instruction at the end.
- Would you classify 500 independent product photos in one request or 500 requests?Separate requests, or a batch job. Independent items gain nothing from sharing a prompt, and bundling them means one oversized request that retries as a whole, fails as a whole, and carries every image's tokens on every turn. Separate calls parallelise, isolate failures, and keep each response easy to attribute to its input.
- How do you keep a long agent loop that takes screenshots from blowing up its context?Prune. Once the model has read a screenshot and stated what it saw, drop that image block from the history and keep the text conclusion instead. Because history is re-sent every turn, a stale screenshot is re-billed indefinitely. Keeping only the most recent screenshot plus text summaries of earlier ones is a common and effective pattern.
- Does Anthropic recommend any particular ordering of image and text blocks?Yes — place images before the text that asks about them, and when there are several, label them in adjacent text blocks. The model reads the array in order, so evidence-then-question grounds the answer better, and explicit labels let both the prompt and the response refer to specific images unambiguously.
saying these in an interview costs you the question
- Expecting a bulk discount for several images in one request
- Appending images after the question and losing grounding
- Bundling unrelated images into one oversized request
- Forgetting that history re-sends every image each turn
- Thinking the assistant can reply with an image block