Which image formats and size limits does Claude's Messages API accept?
answer
- four MIME types, nothing exotic
- vector and phone-camera formats need converting
- five megabytes is the per-image ceiling
- encoding inflates the body by a third
- size limit and resolution cap are separate gates
basics
~20 sFour media types are accepted: image/jpeg, image/png, image/gif and image/webp. Each image may be up to 5 MB through the API, and a single request may carry up to 100 images. Base64 encoding inflates the payload by about a third.
solid answer
~50 sThe accepted `media_type` values are `image/jpeg`, `image/png`, `image/gif` and `image/webp` — nothing else, so SVG, TIFF, HEIC and raw camera formats must be converted before they are sent. Per image the API limit is 5 MB, and a single request may include up to 100 images. Two practical consequences follow. First, base64 encoding grows the payload by roughly 33%, so a file that is comfortably under the limit on disk can still bloat the request; the URL and Files API source shapes avoid inlining bytes altogether. Second, size limits are independent of the resolution cap — a 5 MB file is accepted and then downscaled to the model's supported long edge, so passing the size check says nothing about how many pixels the model actually sees. Validate the real MIME type from file headers rather than trusting the filename extension.
go deeper
Memorise the four accepted media types and the 5 MB per-image ceiling, and know that formats like SVG or HEIC have to be converted before you can send them.
Explain why the declared media_type must match the real bytes, and separate the size limit from the resolution cap — one gates the request, the other gates what the model sees.
Describe a hardened ingest path for untrusted uploads: header-based type detection, conversion, resizing to the model cap, re-encoding, and turning would-be runtime 400s into a deterministic pre-flight step.
Decide where normalisation lives — at the edge, in a shared service, or per caller — and what the contract is when a format cannot be converted, so that limits are enforced once rather than rediscovered in each team's client.
## The accepted media types Four values are legal in the `media_type` field of a base64 image source: `image/jpeg`, `image/png`, `image/gif` and `image/webp`. Anything else is rejected. That list is narrower than what a browser or an operating system will happily display, and the gap is where real pipelines break: `image/svg+xml` is a vector format and is not accepted, iPhone uploads arrive as HEIC, scanned archives are frequently TIFF, and design tools export PSD or AVIF. Every one of those needs a conversion step to PNG or JPEG before it reaches the API. The declared `media_type` must describe the actual bytes. It is a decoding instruction, not a label, so a PNG announced as `image/jpeg` fails rather than being auto-detected. When the file came from a user, sniff the type from the magic bytes in the header; the extension is whatever the uploader typed. The `url` source shape carries no `media_type` at all — the server determines it from the fetch — and the Files API shape derives it from the MIME type recorded at upload time. The rule that the content-block type must match the file type still applies there: images go in an `image` block, while PDFs and text files go in a `document` block. ## The size ceiling and the base64 tax Each image may be up to 5 MB through the API. The subtlety is that base64 is not free: encoding inflates binary data by about a third, so 5 MB of pixels becomes roughly 6.7 MB of characters in the JSON body, and that body also has to carry your system prompt, tool definitions and conversation history. Requests that carry several large images can therefore run into overall request-size pressure well before any individual image is anywhere near its own limit. The two non-inline source shapes exist partly for this reason. A `url` source moves the transfer to Anthropic's side entirely, and a Files API `file_id` moves it to a one-time upload that later requests merely reference. Both keep the message body small. ## Count limits A single request may carry up to 100 images. That is a generous ceiling in absolute terms and an almost irrelevant one in practice, because token cost binds first: a hundred images at the high-resolution per-image ceiling is a very large prompt before you have written a word of text. Treat the count limit as a hard backstop and your token budget as the real constraint. ## Size limits are not resolution limits These are two independent gates and conflating them is the most common misunderstanding here. The 5 MB rule is about bytes in the request. The resolution cap — the model generation's maximum long edge — is about pixels the model sees, and it is applied by downscaling after the request is accepted. A 4 MB, 8000-pixel-wide scan passes the size check and is then resized down before tokenisation. Conversely, a small, highly compressed file can be well under 5 MB and still be at full resolution. The operational reading: client-side resizing is what controls cost and latency, while client-side compression is what keeps you under the size ceiling. You usually want both. ## Building a robust ingest path For anything that accepts user uploads, a defensive pipeline looks like: detect the true type from the header bytes; reject or convert anything outside the four accepted types; strip metadata; resize so the long edge matches the target model's cap; re-encode as JPEG for photographs or PNG for screenshots and diagrams; then check the encoded size and compress further if needed. Doing that work locally turns a class of runtime 400s into a deterministic step you control, and it usually cuts the token bill at the same time because most source images are far larger than the model's cap.
- A user uploads an SVG logo for analysis. What do you do before calling the API?Rasterise it. SVG is not among the accepted media types, so the vector file must be rendered to PNG at a sensible resolution and sent as image/png. Pick the render size deliberately — an SVG has no intrinsic pixel size, so you are choosing the token cost when you choose the raster dimensions.
- Your 4.8 MB PNG is under the per-image limit but the request still feels heavy. Why?Base64 inflates binary data by roughly a third, so that PNG becomes about 6.4 MB of characters inside the JSON body, on top of your system prompt, tools and history. Either compress or resize before encoding, or switch to a url source or a Files API file_id so the bytes never ride inside the message body.
- Does passing the 5 MB size check mean the model sees the image at full resolution?No. Size and resolution are separate gates. A file under 5 MB is accepted, and then downscaled server-side if its long edge exceeds the model generation's cap. Fidelity is governed by pixel dimensions, so control that client-side rather than inferring it from file size.
saying these in an interview costs you the question
- Assuming any format a browser renders will be accepted
- Trusting the filename extension as the media type
- Treating the 5 MB limit as a resolution guarantee
- Ignoring the roughly 33% base64 expansion
- Thinking the image count limit binds before token cost