How does Anthropic's Messages API turn image dimensions into billed input tokens?
answer
- cost follows pixels, not file size
- area over a documented divisor
- server resizes before it counts
- per-image ceiling differs by model generation
- folded into usage.input_tokens
basics
~20 sImage cost scales with pixel area, not file size: Anthropic documents roughly width times height divided by 750 tokens. Oversized images are downscaled server-side first, so you pay for the resized version, capped by the model's resolution limit.
solid answer
~50 sImage tokens track **pixel area**, not bytes on disk — a heavily compressed 200 KB photo and a 3 MB PNG of the same dimensions cost the same. Anthropic's documented approximation is roughly `(width × height) / 750` tokens. Anything above the model's supported resolution is downscaled server-side before tokenisation, so the cap matters more than the original file: models up to the Claude Opus 4.6 / Sonnet 4.6 generation cap the long edge at 1568 px and land near 1600 tokens per image, while the high-resolution tier introduced with Opus 4.7 (and continued on Opus 4.8, Opus 5 and Sonnet 5) allows 2576 px on the long edge and up to about 4784 tokens per image. Those tokens are folded into `usage.input_tokens`, not billed as a separate image line. To price a request exactly, call the token-counting endpoint with the same model and the same content blocks before sending.
code
python · 18 linesimport base64
from anthropic import Anthropic
client = Anthropic()
encoded = base64.standard_b64encode(open("screenshot.png", "rb").read()).decode()
count = client.messages.count_tokens(
model="claude-opus-5",
messages=[{
"role": "user",
"content": [
{"type": "image", "source": {
"type": "base64", "media_type": "image/png", "data": encoded}},
{"type": "text", "text": "Summarise this screen."},
],
}],
)
print(count.input_tokens)go deeper
Remember the direction of the relationship: bigger pixel dimensions mean more tokens, and images are billed as ordinary input. Do not quote file size in megabytes as a cost figure.
Explain the area-over-750 approximation and the server-side downscale step, and be able to say roughly what a 1000x1000 image costs and why compression does not change it.
Show you measure with the count-tokens endpoint against the exact model, prune images from conversation history once consumed, and treat a model upgrade as a cost regression risk for vision traffic.
Own the resolution policy: where the accuracy-versus-token curve flattens for your document or screenshot mix, whether a cheaper model at full resolution beats a stronger one at half, and how that policy is evaluated when a new tier ships.
## Area, not bytes The intuition to unlearn is that a smaller file is a cheaper request. Compression changes the number of bytes on the wire but not the number of pixels the model has to look at, and the model's cost is driven by pixels. A 1024x1024 JPEG saved at quality 40 and the same picture saved as a lossless PNG contain identical dimensions and therefore cost the same in tokens. Cropping and downscaling reduce cost; re-encoding at a lower quality does not. Anthropic documents the relationship as an approximation: tokens are roughly the pixel area divided by 750. A 1000x1000 image is therefore on the order of 1300 tokens, and a 500x500 image is around 330. The divisor is a documented rule of thumb rather than an exact tokenizer, which is why the count-tokens endpoint exists. ## The resolution cap does most of the work Before any of that arithmetic applies, the service resizes images that exceed the model's supported resolution. This is the single most important fact for cost, because it means enormous inputs do not produce enormous bills — but it also means sending enormous inputs is pure waste. You paid the upload bandwidth and the latency for pixels that were thrown away before the model saw them. Two tiers exist as of mid-2026. Models through the Claude Opus 4.6 and Sonnet 4.6 generation resize so the long edge is at most 1568 px, which puts a ceiling of roughly 1600 tokens on any single image. Starting with Claude Opus 4.7, and continuing on Opus 4.8, Opus 5 and Sonnet 5, high-resolution vision raises the long edge to 2576 px and the per-image ceiling to about 4784 tokens. The same picture can therefore cost roughly three times as much on a high-resolution model as on an older one — a real and easily missed regression when you upgrade a vision pipeline by swapping a model string. The upside of the high-resolution tier is not only accuracy on dense documents and screenshots; it also means coordinates the model reports map one-to-one onto the real image pixels, so scale-factor arithmetic that older computer-use code carried around can be deleted. ## Where the tokens show up Image tokens are input tokens. They are added into `usage.input_tokens` on the response and billed at the model's normal input rate. There is no per-image fee and no separate vision multiplier in the price list. Practically, that means a vision workload's cost curve is the same shape as a long-prompt workload's: dominated by what you send, on every single turn, for as long as the image stays in the conversation history. That last clause is the part teams underestimate. In a multi-turn agent, an image sent on turn one is resent with every subsequent request unless you prune it, so its cost is paid again each turn. Dropping stale screenshots out of the history, or replacing them with a short text summary of what was seen, is often a larger saving than tuning resolution. ## Measuring instead of guessing The reliable way to price a vision request is the token-counting endpoint, `POST /v1/messages/count_tokens`, called with the same model identifier and the same `messages` array you intend to send. It returns `input_tokens` for the whole request including images, it is stateless, and it costs nothing to call. Counts are model-specific, which matters precisely because of the two resolution tiers — counting against one model and billing against another gives you the wrong number. Do not reach for a third-party tokenizer library here. OpenAI's tokenizer has no notion of Claude's image handling at all, and it is wrong for Claude's text as well. ## Turning this into a policy A workable default is to downscale on the client so the long edge matches the target model's cap, crop to the region that actually answers the question, and only then send. Where fidelity genuinely matters — dense tables, small type, UI screenshots the model must click into — send at the cap and accept the higher count, because a second round trip caused by an unreadable image costs more than the tokens saved. Measure both variants with count-tokens and with an accuracy eval before freezing the policy.
- You upgrade a screenshot pipeline to a high-resolution model and costs jump. Why?The per-image ceiling moved. Models through the Opus 4.6 generation downscale to a 1568 px long edge and top out near 1600 image tokens; the high-resolution tier from Opus 4.7 onward allows 2576 px and up to roughly 4784 tokens. Images you were already sending at full size now survive resizing at higher fidelity, so the same inputs bill more. Downscale client-side if the extra detail buys nothing.
- Does re-encoding a photo as a smaller JPEG reduce its token cost?No. Token cost follows pixel dimensions, so quality-based compression shrinks the upload without changing the count. Only resizing or cropping reduces tokens. Compression still helps with the per-image size limit and with request latency, so it is worth doing — just not as a billing lever.
- How do images interact with the cost of a long multi-turn conversation?Every image in the history is re-sent, and therefore re-billed as input, on each subsequent request. A ten-turn agent that keeps three screenshots pays for them ten times. Prune images once their information has been extracted, or replace them with a short text description of what was observed.
saying these in an interview costs you the question
- Assuming file size in kilobytes drives the token cost
- Believing images are billed separately from input tokens
- Thinking oversized images are rejected rather than downscaled
- Using an OpenAI tokenizer to estimate Claude image tokens
- Forgetting history re-sends images on every later turn