In a vision-language model, why can one image cost more tokens than a page of text?
answer
- Not one token per image
- Think grid, not picture
- Cost follows area, not side length
- One token per patch, roughly
- Doubling each side quadruples the count
basics
~20 sThe vision encoder cuts the image into a grid of fixed-size patches and emits roughly one token per patch, so token count grows with image area. A 1024x1024 image at patch size 16 is 64x64 = 4096 patches.
solid answer
~50 sImages are not one token. A Vision Transformer splits the image into a grid of non-overlapping patches — 14x14 or 16x16 pixels are typical — and produces an embedding per patch, and the connector hands roughly one decoder token per patch to the language model. So the cost scales with **area**: at patch size 16, a 512x512 image is 32x32 = 1024 tokens, and doubling each side to 1024x1024 quadruples that to 4096. Add high-resolution tiling and a single photo can consume more of the context window than an entire text prompt. Models fight back by pooling or merging neighbouring patches (a 2x2 merge cuts tokens 4x), by resampling to a fixed query count, and by capping how many tiles a single image may use. The practical rule: send the smallest resolution at which the detail you care about is still legible.
code
python · 7 linesdef image_tokens(width, height, patch=16, pool=1):
grid_w, grid_h = width // patch, height // patch
return (grid_w // pool) * (grid_h // pool)
print(image_tokens(336, 336, patch=14)) # 576
print(image_tokens(1024, 1024, patch=16)) # 4096
print(image_tokens(1024, 1024, patch=16, pool=2)) # 1024go deeper
Know that an image is split into patches and costs roughly one token per patch, so images are expensive and larger images are much more expensive. Be able to say a 1024x1024 image is thousands of tokens, not one.
Do the arithmetic out loud: grid = side / patch size, tokens ≈ patches over any pooling factor, and cost scales with area. Explain why sub-patch text is unreadable no matter how you prompt.
Turn the arithmetic into a budget: choose resolution from the smallest feature you must resolve, crop instead of upscaling, and account for images persisting in conversation history and inflating prefill latency.
Own the fleet-level trade: what image resolution policy your product defaults to, how image tokens shape context-window and cost forecasts at volume, and when the answer is a cheaper preprocessing step rather than a bigger context.
## Why the question comes up at all Engineers new to vision-language models routinely underestimate image cost by an order of magnitude, blow their context budget, and then wonder why latency tripled. The arithmetic is simple once you know where the tokens come from. ## Patching: the source of the number A Vision Transformer (ViT) cannot attend over pixels — a 1024x1024 image has over a million of them, and attention is quadratic in sequence length. Instead the image is divided into a grid of non-overlapping square **patches**. Each patch is flattened and linearly embedded into one vector, positional information is added, and the transformer runs over that sequence. Patch size is a property of the encoder, written into names like ViT-L/14 (patch size 14) or ViT-B/16. The grid is therefore: patches = floor(width / patch) * floor(height / patch) Concrete numbers worth memorizing: - 336x336 at patch 14 → 24 x 24 = **576** patches (the classic LLaVA-1.5 configuration) - 512x512 at patch 16 → 32 x 32 = **1024** - 1024x1024 at patch 16 → 64 x 64 = **4096** ## From patches to decoder tokens What the decoder pays for is not patches but the vectors the connector hands it. With a projection-style connector (the dominant design) that is one token per patch, so the two numbers coincide. Two mechanisms break the identity: - **Pooling / patch merging.** Many models merge a 2x2 block of neighbouring patch embeddings into a single vector before projection (pixel-shuffle or pixel-unshuffle style). That is an immediate 4x reduction: 4096 patches become 1024 tokens, at some cost in fine detail. - **Query-based resampling.** A Q-Former-style connector emits a fixed number of vectors — 32 in BLIP-2 — no matter how many patches came in. Cost becomes constant, and detail becomes whatever survives that bottleneck. So the honest formula is: tokens ≈ patches / pooling_factor, unless the connector resamples to a fixed length. ## Area, not side length The consequence people miss is quadratic growth. Token count is proportional to width x height. Doubling the longest side does not double the cost, it quadruples it. Going from a 512px thumbnail to a 2048px scan is a 16x increase. This is why "just send the original" is an expensive default and why every serious pipeline downscales to the smallest resolution at which the target detail is still resolvable. It is also why fine text is the classic failure: if a character is smaller than a patch, the encoder never had a chance to represent it distinctly, and no amount of prompting recovers it. The fix is resolution or tiling, not wording. ## Tiling multiplies the base cost High-resolution handling splits a large image into several encoder-sized tiles and usually adds one downscaled view of the whole image for global context. Each tile costs a full patch grid. A handful of tiles plus a global view can easily land in the thousands of tokens for one picture, which is exactly how a single image out-costs a long text prompt. Implementations cap the tile count for this reason. ## What to do about it - Downscale deliberately: pick the resolution from the smallest feature you must read, not from what the camera produced. - Crop before you send. A crop around the region of interest is strictly cheaper and usually more accurate than a full frame at higher resolution. - Prefer models or settings that pool patches when detail is not the point (a scene-level classification does not need 4096 tokens). - Budget images explicitly. If a conversation carries several images across turns, their tokens stay in the history and are re-charged on every subsequent request unless you drop or summarize them. - Remember latency, not just money: more image tokens means a longer prefill, which shows up directly as time-to-first-token. ## The one-line answer An image becomes a grid of patches, roughly one token each, and the grid grows with area — so a single high-resolution photo is thousands of tokens while a page of prose is several hundred.
- If the model gets a count wrong on a dense photo, is sending a higher-resolution copy always the fix?Not always. Higher resolution helps when the objects were smaller than a patch, but it multiplies tokens quadratically and can push the image sequence long enough that the model attends poorly across it. Cropping to the region of interest is usually the better trade: it raises effective resolution on what matters while lowering total tokens. Frontier reasoning models now do this themselves by zooming into regions during reasoning.
- How does patch merging or pooling change what the model can still see?Merging a 2x2 block into one vector cuts image tokens fourfold but averages away sub-patch structure — thin strokes, small digits, hairline table rules. It is a good trade for scene-level understanding and a bad one for dense documents or charts. This is the same space-versus-detail dial as choosing a lower input resolution, applied after the encoder rather than before it.
- Why do images in a multi-turn conversation cost more than their one-time price?Because they stay in the conversation history. Every subsequent request re-sends the whole transcript, so a 4000-token image is re-charged on each turn and keeps occupying the context window. Long image-heavy sessions therefore grow superlinearly in cost. Mitigations are dropping stale images, replacing them with a text description of what was extracted, or keeping only the most recent frames.
Think of the encoder as reading the image through graph paper: it sees one square at a time, and a bigger sheet at the same square size simply means far more squares to read.
saying these in an interview costs you the question
- Assuming an image costs a flat handful of tokens
- Thinking token count scales with file size in bytes
- Believing doubling image width only doubles the cost
- Saying resolution has no effect on token spend
- Assuming images are re-encoded for free on later turns