How do image detail levels change what a vision model sees and what each image costs?
answer
- One knob trades resolution against price
- Low is a fixed cheap budget
- Cost grows with processed pixels
- Halving dimensions quarters the cost
- Downscaled small text does not blur, it vanishes
basics
~20 sA low detail setting collapses the image to a small fixed token budget — cheap and fast, but fine text and precise layout are gone. Higher settings preserve resolution and cost several times more, because image tokens scale with the pixels actually processed.
solid answer
~50 sVision APIs expose a resolution knob — OpenAI's `detail` parameter, with `low`, `high`, `original` and `auto` as of mid-2026 — and it is the main lever on both quality and price for an image request. At `low` the image is reduced to a small fixed token cost regardless of its original size: fine for "is this a login screen or a settings screen?", useless for reading an 11px validation message. Higher settings keep the resolution, and cost rises with the number of image patches the model processes, so a 4K screenshot can cost many times a thumbnail. `original` exists for dense, spatially sensitive images — computer-use screenshots, dashboards — where downscaling destroys the answer. The practical discipline is to classify requests: route coarse questions to `low`, reserve full detail for tasks that genuinely need small text or precise positions, and downscale or crop client-side before sending rather than paying for pixels the question does not use.
go deeper
Know there is a detail or resolution setting, that low is cheap and coarse, and that small text needs the higher settings to be readable at all.
Explain that cost tracks processed pixels — so halving dimensions roughly quarters the price — and that low detail loses fine text outright rather than blurring it.
Show you route by request class, crop client-side where the region is known, and can predict the bill for a screenshot-heavy pipeline before it ships.
Own the cost-quality frontier: which question classes justify full detail, what the per-image budget is across the product, and how image spend behaves as agent loops lengthen.
## The knob and what it does Every mainstream vision API gives you some control over how much of an image reaches the model. OpenAI's version is the `detail` parameter, which as of mid-2026 accepts `low`, `high`, `original` and `auto`. The semantics generalize even where the parameter name does not: - **low** — the image is reduced to a small, fixed token budget. Cost is constant no matter how large the source image was. The model gets the gist: layout at a glance, dominant colours, obvious shapes, roughly what a squinting human sees from across the room. - **high** — the image is processed at meaningful resolution. Cost scales with the image's dimensions, so a large screenshot costs many multiples of a thumbnail. - **original** — added for dense, spatially sensitive input such as computer-use screenshots, where any resampling loses the answer. Most expensive. - **auto** — the provider picks based on the image. ## The cost model, as a budgeting decision On newer models, image billing is patch-based: the image is divided into small fixed-size squares (32×32 pixels on OpenAI's newer models), the patch count is multiplied by a per-model factor, and that is your image token count. Two consequences matter for budgeting. First, **cost is quadratic in linear resolution**. Doubling both width and height quadruples the patch count. Halving a screenshot's dimensions before sending cuts its cost to roughly a quarter — usually the single largest saving available on a vision request, and one you control client-side without touching any API parameter. Second, **cost is per image, per request**. In an agent loop that re-examines the same screenshot across five turns, you pay for it five times unless the conversation prefix is cached. A pipeline processing 400 screenshots at full detail has a very different bill from one processing them at `low`, and the difference is often the deciding factor in whether the feature ships. ## Choosing per request class, not per project The mistake is picking one setting for the whole system. Classify the questions instead: - **Coarse classification** — "which screen is this?", "did the app crash or render?", "is there a modal open?" — `low` is enough and is dramatically cheaper. - **Reading small text** — a toast message, a truncated button label, an 11px validation error. `low` will not merely be less accurate here; the pixels carrying the text no longer exist after downscaling, and the model will confabulate plausible text rather than report that it cannot read it. This is the failure that surprises people: there is no "blurry" signal in the output. - **Precise spatial questions** — element coordinates, alignment defects, whether a control overlaps another. Needs full detail, and `original` when the downstream consumer acts on positions. - **Charts and dense dashboards** — gridlines and axis labels are small text. Treat these as the fine-text case. ## Cheaper alternatives to raising detail Before paying for `original` on a 4K screenshot, consider whether the question needs the whole image. Cropping to the region of interest client-side — the error dialog, the header bar — often gives *better* accuracy than full detail on the whole frame, at a fraction of the cost, because the region occupies more of the model's effective resolution. If you know the layout is fixed, that crop is deterministic and free. If you do not know where to look, the model can crop and zoom for you during reasoning, at the cost of extra passes. Similarly, resizing to the smallest resolution at which the target text is legible to a human is a reliable rule of thumb: if you cannot read it in the resized image, neither can the model. ## Verifying rather than guessing Token counts for images are predictable from dimensions and the detail setting, and providers document the formula and offer counting endpoints. Measure the cost of a representative request before extrapolating a monthly bill from an assumption; per-model multipliers differ enough that an estimate carried over from another model can be off by several times. ## Version caveat The four-level `detail` vocabulary and patch-based billing reflect mid-2026 practice; earlier models billed images via 512-pixel tiles and offered fewer levels. The durable ideas — a resolution knob, cost scaling with processed pixels, and the silent loss of fine text at low settings — outlive any particular parameter spelling.
- A model reads an error message that was never legible at the resolution you sent. What happened?It confabulated. Downscaling does not leave the model a visible smudge it can flag as unreadable; the text is simply gone, and the model completes with the most plausible string given the surrounding UI context. That is why low detail on fine-text tasks fails silently — you get a confident wrong answer, not an abstention. Verify by checking whether a human can read the same resized image.
- Would cropping to the error dialog beat sending the full screenshot at maximum detail?Usually yes, on both accuracy and cost. A crop makes the region of interest occupy most of the frame, so it survives whatever resampling happens, and you pay for far fewer patches. Full detail on a 4K frame spends most of its budget on pixels irrelevant to the question. Crop when you know where to look; raise detail when you do not.
- How does image cost behave in an agent loop that re-examines the same screenshot over several turns?You pay for the image on every turn it remains in the sent context, unless the provider's prompt caching covers that prefix. Long screenshot-heavy loops are therefore expensive by default. Mitigations are dropping images from history once their finding is recorded, replacing them with a textual summary, or arranging the conversation so the image sits in a cacheable prefix.
saying these in an interview costs you the question
- Thinking low detail just makes answers slightly worse
- Assuming image cost depends on file size in bytes
- Setting one detail level for every request class
- Believing a downscaled image looks blurry to the model
- Sending 4K screenshots when a crop would answer the question