skip to content

Image Understanding in Practice

Passing images to a model and getting reliable answers back: image blocks, detail levels, token cost, and the counting or fine-text failures that models now fix by zooming into a region themselves.

on this pageshow

questions

5

How do image detail levels change what a vision model sees and what each image costs?

level: middleimportance: must knowfreq 57%

answer

  1. One knob trades resolution against price
  2. Low is a fixed cheap budget
  3. Cost grows with processed pixels
  4. Halving dimensions quarters the cost
  5. Downscaled small text does not blur, it vanishes

basics

~20 s

A low detail setting collapses the image to a small fixed token budget — cheap and fast, but fine text and precise layout are gone. Higher settings preserve resolution and cost several times more, because image tokens scale with the pixels actually processed.

solid answer

~50 s

Vision APIs expose a resolution knob — OpenAI's `detail` parameter, with `low`, `high`, `original` and `auto` as of mid-2026 — and it is the main lever on both quality and price for an image request. At `low` the image is reduced to a small fixed token cost regardless of its original size: fine for "is this a login screen or a settings screen?", useless for reading an 11px validation message. Higher settings keep the resolution, and cost rises with the number of image patches the model processes, so a 4K screenshot can cost many times a thumbnail. `original` exists for dense, spatially sensitive images — computer-use screenshots, dashboards — where downscaling destroys the answer. The practical discipline is to classify requests: route coarse questions to `low`, reserve full detail for tasks that genuinely need small text or precise positions, and downscale or crop client-side before sending rather than paying for pixels the question does not use.

go deeper

for a junior

Know there is a detail or resolution setting, that low is cheap and coarse, and that small text needs the higher settings to be readable at all.

for a middle

Explain that cost tracks processed pixels — so halving dimensions roughly quarters the price — and that low detail loses fine text outright rather than blurring it.

for a senior

Show you route by request class, crop client-side where the region is known, and can predict the bill for a screenshot-heavy pipeline before it ships.

for a principal

Own the cost-quality frontier: which question classes justify full detail, what the per-image budget is across the product, and how image spend behaves as agent loops lengthen.

## The knob and what it does Every mainstream vision API gives you some control over how much of an image reaches the model. OpenAI's version is the `detail` parameter, which as of mid-2026 accepts `low`, `high`, `original` and `auto`. The semantics generalize even where the parameter name does not: - **low** — the image is reduced to a small, fixed token budget. Cost is constant no matter how large the source image was. The model gets the gist: layout at a glance, dominant colours, obvious shapes, roughly what a squinting human sees from across the room. - **high** — the image is processed at meaningful resolution. Cost scales with the image's dimensions, so a large screenshot costs many multiples of a thumbnail. - **original** — added for dense, spatially sensitive input such as computer-use screenshots, where any resampling loses the answer. Most expensive. - **auto** — the provider picks based on the image. ## The cost model, as a budgeting decision On newer models, image billing is patch-based: the image is divided into small fixed-size squares (32×32 pixels on OpenAI's newer models), the patch count is multiplied by a per-model factor, and that is your image token count. Two consequences matter for budgeting. First, **cost is quadratic in linear resolution**. Doubling both width and height quadruples the patch count. Halving a screenshot's dimensions before sending cuts its cost to roughly a quarter — usually the single largest saving available on a vision request, and one you control client-side without touching any API parameter. Second, **cost is per image, per request**. In an agent loop that re-examines the same screenshot across five turns, you pay for it five times unless the conversation prefix is cached. A pipeline processing 400 screenshots at full detail has a very different bill from one processing them at `low`, and the difference is often the deciding factor in whether the feature ships. ## Choosing per request class, not per project The mistake is picking one setting for the whole system. Classify the questions instead: - **Coarse classification** — "which screen is this?", "did the app crash or render?", "is there a modal open?" — `low` is enough and is dramatically cheaper. - **Reading small text** — a toast message, a truncated button label, an 11px validation error. `low` will not merely be less accurate here; the pixels carrying the text no longer exist after downscaling, and the model will confabulate plausible text rather than report that it cannot read it. This is the failure that surprises people: there is no "blurry" signal in the output. - **Precise spatial questions** — element coordinates, alignment defects, whether a control overlaps another. Needs full detail, and `original` when the downstream consumer acts on positions. - **Charts and dense dashboards** — gridlines and axis labels are small text. Treat these as the fine-text case. ## Cheaper alternatives to raising detail Before paying for `original` on a 4K screenshot, consider whether the question needs the whole image. Cropping to the region of interest client-side — the error dialog, the header bar — often gives *better* accuracy than full detail on the whole frame, at a fraction of the cost, because the region occupies more of the model's effective resolution. If you know the layout is fixed, that crop is deterministic and free. If you do not know where to look, the model can crop and zoom for you during reasoning, at the cost of extra passes. Similarly, resizing to the smallest resolution at which the target text is legible to a human is a reliable rule of thumb: if you cannot read it in the resized image, neither can the model. ## Verifying rather than guessing Token counts for images are predictable from dimensions and the detail setting, and providers document the formula and offer counting endpoints. Measure the cost of a representative request before extrapolating a monthly bill from an assumption; per-model multipliers differ enough that an estimate carried over from another model can be off by several times. ## Version caveat The four-level `detail` vocabulary and patch-based billing reflect mid-2026 practice; earlier models billed images via 512-pixel tiles and offered fewer levels. The durable ideas — a resolution knob, cost scaling with processed pixels, and the silent loss of fine text at low settings — outlive any particular parameter spelling.

  • A model reads an error message that was never legible at the resolution you sent. What happened?
    It confabulated. Downscaling does not leave the model a visible smudge it can flag as unreadable; the text is simply gone, and the model completes with the most plausible string given the surrounding UI context. That is why low detail on fine-text tasks fails silently — you get a confident wrong answer, not an abstention. Verify by checking whether a human can read the same resized image.
  • Would cropping to the error dialog beat sending the full screenshot at maximum detail?
    Usually yes, on both accuracy and cost. A crop makes the region of interest occupy most of the frame, so it survives whatever resampling happens, and you pay for far fewer patches. Full detail on a 4K frame spends most of its budget on pixels irrelevant to the question. Crop when you know where to look; raise detail when you do not.
  • How does image cost behave in an agent loop that re-examines the same screenshot over several turns?
    You pay for the image on every turn it remains in the sent context, unless the provider's prompt caching covers that prefix. Long screenshot-heavy loops are therefore expensive by default. Mitigations are dropping images from history once their finding is recorded, replacing them with a textual summary, or arranging the conversation so the image sits in a cacheable prefix.

saying these in an interview costs you the question

  • Thinking low detail just makes answers slightly worse
  • Assuming image cost depends on file size in bytes
  • Setting one detail level for every request class
  • Believing a downscaled image looks blurry to the model
  • Sending 4K screenshots when a crop would answer the question

context

open as a page

When sending images to a vision model, how do you choose between base64, a URL, and an uploaded file?

level: middleimportance: must knowfreq 62%

basics

~20 s

Inline base64 suits one-off images but inflates the request by roughly a third. A URL keeps requests small when the image is already hosted and reachable by the provider. An uploaded file reference wins when the same image is sent many times across requests.

open as a page

What kinds of image questions do vision models still get wrong, and how do you design around it?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Counting repeated objects, reading dense small text, precise spatial relations, and pulling exact values off unlabelled charts remain unreliable — and the model states wrong answers fluently rather than abstaining. Design around it by cropping, decomposing into enumerable steps, and requiring ranges instead of false precision.

open as a page

How do you ask a vision model to compare two screenshots in one request?

level: juniorimportance: should knowfreq 44%

basics

~20 s

Put both images in the same message, each preceded by a short text label such as "Before:" and "After:", then ask the comparison question. Images are positional, so labelling them in text is what lets the model refer to each one reliably.

open as a page

What is thinking with images, and when does it beat pre-cropping a screenshot yourself?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Thinking with images means the model crops, zooms and re-inspects regions of an image during its own reasoning instead of answering from one downscaled glance. It beats manual pre-cropping when you do not know in advance where the answer lives.

open as a page