skip to content

What is thinking with images, and when does it beat pre-cropping a screenshot yourself?

level: seniorimportance: should knowfreq 34%

answer

  1. The image becomes something to interrogate
  2. Zoom spends a fresh token budget on less area
  3. Crop when you know where to look
  4. Pass count varies, so cost is a distribution
  5. Coarse crop plus model zoom is the hybrid

basics

~20 s

Thinking with images means the model crops, zooms and re-inspects regions of an image during its own reasoning instead of answering from one downscaled glance. It beats manual pre-cropping when you do not know in advance where the answer lives.

solid answer

~50 s

Frontier reasoning models can now treat an image as something to interrogate rather than to read once. Mid-reasoning, the model issues operations on it — crop this region, zoom, rotate — and re-inspects the result, which is why it can recover an 11px error string from a 4K screenshot that a single full-frame pass would have lost. As of mid-2026 this is the standard mitigation for small text, counting and dense-chart failures, displacing hand-rolled tiling pipelines. The trade is control. Each inspection adds image tokens and a reasoning round trip, the number of passes varies run to run, and you cannot predict cost or latency per request as tightly. Manual cropping stays better when the region of interest is known and fixed — a dialog that always appears in the same place — because a deterministic crop is cheaper, faster and reproducible. Use the model's own zooming when the search is open-ended.

go deeper

for a junior

Know that some models can zoom into part of an image while reasoning instead of answering from one whole-image look, which helps with tiny text.

for a middle

Explain why zooming helps — a crop gets its own token budget, so a small region is seen at far higher effective resolution — and that each pass adds cost and latency.

for a senior

Argue the trade against deterministic cropping: fixed layouts and tight budgets favour your own crop, open-ended search favours the model, and most pipelines hybridize the two.

for a principal

Own the variance this introduces into cost and latency SLOs, the observability needed to debug where the model looked, and where the platform's effort controls cap exploration.

## What the technique is The naive way to answer a question about an image is one pass: encode the whole frame at whatever resolution the detail setting allows, then generate an answer. Everything below the effective resolution is unavailable, and everything the model failed to attend to on that single pass is gone. Thinking with images breaks that. During its reasoning, the model acts on the image — cropping to a region, zooming in, rotating — and then looks again at the result, folding what it finds into the ongoing reasoning trace. Functionally it is tool use where the tool operates on the pixels. A 4K QA screenshot no longer has to be understood all at once: the model localizes the error dialog, zooms into it, reads the 11px text at usable resolution, and answers from that. As of mid-2026 this is live in frontier reasoning models and in a growing set of open research systems, and it has become the expected answer to "how do you handle small text and counting failures?" — replacing the older answer of building a tiling pipeline yourself. ## Why it works The underlying constraint has not changed: an image gets a finite token budget, and detail below what that budget represents is lost. Zooming spends a fresh budget on a smaller area, so the same tokens now describe far fewer pixels each. A crop that covers 5% of the frame gives that region roughly twenty times the effective resolution of the full-frame pass. It also helps for reasons beyond resolution. Counting a crowded shelf becomes counting several sparse regions in sequence, which is a much easier task. Reading a cluttered dashboard becomes reading one chart at a time. Attention no longer has to be split across a whole busy frame. ## What it costs **Tokens.** Every re-inspection is another image in the context. A run that zooms four times costs roughly five images, not one. On a high-volume pipeline that difference dominates the bill. **Latency.** Each pass is a reasoning round trip. Response time grows with the number of inspections, and the number is not fixed. **Determinism.** Two runs on the same screenshot may inspect different regions and take different numbers of passes. Cost per request becomes a distribution rather than a number, which complicates budgeting and makes latency SLOs harder to hold. **Observability.** You are relying on the model's choice of where to look. When it answers wrongly, the useful diagnosis is which regions it inspected — so capture that from the reasoning trace where the provider exposes it, or you are debugging blind. ## When to crop yourself instead Deterministic pre-cropping wins whenever you already know where to look: - **Fixed layouts.** In a mobile-app QA pipeline the error dialog, the toast area and the nav bar are at known coordinates. Cropping to them is free, exact and reproducible, and it removes a whole class of "the model looked at the wrong thing" failures. - **Tight cost or latency budgets.** One crop at a known resolution is a fixed price. Agentic zooming is not. - **High volume, narrow task.** Reading one field from a million screenshots does not need open-ended visual search; it needs the same crop a million times. - **Auditability.** When a decision must be reproducible after the fact, a deterministic crop is far easier to defend than "the model chose to look here". Use the model's own inspection when the search is genuinely open-ended: an unfamiliar interface, a failure you cannot localize in advance, a document-like screenshot whose structure varies, or a question that requires comparing several distant regions whose positions are not known. ## The hybrid, which is usually the answer Most production designs combine them. Crop deterministically to a coarse region you are certain about — the app window, the content pane, the chart area — to remove irrelevant pixels and cut cost, then let the model zoom within it for whatever fine detail the question needs. You get the cheap, reliable reduction from your own knowledge of the layout and the adaptive search from the model's. A second useful pattern is capping the exploration. Where the platform exposes an effort or reasoning-budget control, a low setting keeps the pass count down for easy screenshots and reserves deep inspection for the ones that need it. Then measure: track accuracy and cost per task type with and without model-driven zooming, because on coarse questions the extra passes buy nothing and you are paying for exploration you did not need.

  • How does this change how you budget a screenshot-heavy pipeline?
    Per-request image cost stops being a single number and becomes a distribution, because the model decides how many inspections to make. Budget on a measured p50 and p95 rather than on one image's price, cap exploration with whatever effort or reasoning-budget control the platform offers, and pre-crop deterministically wherever the layout is known so the variable part starts from a smaller frame.
  • The model zoomed into the wrong region and answered incorrectly. How do you debug it?
    Capture which regions it inspected from the reasoning trace where the provider exposes them; without that you cannot distinguish a localization failure from a reading failure. If it consistently mislocalizes, the usual fixes are to pre-crop to a coarser region you are confident about, or to state in the prompt where the relevant content tends to appear so the search starts in the right place.
  • Does model-driven zooming remove the need to send images at high detail?
    Not entirely. The first pass still has to be good enough to localize the region worth zooming into — if the whole frame arrives as a thumbnail, the model may not see that there is an error dialog at all. In practice you send a reasonable full-frame resolution for localization and let the zoom passes supply the fine detail, rather than paying maximum detail on the whole image.

saying these in an interview costs you the question

  • Assuming per-request image cost is fixed when zooming is enabled
  • Using model-driven zoom where a fixed crop would do
  • Believing zooming increases the image's real resolution
  • Skipping capture of which regions were inspected
  • Thinking it eliminates counting and small-text errors entirely

context