skip to content

What kinds of image questions do vision models still get wrong, and how do you design around it?

level: seniorimportance: must knowfreq 60%

answer

  1. Fluency is not calibrated to the pixels
  2. Counting fails hardest under crowding
  3. Missing text is absent, not blurry
  4. False precision from unlabelled charts
  5. Enumerate, then count the list

basics

~20 s

Counting repeated objects, reading dense small text, precise spatial relations, and pulling exact values off unlabelled charts remain unreliable — and the model states wrong answers fluently rather than abstaining. Design around it by cropping, decomposing into enumerable steps, and requiring ranges instead of false precision.

solid answer

~50 s

Four failure families survive in frontier models. **Counting**: asked how many facings of the blue box are on a shelf, a model will answer "11" with full confidence when the truth is 14, because it pattern-matches a quantity rather than enumerating. **Dense small text**: labels below the effective resolution are not blurred, they are absent, and the model completes plausible text. **Fine spatial relations**: which element is above, overlapping or aligned with which, especially in cluttered UI. **Unlabelled charts**: asked for Q3 revenue from a bar chart with no data labels, the model reports a suspiciously exact figure it interpolated. The common thread is that fluency is not calibrated to visual evidence, so there is no confidence signal to trigger on. Mitigations are structural: crop to the region, make the model enumerate (list each item, then count the list), require ranges or "cannot determine" as valid outputs, cross-check numbers against a non-visual source, and audit a labelled sample so you know your real accuracy per task type.

go deeper

for a junior

Know the standard weak spots — counting many similar objects, tiny text, exact chart values — and that the model answers confidently even when it is wrong.

for a middle

Explain why there is no usable confidence signal, and show the prompt-level fixes: crop to the region, make the model enumerate then count the list, allow an explicit cannot-determine answer.

for a senior

Demonstrate a verification design that never relies on the model flagging itself: cross-source checks, repeated sampling with disagreement routing, and accuracy tracked per task type rather than in aggregate.

for a principal

Own where automation stops — which visual judgments the product is allowed to act on unreviewed, what the human-review threshold costs, and how that boundary is revisited as models improve.

## The failure families **Counting.** This is the most reliably broken capability and the one most often assumed fixed. Ask for the number of product facings on a retail shelf photo and you get a fluent integer that is frequently off by a few — 11 where the truth is 14. The error is not random noise from a hard image; it is systematic under crowding, occlusion and repetition, exactly the conditions that make counting worth automating. Nothing in the output distinguishes a careful count from a guess. **Dense small text.** Covered by the resolution the model actually processed. Below that threshold, characters do not degrade gracefully — they are gone, and the model fills the gap with text that fits the surrounding context. A QA bot reading a toast notification at low resolution will report a message that sounds exactly like your product's copy and is not what was on screen. **Fine spatial relations.** "Is the submit button inside the modal or behind it?", "is this label misaligned with the field?", "which of these two icons is further left?" Models do far better at what is present than at precise relative geometry, and cluttered interfaces are the worst case. Anything downstream that consumes coordinates should treat them as hints, not measurements. **Unlabelled charts.** A quarterly revenue bar chart with gridlines but no data labels forces the model to read values off the axis. The right answer is an approximation with stated uncertainty ("roughly 4.2 to 4.5 million"); the common answer is "4.37 million", a precision the image cannot support. False precision is worse than a wide range because it survives review — a number with two decimal places reads as measured. ## Why there is no confidence signal All four share one root cause for the engineer: the model's linguistic fluency is decoupled from the strength of the visual evidence. It produces an equally well-formed sentence whether it enumerated carefully or pattern-matched. Asking "how confident are you?" produces another fluent sentence, not a calibrated probability. Any design that assumes the system will flag its own uncertainty is unsound. ## Designing around it **Crop before you ask.** If the answer lives in one region, send that region. The region then occupies most of the frame and survives resampling, and you pay for far fewer pixels. Where the layout is fixed — a known dialog position in a QA screenshot — the crop is deterministic. **Force enumeration.** Instead of "how many facings?", ask for a list: one line per item with a short description or an approximate position, then count the list length in your own code. This converts an unverifiable single number into a structure you can inspect, deduplicate and sanity-check. It is slower and costs more tokens, and it is the single most effective counting mitigation available at the prompt level. **Make abstention legal.** If your output schema has no way to say "cannot determine from this image", the model will always produce a value. Add the option explicitly and instruct that it be used. You will get some over-abstention; you will also stop shipping confabulated strings. **Demand ranges for chart reads.** Require a low and high bound plus the basis ("between the 4M and 4.5M gridlines"). A range you can act on beats a point estimate you cannot trust, and the stated basis makes the reasoning auditable. **Cross-check against non-visual truth.** If a number exists in a database, a log or an API, get it from there and use the image only to locate or contextualize it. Reading a value off a picture when the source of truth is queryable is a design smell. **Sample repeatedly where it matters.** For high-stakes counts, ask several times and compare. Agreement across independent samples is a weak but real signal; disagreement is a strong one and should route to human review. **Measure per task type.** Aggregate accuracy hides everything. Track counting, text extraction, spatial and chart-reading tasks separately against a labelled sample, because they fail at very different rates and only the per-type numbers tell you which parts of the product are safe to automate. ## Where the field is moving Frontier reasoning models increasingly mitigate small-text and counting failures themselves by inspecting regions of the image during reasoning rather than answering from one downscaled glance. That shifts the balance but does not eliminate the failure families, and it introduces its own cost and determinism trade-offs. The engineering posture stays the same: assume confident wrong answers are possible on these four task types, and build verification that does not depend on the model noticing its own uncertainty.

  • Why does asking the model "how confident are you in that count?" not help?
    Because the answer is generated the same way the count was — fluently, from context, not from a calibrated estimate of visual evidence. Self-reported confidence on vision tasks correlates poorly with correctness and can be anti-correlated when the image is degraded, since a clean-looking downscaled frame gives no cue that detail was lost. Use structural checks — enumeration, repeated sampling, cross-source verification — instead.
  • A dashboard screenshot has a chart with no data labels. What should the output contract look like?
    A bounded estimate plus its basis: a low and high value, the gridline interval used, and an explicit "cannot determine" option. Reject point estimates with more precision than the gridlines support. If the underlying series is available from a data source, read the picture only to identify which series and period is being asked about, and take the number from the source.
  • When is a counting task safe to automate despite these limits?
    When the count is small and uncrowded, when an approximate answer is genuinely useful, or when a downstream check catches errors cheaply — for example, a count that is reconciled against inventory records, where a mismatch triggers review rather than acting directly. Automate the enumeration and let a system with real ground truth own the number.

saying these in an interview costs you the question

  • Trusting a model's stated confidence on a count
  • Reading exact values off an unlabelled chart
  • Assuming missing small text produces a blurry or hedged answer
  • Reporting one aggregate accuracy across all image task types
  • Acting on model-reported coordinates as if measured

context