skip to content

How do you ask a vision model to compare two screenshots in one request?

level: juniorimportance: should knowfreq 44%

answer

  1. Content blocks are ordered, not named
  2. Position is the only addressing scheme
  3. Label each image in adjacent text
  4. Question goes after the evidence
  5. Batching saves prompt tokens, not image tokens

basics

~20 s

Put both images in the same message, each preceded by a short text label such as "Before:" and "After:", then ask the comparison question. Images are positional, so labelling them in text is what lets the model refer to each one reliably.

solid answer

~50 s

A chat message body is an ordered list of blocks, and image blocks sit in that list alongside text. To compare two screenshots you send **one** user message containing: a text block labelling the first image, the first image, a text block labelling the second, the second image, then the question. The labels matter — without them the model has to refer to images by position ("the second image"), which it does inconsistently, and any answer that mixes the two up is hard to spot. Keep the question last so it applies to everything above it. Cost-wise, both images are charged as images regardless of how you split the request; putting them in one call saves you re-sending the system prompt and instructions, not image tokens. If you need a structured diff, ask for a fixed output shape (e.g. a list of changed regions) rather than free prose.

go deeper

for a junior

Know that images and text are ordered blocks in one message, and that you label each image with a short text line before it so the model can refer to them.

for a middle

Explain why positional reference is unreliable and how labelling plus a fixed output schema makes a comparison checkable rather than merely plausible.

for a senior

Show judgment about context hygiene: which screenshots to retain, when to summarize and drop, and why composites destroy exactly the detail a QA diff depends on.

for a principal

Own the interface contract — what the comparison step must emit for downstream automation, and how many images per request the pipeline is allowed to spend before it needs a different design.

## What a multi-image request actually is Modern chat APIs model a message body as an ordered sequence of content blocks. Most blocks are text; some are images, supplied inline, by URL, or by a reference to a previously uploaded file. Nothing about the format restricts you to one image per message — you can interleave several images with text in a single user turn, and the model sees them in the order you sent them, as part of one prompt. That ordering is the entire addressing scheme. There is no filename, no variable name, no handle the model can quote back at you. If you send two screenshots and ask "what changed in the second one?", the model has to resolve "second" positionally, and its answer gives you no way to tell whether it resolved it correctly. ## Label the images in adjacent text The fix is to name each image in the text block immediately preceding it: - text: "BEFORE — build 4.2.1, checkout screen:" - image: screenshot A - text: "AFTER — build 4.2.2, same screen:" - image: screenshot B - text: "List every visual difference between BEFORE and AFTER. For each, name the UI element and describe the change." Now the model has stable tokens ("BEFORE", "AFTER") to reason with and to quote in its answer, and you can check that its answer is anchored to the right image. This is ordinary prompt hygiene, not a vendor trick, and it survives changes to any provider's API shape. ## Where the question goes Put the instruction last, after the images. A question placed before the images is still visible to the model, but a reader-friendly ordering — context, evidence, task — is easier to extend when you later add a third image, and it keeps the task statement adjacent to the output the model is about to produce. If you have long standing instructions (an output schema, a rubric for what counts as a difference), those belong in the system prompt so they are stable across every screenshot pair. ## Cost and limits Each image is billed as an image. Two images in one request cost roughly what two images cost in two requests; the saving from batching them into one call is that the system prompt, instructions and any shared context are sent once instead of twice. What you get in exchange is the thing you actually wanted: the model can only compare two images if both are in the same context window. Providers do cap how many images a single request may carry and how large each one may be, and long conversations that accumulate many screenshots will consume context quickly. Two practical consequences: drop stale screenshots out of the conversation once they are no longer needed, and do not assume you can dump fifty images into one turn — check the provider's per-request limits before designing around it. ## Asking for output you can act on Free-prose comparisons read well and are hard to consume. If the comparison feeds a pipeline — a QA bot filing bug reports, for example — ask for a fixed shape: a list of objects with an element name, a change description and a severity. Structured output also makes the failure visible: an empty list is an explicit "I found nothing", whereas prose that says "the screens look broadly similar" is unfalsifiable. ## Common mistakes Sending each screenshot in its own request and asking the model to remember the first one only works if both requests are turns in the same conversation, and even then the first image must still be in context. Referring to "image 1" when you never wrote "image 1" anywhere in the prompt is the most common bug, and it fails silently — the model answers confidently about whichever image it guessed. Finally, resist stitching two screenshots into one composite image to "save" a slot: you halve the effective resolution of each, which is exactly the wrong trade when the difference you are hunting is a truncated label or an 11px error string.

  • Why not stitch the two screenshots side by side into one image instead?
    Because a composite has the same pixel budget as a single image, so each half is effectively downscaled. Any difference that lives in fine detail — a truncated button label, an 11px validation message — is the first thing to disappear. Composites are fine when you only care about layout at a glance; they are the wrong choice for a QA diff that hinges on small text.
  • If the conversation already contains ten screenshots, what problems should you expect?
    Context pressure and drift. Every retained image keeps consuming context, and the model's attention across many similar screenshots degrades — it starts conflating them. Drop images out of history once their turn is answered, keep only a textual summary of what was found, and re-send an image only when a later question genuinely needs it.

saying these in an interview costs you the question

  • Assuming the model can name images without text labels
  • Thinking two images in one call cost less than two calls
  • Sending images in separate requests and expecting recall
  • Stitching screenshots into a composite to save a slot
  • Believing image order in the request does not matter

context