skip to content

Instructions in a Picture

A model reads an image at full fidelity while the person who pasted it read it at a glance, and a screenshot carries no provenance. Interviewers raise it because pasted images look like trusted input.

on this pageshow

explore

questions

3

In an assistant that reads pasted screenshots, why is text inside the image untrusted input?

level: juniorimportance: must knowfreq 68%

answer

  1. who sent it versus who wrote it
  2. a paste is a delivery act
  3. the model reads the whole frame
  4. a glance is not a read

basics

~20 s

Pasting an image vouches for why the user wanted it, not for who wrote the words in it. A multimodal model reads every legible sentence in the frame, so a planted directive reaches the context as content the model may follow.

solid answer

~50 s

Two different things get confused here. The turn is trusted because an authenticated user sent it; the words inside the picture were authored by whoever produced the picture, and the user only pasted it. A multimodal model reads the frame and produces a reading of everything legible in it, and nothing in an image marks which region the user meant. So a directive placed in the picture arrives in the model's context as readable content, competing with the application's own instructions the way any injected span does. The channel also carries unusually little provenance: nothing fetched it, no source identifier travels with it, and an input screen matching on characters sees an image rather than a sentence. The only thing between the picture and the model is the pasting user's glance, and a glance is not a read.

go deeper

for a junior

Be ready to say plainly that a pasted image is untrusted input: the account is trusted, the words inside the picture are not, and a multimodal model reads them as content it can follow.

for a middle

Explain the mechanics: the model produces a reading of the whole frame with no marker for the region the user meant, and a pasted picture carries far less provenance than a fetched page or a retrieved chunk.

for a senior

Show triage discipline. State what an observed obedience proves, which is that the span reached the context, and what a refusal does not prove, which is that anything inspected the image at all.

for a principal

Own the classification call: whether attacker-authored content arriving inside a trusted user's own turn is counted as an injection class at all, because that choice decides whether the pattern ever appears in what anyone reviews later.

## The two claims a pasted image makes, and the one it does not Picture an assistant that holds one person's mailbox, calendar and contacts. The user is sent something, screenshots it, pastes the screenshot into the conversation and types "deal with this". The application now has one turn with two parts: a short instruction the user wrote, and an image the user did not write. Almost everything that goes wrong here follows from collapsing those two into "the user's message". ### Trust attaches to the sender, not the author The turn arrives from an authenticated session and the application's message structure marks it as coming from the user. That structure records which participant transmitted the content. It makes no claim whatsoever about who authored the bytes inside it. Forwarding is the everyday version of this: a colleague you trust can hand you a document written by anyone at all, and your trust in the colleague says nothing about the document. A paste is a delivery act. ### A multimodal model reads the frame, not the point Given a picture, the model produces a reading of what is legible in it: the message the user cared about, the interface chrome around it, a footer strip, a caption, a quoted tail, anything resolvable at the fidelity the pipeline handed over. There is no marker inside an image saying "this region is the part the user meant". A sentence written in the imperative and addressed to an assistant therefore sits in the model's context as readable content, alongside the application's own instructions, and the model's preference for its own instructions is a trained tendency rather than an enforced boundary. ### This channel carries less provenance than the ones people worry about Compare how untrusted text usually reaches a model: | how the text arrived | what travels with it | | --- | --- | | a retrieved chunk | a source identifier and index metadata | | a fetched page | a URL, and a fetch some code initiated | | an uploaded file | an upload event, and whatever runs on uploads | | a pasted screenshot | only the fact that this account pasted it | Nothing fetched the screenshot, nothing indexed it, and a screen that matches on characters sees an image rather than a sentence. What stands between the picture and the model is the user's own glance before they pasted, and looking at an image is not the same act as reading the text in it: people paste an image for one thing in it and their attention stays on that thing. ### Why this is an injection question and not a content question Injection aims at the application's own instructions using data the application feeds the model. Jailbreaking aims at the model's trained refusal. The picture in this scenario does not need the model to say anything it would otherwise refuse; it needs the assistant to do something different with a mailbox it already has permission to use. By channel the span arrives inside the user's own turn, which superficially looks direct; by authorship it is second-order, because the span was written by whoever built the image and the trusted user is the delivery mechanism. In a report, what matters is the target and the authorship, not which turn carried it. ### What obedience proves, and what it does not If the assistant follows the planted sentence, that proves the span reached the context and was read as directive. It does not prove that any stage was bypassed, that no screening ran, or that the application's instructions were removed. Conversely, if the model declines, that is the answering model refusing to produce something; it is not evidence that a screening layer blocked the image. A refusal and a block are different events with different shapes, and telling them apart is a measurement rather than an assumption. ### Where the construction stops working The crop the user takes may exclude the region entirely. The user's reason for pasting may be the image's text itself, in which case they read it. The deployment may never pass the picture to the model. The payoff may need a capability the assistant does not hold. Each is a separate failure mode with its own rate, which is why a single successful paste is one observation against one configuration rather than a reliable finding. ### What a finding should actually say Name what carried the span (a picture inside the user's own turn), what it was aimed at (the application's instructions), what was in the way (only the pasting user's attention), and what was observed, how many times, against which configuration. Anything stronger overstates a probabilistic result.

  • Would you call this direct or indirect prompt injection?
    By channel it arrives in the user's own turn, which looks direct. By authorship it is second-order: the span was written by whoever produced the image, and the trusted user is only the delivery mechanism. What matters in a report is that the target is the application's own instructions and the content is attacker-authored, so treat it as indirect regardless of which turn carried it.
  • The assistant declined to act on the picture. Does that mean a screening layer blocked it?
    No. A refusal is the answering model declining to produce something; a block by a screening layer is a separate event with a different shape, often different latency and different text, sometimes with no partial output at all. One refusal on one paste says nothing about whether anything inspected the image.
  • Why does the assistant's message-role structure not help here?
    Roles record which participant a turn came from, not who authored the content inside it. A user role on a turn containing a pasted picture asserts that this account sent this content; it makes no claim about the picture's origin. Any trust an application places in a role is trust in the sender alone.

A courier you trust hands you an envelope. You trust the courier; you have no idea who wrote what is inside it.

saying these in an interview costs you the question

  • Says the user vetted the image by choosing to paste it
  • Assumes everything in a user turn was written by the user
  • Thinks a picture can carry content but not instructions
  • Treats a model refusal as proof a screen blocked the image
  • Claims message roles establish who authored the content

context

open as a page

Why does directive text placed in a pasted screenshot survive the pasting user's glance?

level: middleimportance: should knowfreq 52%

basics

~20 s

Because salience and legibility are separate thresholds. The user looks at the image for the one thing they pasted it for, and their attention stops there, while the model reads everything in the frame that is resolvable at all.

open as a page

A pasted screenshot's instruction is obeyed, but the assistant's draft still needs a confirm click — what does that constrain about which argument values survive?

level: seniorimportance: should knowfreq 41%

basics

~20 s

The click reviews a rendered surface, not the call. Values the draft shows have to look like what the user asked for, so a value the user never chose survives only where the rendering rolls it up, truncates it or omits it.

open as a page