skip to content

Why does directive text placed in a pasted screenshot survive the pasting user's glance?

level: middleimportance: should knowfreq 52%

answer

  1. two thresholds, not one
  2. seeing is not reading
  3. they pasted it for one thing
  4. unremarkable to the eye, resolvable to the model
  5. placement is cheap, faintness is expensive

basics

~20 s

Because salience and legibility are separate thresholds. The user looks at the image for the one thing they pasted it for, and their attention stops there, while the model reads everything in the frame that is resolvable at all.

solid answer

~40 s

Human reading of a picture is task-directed: someone pastes a screenshot because of a specific message or number in it, and that region takes their attention while the rest is categorised at a glance and skipped. Regions that look like furniture, such as a footer strip, a caption, a quoted tail or an interface panel, are skipped hardest, because in every other image they are noise. A model does none of that; it reads whatever is resolvable in the frame it receives. So the person placing the text works two independent thresholds at once, and the cheap levers are placement and framing, which cost nothing in legibility. Shrinking the type or fading the contrast is the expensive lever: every step toward being unremarkable to an eye moves toward being unresolvable to the model.

go deeper

for a junior

Know the core distinction: someone can look straight at an image and never read a sentence in it, because looking at a picture and reading its text are different acts.

for a middle

Explain both thresholds and the asymmetry between the levers. Placement and looking like furniture cost nothing in legibility; shrinking and fading spend legibility to buy inattention.

for a senior

Be able to price the construction: state which lever it depends on, which parts of the delivery path are outside anyone's control, and therefore what a single successful paste is actually evidence of.

for a principal

Own the reporting standard for carrier-dependent findings, so that a result leaning on faint type and a result leaning on placement are not written up as the same claim.

## Two thresholds that have almost nothing to do with each other **Legibility** is whether marks are resolvable as characters by whatever is doing the reading. **Salience** is whether a region recruits a person's attention in the first place. They are independent, and the whole of this construction lives in the gap between them. The common wrong answer is "a user would obviously notice text in the image". Notice is a salience claim, and it is asked of a person whose salience budget was already spent before they looked. ### How a person actually reads a picture they pasted Someone screenshots a message and pastes it into an assistant because of one thing: the request in it, the error, the amount, the date. Their eyes go to that region and the rest of the frame is sampled, not read. Reading is a decision, and it is made per region, at a glance, on the basis of what the region looks like. Blocks that resemble furniture, such as a footer, a signature block, a caption strip, a quoted-reply tail or a panel of interface chrome, get categorised as noise and skipped, because in nearly every other picture that is exactly what they are. None of this is carelessness. A person who read every word of every screenshot they forwarded would get nothing else done in a day. ### What the model does instead A multimodal model has no task-directed eye and does not skim. The pipeline hands it a frame and it produces a reading covering whatever is resolvable at the fidelity it received. The user's purpose is invisible to it, so the region the user cared about and the region they never looked at arrive on identical footing. ### The two-sided budget Someone placing directive text is optimising against both thresholds at once, and the levers do not behave the same way: | lever | effect on the eye | effect on the model | | --- | --- | --- | | placement away from the region of interest | much less salient | none | | framing it as furniture (footer, caption, quoted tail) | much less likely to be read | none | | reducing size | less salient | less resolvable | | reducing contrast | less salient | less resolvable | The first two are cheap and the last two are expensive. That asymmetry is the answer an interviewer is scoring: the good construction relies on where the span sits and what it looks like it is, not on making it faint. Faintness spends legibility to buy invisibility and lands in a narrow band, and the band moves with things nobody on the attacking side controls, because the user chooses the crop and the moment they paste, the client renders the picture at some size, and the pipeline scales it before the model sees it. ### What is actually in the way One obstacle only: the pasting user's glance. Nothing fetched the picture, nothing indexed it, and there is no source to check. The trusted user is simultaneously the delivery mechanism and the only thing standing between the span and the model, which is why the construction is priced in attention rather than in bypasses. ### Where it stops working * The user's reason for pasting is the image's text itself, as when they want something proofread or transcribed. Then they read it. * The crop excludes the region. The user chose the crop, and this is the single largest source of variance. * The client renders the image large enough, or the layout places the region prominently enough, that it stops looking like furniture. * The frame handed to the model is too coarse to resolve the marks, so it fails in the other direction. These have different rates and different causes, and they are worth separating in a report. A person who says "it works" without stating which of these they controlled has described a single observation, not a method. ### Pricing it The honest claim is about which lever the construction leans on. A span that depends on placement and framing tolerates rescaling and re-encoding and survives across pastes. A span that depends on small faint type is a coin flip that degrades the moment anything in the delivery path changes, and reporting the second as though it were the first is the mistake that gets a finding thrown back.

  • What does the user's own reason for pasting do for the person who built the image?
    It fixes their attention in advance. Task-directed reading means the region they pasted the picture for absorbs the looking, and everything else is categorised rather than read. The construction does not have to defeat attention; it only has to sit outside where the attention was already going.
  • Why is a picture a better carrier here than the same sentence typed into the turn?
    Typed characters are read by the person composing the turn and are visible to anything matching on characters. A picture's words are neither until a model reads the frame, so the span crosses a stretch of the pipeline as pixels and only becomes a sentence at the point where it can act.
  • Where does this construction stop working?
    When the user's purpose is the text itself, when their crop excludes the region, when the rendering makes the region conspicuous rather than furniture, or when the frame reaching the model is too coarse to resolve the marks. Those are four different failure modes with four different rates, and only the last one is a legibility failure.

A proofreader and a photocopier look at the same page. One decides what is worth reading; the other reproduces everything on it.

saying these in an interview costs you the question

  • Says a user would obviously notice text in an image
  • Treats having seen the image as having read it
  • Assumes smaller and fainter is always better for the attacker
  • Claims one successful paste makes the carrier reliable
  • Confuses low contrast with being invisible to the model

context