skip to content

Why is an internal handoff's notes field still untrusted when our own component wrote it from attacker-supplied text?

level: juniorimportance: should knowfreq 58%

answer

  1. who wrote it, not what it says
  2. the check ran once, at arrival
  3. typed members constrain, prose does not
  4. the label changed, the wording did not

basics

~20 s

Provenance is not integrity. The notes member is free text one component composed out of attacker-supplied input, so the attacker's wording crossed the hop intact. What changed at the boundary was the label on the text, not the text.

solid answer

~50 s

Because the only thing the handoff establishes is *which component emitted the field*, and that is a claim about authorship, not about content. The arriving document was screened once, at the customer's front door; everything past that point is internal by construction. Upstream, a component reads that document and composes a handoff record whose typed members — a severity enum, an asset id, a timestamp — are constrained to values somebody could enumerate in advance. The `notes` member is not, which is exactly why it exists: no one could enumerate prose. So the attacker's directive-shaped wording rides through the component that summarised it and arrives at the next component wearing an internal label. The next component reads it as free text in its own context, and free text in context is directive-eligible. Nothing about passing through a model raises the trust of what came out.

go deeper

for a junior

Be ready to say, in one sentence, that a field's provenance records who emitted it and never what it contains. Know that a check placed at the arrival boundary does not run again on messages between components inside it.

for a middle

An interviewer expects you to explain the record's shape: typed members are constrained to enumerable values, the prose member is not, and that is why the prose member is the one that carries an attacker's wording across the hop.

for a senior

Show that you can locate the defect at the hop where the trust label changes hands rather than at the component that misbehaved, and that you can say what the attacker had to pay to make their text survive the transformation.

for a principal

Own the framing that provenance and integrity are separate assurances, and that a pipeline diagram showing who talks to whom answers neither. Be able to state which claim the organisation currently has evidence for.

## The situation Picture an assistant embedded inside somebody else's product: a vendor supplies a component that reads customer-submitted material, and it hands its findings to the customer's own automation, which then acts. Two organisations, two release cadences, one conversation. Between them sits a handoff record — mostly typed, with one prose member (call it `notes`, `rationale` or `summary`) because some part of the finding could never be reduced to an enumeration. An attacker submits the material. It passes the front-door check at arrival on the customer boundary. The vendor component reads it and writes its finding. Its directive-shaped wording — text that reads as an instruction rather than as a fact — comes back out inside `notes`. The downstream component reads `notes` as prose in its own context and acts. ## Provenance is not integrity The single most common mistake here is to hear *our own component wrote that field* as a statement about the content. It is not. It is a statement about **authorship**: which process serialised the bytes. Integrity is a different claim — that the content is what somebody authorised, unmodified and uninfluenced by a party who should not have had a say. A field composed from attacker-influenced input has attacker influence in it whoever serialised it. The same distinction shows up everywhere in security, but LLM pipelines make it unusually easy to lose, because the intermediate component is a model. A model reads text and writes text; passing prose through it does not launder it. If anything the reverse is true — a component whose job is to *restate* what it read is a component that carries wording forward on purpose. ## Why the check did not cover it The front-door check ran once, on the artefact that arrived, at the boundary where trusted and untrusted were defined. Its scope was the arrival channel. It was never applied to the handoff, because by the pipeline's own topology the handoff is not an arrival — it is an internal message between two components that are both inside the boundary. The check did not fail; it was not present at that hop. Naming that is the whole of the junior answer: the attacker did not defeat the screen, they arranged for their text to be re-emitted downstream of it. ## Typed members versus the prose member It is worth being precise about which part of the record the attacker actually influences: | member | what constrains it | what an attacker can put there | | --- | --- | --- | | `severity` | an enumeration | one of the allowed values | | `asset_id` | an integer key | a value in range | | `notes` | nothing — it is prose | anything the upstream component will restate | The typed members are constrained because their values were enumerable when the schema was written. The prose member exists precisely because its values were not. That asymmetry is not an oversight; it is what makes a free-text member useful, and it is why the free-text member is the interesting one. ## What the reading component does with it When the downstream component builds its own prompt, `notes` becomes ordinary text sitting in a context window. A model has no mechanism that distinguishes *text my operator wrote* from *text that arrived in a field my operator trusts*; instruction preference is a trained tendency, and an internal-looking field is not marked. So the reading component is free to treat the span as something addressed to it. ## What it cost the attacker, and where it stops This is not free. The attacker has to get content into the upstream component's input at all, and their wording has to survive whatever transformation that component applies — a component that emits only enumerated members carries nothing forward, and a component that renders its output to a human rather than feeding it to a second model has a person, not a parser, at the far end. The construction is also blind: the attacker usually cannot see the handoff record, so they are writing for a transformation they cannot observe. ## Saying it in an interview The sentence to land is: *the label on that field records who wrote it, not what it says, and the attacker's text is what it says.* Everything else follows.

  • The upstream component summarises rather than copies. Doesn't paraphrasing destroy the attacker's wording?
    Not reliably. A summariser's job is to carry forward what is salient, and directive-shaped content is salient — it reads like the most important thing in the document. The attacker is also writing for the summary rather than for the original, so the survival of the span through restatement is the property they are selecting for, not an accident.
  • Which members of the handoff record does the attacker actually get to influence?
    Realistically only the unconstrained one. Enumerated and typed members admit a value from a fixed set, so at most the attacker steers which allowed value appears, which is a nudge and not a channel. The prose member admits arbitrary text, so it is the member that carries wording across the hop, and it is the member the next component reads as free text.
  • The downstream component only ever receives messages from our own vendor component. Does that narrow the exposure?
    It narrows who can *send*, not what the message can *contain*. Connectivity and integrity are separate properties: a closed set of senders says nothing about whether one of those senders is restating text an outsider composed. Treating a fixed sender list as an integrity claim is the standard wrong answer in this shape.

A letter you receive, retype and file in your own filing cabinet is now in your handwriting and your cabinet. Neither of those facts makes its contents yours or true.

saying these in an interview costs you the question

  • Says the field is internal, therefore trusted
  • Confuses who wrote a field with what it contains
  • Assumes summarising strips directive-shaped content
  • Believes passing text through a model raises its trust
  • Treats a fixed set of senders as an integrity guarantee

context