In an invoice-processing agent, how does a supplier-controlled document field reach a tool call's argument slot?
answer
- no imperative span anywhere in it
- data at every stage of the pipeline
- the extractor copies, it does not vouch
- authoring one field is the whole move
basics
~20 sBy being read as data the whole way. An extraction step lifts the field from the ingested document, a mapping step places it in the argument the operation takes, and no stage asks who wrote it.
solid answer
~50 sThe path is ordinary data flow, which is why it is easy to miss. An invoice arrives in the queue from an outside party. Something extracts its fields — a model reading the document, a deterministic parser, or both. The workflow then maps extracted fields onto the arguments of the disbursement call: this field becomes the reference, that one becomes the account. The value is data at every hop; nobody obeyed anything, and there is no imperative span for a screen to find. The attacker's whole contribution is authorship of a field on a document the pipeline was designed to ingest. What it costs them is guessing which document field lands in which argument slot, and a queue that processes every submission gives cheap attempts. Whether a model or a regex did the lifting, the value's origin is outside the organisation.
go deeper
Be able to describe the path in order: document arrives, fields are extracted, fields become arguments, the call runs. Knowing that no instruction is needed anywhere in that chain is the point to carry away.
Explain why an input screen looking for directive language finds nothing here, and why swapping a model for a parser changes nothing about trust. Expect to be asked which stage you would examine first.
Demonstrate the inventory habit: trace every argument back to whether its value is document-sourced or state-sourced, and rank the document-sourced ones by what the operation does with them.
Own the framing that a queue accepting outside documents has already granted the submission channel; the open question is which of the agent's operations are allowed to take their targets from what arrives in it.
## From a field somebody else wrote to an argument your agent runs on ### The pipeline, stated plainly A procurement and expense workflow with no human in the loop looks roughly like this. A document lands in a queue — an invoice, submitted by an outside party, because that is what the queue is for. An extraction stage turns the document into structured fields. A matching stage lines those fields up against an internal purchase order. A tool call then files a disbursement request, and its arguments are populated from what the earlier stages produced. Somewhere in that mapping, at least one argument gets its value from the document rather than from internal state. That is the whole surface. An attacker who can author that field on a document the queue accepts has supplied an argument to an operation your organisation performs on its own authority. ### Nothing here is an instruction The common mental model of attacks on LLM applications is a directive span hidden in content — text written in the imperative, aimed at the model, competing with the application's own instructions. That model is not what is happening here, and expecting it is why this path gets overlooked. The field in question reads as data to every stage that touches it. It is data to the extractor, which is looking for a value of that kind in that position. It is data to the schema that types it. It is data to the mapping that places it. An input screen hunting for imperative language finds nothing, because there is nothing of that shape to find. The construction does not need the model to change its mind about its task; it needs the workflow to do exactly what it was built to do. ### Model-mediated or deterministic — same trust property Two variants show up in real systems. - **A model transcribes the field.** The assistant reads the document and emits the value into a structured output that becomes the call's arguments. - **A deterministic extractor copies it.** A parser locates the field by position or label and hands it on with no model in the loop for that hop. People instinctively treat the second as safer because no model is exercising judgment. For this class it makes no difference. The property that matters is where the value originated, and in both cases it originated with whoever authored the document. A model in the path adds fuzziness about *which* field gets picked up; it does not add or remove the trust problem. Conversely, removing the model does not remove it either. ### Which fields are worth authoring At the class level, the interesting fields are the ones a pipeline forwards *verbatim* into an operation, rather than the ones it summarises, aggregates or merely displays. A value that only ever reaches a human-readable summary has no effect to influence. A value that reaches an argument on an operation with consequence — what the money goes to, which record is amended, which scope the action covers — is the one that turns text into effect. Free-text fields tend to be the least interesting: they are usually rendered, not acted on. Identifier-shaped fields tend to be the most interesting, precisely because identifiers are what operations act through. ### What it costs the attacker Three things, none of them expensive in a workflow of this shape: 1. **A submission channel.** Already granted — the queue accepts documents from outside parties by design. 2. **A working guess at the mapping.** Which field becomes which argument is not usually published, but it is stable, and a queue that processes every submission gives repeated attempts with observable outcomes. 3. **Plausibility.** The value must be one the downstream stages accept without incident, which for identifier-shaped fields means correctly formed. Notably absent from that list: any need to defeat a model's refusal behaviour, or to place text where a model will read it as directive. ### Where the path closes The surface exists only where an argument's value is sourced from the document. Where the workflow derives that argument from internal state instead — the purchase order's own record of who gets paid — the document has nothing to contribute to it, and authoring the field changes nothing but the paperwork. Mapping which arguments are document-sourced and which are state-sourced is the actual analysis; everything else is commentary. ### Direction of the claim If you observe an attacker-authored value on a real call, that proves the mapping forwarded the field. It does not prove the model was hijacked, that any instruction was obeyed, or that a screening layer failed — there may have been nothing for one to catch.
- Does it matter whether a model or a regex lifted the field out of the document?Not for trust. The property that decides the outcome is where the value came from, and in both cases that is the outside party who authored the document. A model in the path makes which field gets picked less predictable; it neither creates nor removes the exposure. Teams that removed the model and declared the path safe have mistaken determinism for provenance.
- How would you map which of an agent's arguments are exposed this way?Trace each argument backwards to where its value is produced. Arguments sourced from internal state — a purchase order record, the identity that opened the request — are not exposed by this path. Arguments sourced from ingested content are, whichever component did the lifting. That inventory is the finding; the rest is judging which of them carry consequence.
- Why do free-text document fields tend to be less interesting here than identifier-shaped ones?Free text is usually rendered or summarised rather than acted on, so influencing it moves nothing. Identifiers are what operations act through — they select the record, the account, the scope. An argument that selects what an effect lands on is worth more than one that decorates it.
saying these in an interview costs you the question
- Assumes an attack must contain instructions aimed at the model
- Believes a deterministic parser makes ingested values trustworthy
- Thinks an input screen would catch a plain data field
- Cannot say which arguments come from documents and which from internal state
- Confuses the value's format with the value's origin