How does an isolated extraction pass keep untrusted text out of a tool-calling model?
answer
- reader has no tools, no credentials
- prose stops at the boundary
- typed record crosses, free text does not
- enums and bounds shrink the channel
- one open string field undoes it
basics
~20 sA first model with no tools reads the untrusted document and emits only a fixed, typed record. The privileged model that can call tools sees that record, never the free text — so an injected imperative has no channel into the component holding the capabilities.
solid answer
~50 sThe pattern splits one model call into two components with different privileges. A quarantined reader holds no tools, no credentials and no memory of the user's goal; its whole job is to turn the untrusted document into a schema-constrained record — say nine typed fields for a résumé: years of experience as a bounded integer, degree level as an enum, skills as a capped array of short strings. The orchestrating model that can actually act receives only that validated record. An injected instruction inside the document can now do at most one thing: distort a field value within its declared type. It cannot reach the deciding model as prose, because prose never crosses the boundary. The cost is real — an extra call, lost nuance, and a schema you must design to be complete enough for the task. And the pattern collapses the moment you add a free-text field copied from the source.
code
json · 15 lines{
"type": "object",
"additionalProperties": false,
"required": ["years_experience", "degree_level", "top_skills"],
"properties": {
"years_experience": { "type": "integer", "minimum": 0, "maximum": 60 },
"degree_level": { "enum": ["none", "bachelor", "master", "doctorate"] },
"top_skills": {
"type": "array",
"maxItems": 8,
"items": { "type": "string", "maxLength": 32 }
},
"source_document_id": { "type": "string", "pattern": "^doc_[0-9a-f]{16}$" }
}
}go deeper
Know the shape: one model reads the risky document and returns fixed fields, a second model does the work using only those fields, so the risky text never reaches the part that can act.
Explain why the schema carries the security property — enums, bounds and no additional properties limit how much an attacker can influence — and why validation must also happen in code after decoding.
Show where it applies and where it does not: structured intake yes, open-ended reasoning over prose no. Be able to argue against adding a free-text field and to describe passing source content by reference instead.
Frame it as buying a boundary with product expressiveness, and own the governance consequence: schema changes become security reviews, and some features cannot be built this way at all.
## The idea in one line If attacker-controlled text never reaches the component that holds the capabilities, it cannot instruct that component. Isolation implements the boundary the prompt could not enforce, by separating *reading untrusted content* from *being allowed to act*. ## The two components **The quarantined reader.** It receives the untrusted document and nothing else of consequence: no tools, no credentials, no network, no knowledge of what will be done with the answer. Its output is constrained to a schema — enforced by the decoding layer, then validated again in code. It is fully expendable: assume it is compromised on every call, because that assumption is what makes the design sound. **The privileged orchestrator.** It talks to the user, holds the tools, and receives only the validated record. It never sees the source prose. In the recruiting-screener case, the reader turns a candidate PDF into a fixed record — years of experience (integer, bounded), degree level (enum), skills (array of short strings, capped), location (enum of regions), and so on. The 1-point white text urging a strong-hire rating has nowhere to go: it is not a field, and the fields it could influence are typed. The worst outcome is a résumé that claims more experience than it has, which is résumé fraud — an ordinary product problem with ordinary controls — not remote control of an agent. ## Why the schema is the whole design The security property comes from the constraint, so the constraint has to be genuine: - **Prefer enums and bounded numerics** to strings. An enum of five values carries at most log2(5) bits of attacker influence. - **Cap array lengths and string lengths.** A `maxItems` of 8 and a `maxLength` of 32 turn a smuggling channel into a trickle. - **Forbid additional properties**, and validate in code after decoding rather than trusting the model to have obeyed. - **Treat any free-text field as a hole.** A `notes: string` copied from the source reopens the exact channel the design closed. If a human must eventually read the original prose, pass it by reference — an identifier the orchestrator can hand to a rendering surface — rather than as content in the orchestrator's context. That last technique generalises: values that must flow through the system but need not be *reasoned about* can travel as opaque references, with the privileged component manipulating handles rather than text. ## What it costs **Latency and money.** Two calls instead of one, and the reader must process the full document. **Lost nuance.** Anything the schema did not anticipate is gone. If the task genuinely needs open-ended judgment over the source — summarising a legal opinion, assessing writing quality — the schema cannot be made complete, and the pattern either fails to do the job or degenerates into a free-text field. **Schema design becomes a product surface.** Every new requirement is a schema change, and the temptation to add "just one" open field is constant. The discipline is to route new requirements through typed fields or through references, and to review any proposed free-text field as a security change. **It does not cover everything.** Field values still influence downstream decisions, so a corrupted record can steer an action even without instructing the model. Isolation reduces the *channel*, not the *influence*; it composes with approval on side effects and with limiting what the orchestrator can invoke. ## When to reach for it Use it where the untrusted input is high-volume and the required output is structured: document intake, screening, triage, classification, extraction pipelines. Skip it where the value of the feature *is* open-ended reasoning over untrusted prose — there you accept that the model reads the text and put your weight on capability limits and approval instead. A useful middle path is partial isolation: the reader emits a typed record plus a short, marked quotation whose length is capped and which is spotlit when it reaches the orchestrator. That is weaker, and it should be described honestly as weaker, but it keeps the schema from swallowing the product. ## Interview framing Say what the boundary is (prose stops at the reader), what the reader is allowed to have (nothing), what makes the record safe (types, bounds, no free text), and what it costs (a call, and nuance). Then name the limit: the pattern shrinks the channel to the size of your schema, which is why an unconstrained string field undoes it entirely.
- The product team wants a one-sentence rationale field from the source document. How do you handle it?Treat it as a security change, not a copy change. Options in order of preference: derive it as an enum of reason codes; keep it as a reference the user interface renders from the source, never entering the orchestrator's context; or, if it must be text, cap its length hard, spotlight it on arrival and accept that the orchestrator is now reading attacker-influenceable prose — which means the capability limits behind it carry the weight.
- Does the extraction model need to be smaller or cheaper than the orchestrator?Not necessarily. The property that matters is privilege, not size: no tools, no credentials, no goal context. In practice a cheaper model is often adequate because the task is bounded extraction, and that helps with the extra call's cost — but choosing a small model for security reasons is a category error, since a compromised reader is assumed either way.
- What residual influence does an attacker keep once the schema is enforced?The field values themselves. A document can misstate years of experience or claim a skill, and downstream logic acting on those values is steered accordingly. That is ordinary data fraud with ordinary defences — cross-checks, verification, human review of outliers — and it is categorically different from the model being instructed, which is what the isolation removed.
saying these in an interview costs you the question
- Including a free-text field copied from the untrusted source
- Giving the extraction model tool access for convenience
- Trusting the decoded output without validating it in code
- Calling the pattern a complete fix for prompt injection
- Assuming a smaller extraction model is what makes it safe