How do CaMeL-style information-flow controls go beyond the Dual-LLM pattern?
answer
- a handle you never read is not a policy
- tags that travel with the value
- provenance plus permitted readership
- the check happens where the action happens
- guarantees on a subset, not on everything
basics
~20 sDual-LLM keeps untrusted text away from the privileged model by passing symbolic references, but nothing checks what the privileged program then does with those values. CaMeL and FIDES attach provenance and permission labels to every value and enforce policy at the point of action.
solid answer
~50 sDual-LLM, the older pattern, splits the system in two: a privileged model that sees only trusted input and orchestrates, and a quarantined model that handles untrusted content and returns results the privileged side references symbolically rather than reading. That removes the direct path from attacker text to the orchestrator, but it stops there — it says nothing about whether a value derived from untrusted content may be used as an email recipient or written to a public repository. CaMeL (Google DeepMind) generalises the idea: the privileged model emits code in a restricted interpreter, the quarantined model only parses untrusted data into typed values, and the interpreter tracks a capability on each value — where it came from and who is allowed to see it — enforcing a policy when the value reaches a tool. FIDES (Microsoft Research) does the same with explicit confidentiality and integrity labels and deterministic information-flow rules. The result is a checkable guarantee on a subset of tasks rather than a discipline you hope holds.
code
python · 23 linesclass Value:
def __init__(self, data, readers, source):
self.data = data
self.readers = frozenset(readers) # capability: who may receive this
self.source = source # provenance
def derive(self, data):
# derived values inherit the capability, so taint propagates
return Value(data, self.readers, self.source)
def send_email(body: Value, to: str):
if to not in body.readers:
raise PermissionError(f"policy: {to} may not read data from {body.source}")
return f"sent to {to}"
payroll = Value("Q3 payroll table", readers={"[email protected]"}, source="hr-drive")
print(send_email(payroll, "[email protected]"))
summary = payroll.derive("payroll summary")
try:
send_email(summary, "[email protected]")
except PermissionError as err:
print(err)go deeper
Know that Dual-LLM keeps the untrusted text away from the model that holds the tools, and that newer systems additionally tag data with where it came from and who may see it.
Explain the symbolic-handle mechanism and its gap: passing a value the orchestrator never read still permits that value to become a recipient or a destination. Say that capabilities propagate through derived values.
Contrast containment with authorisation, describe the interpreter enforcing policy at the tool call site, and be candid about the cost — policies to author, tasks that must escalate to a human, restricted expressiveness. Say when the cheaper patterns suffice.
Judge whether the engineering investment is justified for your risk profile, and decide who owns the policy set as it drifts. Be able to argue that the guarantee covers a task subset, and what the organisation does about the remainder.
## Dual-LLM, briefly Dual-LLM is the ancestor of this family. You run two model instances with different privilege. The **privileged** one sees the user's request and can call tools; it never sees attacker-influenceable text. The **quarantined** one sees untrusted content but has no tools. When the privileged side needs untrusted data processed, it dispatches to the quarantined side and receives back a handle — `$VAR_1` — that it can pass around without ever reading. Strict data formatting on the quarantined output keeps the handle from smuggling instructions. This buys a real property: injected text has no route into the reasoning that chooses actions. ## Where Dual-LLM stops It is a *containment* pattern, not an *authorisation* one. Once `$VAR_1` exists, nothing in the design constrains its use. If the orchestrating program says `send_email(to=$VAR_2, body=$VAR_1)` and `$VAR_2` was extracted from an attacker's message, the attacker chose the recipient without the privileged model ever reading a hostile word. Three concrete gaps follow: - **No provenance.** The system cannot answer "which values in this call touched untrusted data?" - **No policy.** There is no place to state "payroll data may only be sent to internal HR addresses." - **No propagation.** Values derived from a tainted value inherit nothing; concatenate, summarise or re-parse it and the taint disappears. ## What CaMeL adds CaMeL — from the paper *Defeating Prompt Injections by Design* — restructures the same split around a custom interpreter. 1. The **privileged model (P-LLM)** never sees untrusted content. It emits a program in a restricted, side-effect-free Python-like language expressing the whole task. 2. The **quarantined model (Q-LLM)** is used only as a parser: it turns unstructured untrusted text into typed values. It has no tools and cannot cause an action. 3. A **custom interpreter** runs the program. Every value carries a *capability*: its provenance (which sources it derived from) and its readership (who may receive it). Capabilities propagate through operations, so a summary of a tainted document is itself tainted. 4. At each tool call, a **security policy** written in ordinary code checks the capabilities of the arguments against what the tool is about to do. A violation is refused, or escalated to the user. The crucial move is that both the control flow and the data flow are now expressed in a formal language a checker can reason about, instead of living inside the model's head. The published evaluation reports solving a large majority of AgentDojo tasks — roughly three quarters — while holding a provable security property, which is the honest headline: strong guarantees on a *subset* of tasks, not on everything. ## What FIDES adds FIDES approaches the same target from classical information-flow control. Each piece of data carries **confidentiality** and **integrity** labels; the planner is constrained so that untrusted (low-integrity) data cannot influence security-relevant decisions, and confidential data cannot flow to a low-confidentiality sink. Where a task would violate a label, the system can fall back to a restricted operation — for instance letting a model choose among pre-approved options rather than emitting a free value — so that useful work still happens under the constraint. The deterministic label check is the boundary; the model remains a probabilistic component inside it. ## What you give up This is not free, and a good answer says so. - **Policies must be written.** Someone has to state, per tool, what provenance and readership are acceptable. That is real engineering and it drifts. - **Some tasks become impossible without a human.** "Reply to whoever emailed me" is genuinely a case where untrusted data must determine a recipient. Under a label system that requires an explicit user confirmation or a pre-approved recipient set, not a silent allow. - **Expressiveness limits.** The privileged program runs in a restricted language, so agents that need arbitrary dynamic behaviour do not fit cleanly. - **Two model calls plus interpretation** cost latency and tokens. ## When to reach for it Use the cheaper patterns — an absent capability, a frozen plan, isolated map-reduce — when they contain the risk, because they are far less machinery. Reach for information-flow control when the task *genuinely* requires untrusted data to influence a privileged action, and you need to state precisely which influences are acceptable. That is the situation the earlier patterns cannot express and this family can. As of 2026 these systems are research-grade and adopted piecemeal; the interview-relevant point is that they show the direction of travel — from hoping the model resists, to making unsafe flows unrepresentable.
- A user legitimately asks the agent to reply to whoever emailed them — how does a capability system handle that?It refuses to treat the recipient as a free value derived from untrusted text, and instead narrows the choice. Either the address comes from trusted transport metadata rather than the message body, or the model selects from a pre-approved set, or the flow is escalated for an explicit human confirmation naming the recipient. The point is that the unsafe flow is made explicit rather than silently allowed.
- Why does emitting a program rather than a sequence of tool calls matter for this family of defenses?Because a program is a formal artifact you can analyse. Both the control flow and the data dependencies are visible before anything executes, so an interpreter can propagate labels through operations and check them at each call site. A stream of model-chosen tool calls gives you no such structure — you can only inspect each call in isolation, with no record of which values it derived from.
- CaMeL reports solving roughly three quarters of a benchmark's tasks with its guarantee — is that a good result or a bad one?It is an honest one. The unsolved remainder is largely tasks whose intent genuinely requires untrusted data to drive a privileged action, which no sound policy can allow silently. Read it as a map of where automation must stop or ask a human, not as a defect rate. The alternative — full task coverage with no guarantee — is what the field is trying to move away from.
saying these in an interview costs you the question
- Thinks Dual-LLM alone prevents exfiltration once handles are passed around
- Assumes taint disappears when a tainted value is summarised or reformatted
- Describes information-flow control as a model behaviour rather than interpreter-enforced
- Claims these systems give a total guarantee across all agent tasks
- Treats capability labels as a substitute for removing an unnecessary capability