skip to content

Under the Agents Rule of Two, how would you restructure a triage agent that reads public issues while holding private-repo access?

level: principalimportance: should knowfreq 36%

answer

  1. at most two of three, per session
  2. untrusted in, secrets, and the power to act
  3. split the run, not the product
  4. the handoff is a schema, not prose
  5. human approval is the weak escape hatch

basics

~20 s

The Rule of Two allows an agent session at most two of: untrusted input, sensitive access, and the ability to change state or communicate externally. A triage agent holding all three should be split into a public-only reader that emits a structured result and a privileged session that never reads issue text.

solid answer

~50 s

Meta's Agents Rule of Two is a pick-two constraint: within one session an agent should hold at most two of [A] processing untrusted input, [B] access to sensitive systems or private data, and [C] the ability to change state or communicate externally — and if it needs all three, the session must be gated on human validation rather than run autonomously. A repo-triage agent that reads public issues (A), can read private repositories (B) and can open pull requests (C) has all three, which is exactly the shape behind the reported incident where text in a public issue drove an agent to leak private source into a public PR. The restructure is to split sessions: an untrusted-triage session reads issues with public-only credentials and emits a typed record — issue id, category, affected component — and a second, privileged session consumes only that record, never the raw text. Neither session holds three legs, and the boundary is credentials and schema, not the prompt.

go deeper

for a junior

Know the three properties and that an agent session should hold at most two of them. Give one example, such as an agent that reads public tickets and can also push code.

for a middle

Explain why the session is the unit of analysis and describe a concrete split with a typed handoff between an untrusted reader and a privileged actor. Note that the third leg covers any state change, not only data leaving.

for a senior

Redesign a real workflow under the constraint and defend the handoff schema as the boundary. Be explicit that human approval degrades under volume, and place it only on rare, legible, high-consequence actions.

for a principal

Own it as an org-wide intake gate: per-session capability declarations, credentials issued per session, and a blocking review finding for three-leg designs with time-boxed exceptions. Be honest that it is a heuristic and pair it with stronger controls where a guarantee is required.

## The rule The Agents Rule of Two states that an agent session should satisfy no more than two of three properties: - **[A] It processes untrusted input** — content an outsider can influence. - **[B] It has access to sensitive systems or private data.** - **[C] It can change state or communicate externally** — write, deploy, pay, post, send. If a workflow genuinely needs all three, the rule says the session does not run autonomously: it requires human validation, or it is restarted with a narrower scope. It is deliberately close in spirit to the lethal trifecta, and differs in one important way: the third leg is **any** state change or external communication, not just exfiltration. That widens the net to cover an agent with no secrets at all that can still delete a table or post publicly on your behalf. ## Why the unit is the session Products do many things; sessions are where capability actually accumulates. A code assistant may, across its lifetime, read public issues, read private code, and open pull requests — that is fine. What must not happen is one *run*, with one context and one credential set, holding all three simultaneously, because prompt injection operates inside a run. Making the session the unit is what turns the rule from a slogan into something you can enforce: sessions are where you attach credentials, tool catalogues and egress policy. ## The triage case The reported MCP-connected repository incident is the canonical illustration. An agent had access to a developer's private repositories and was pointed at a public repository's issues. An issue authored by an outsider contained text addressed to the agent; the agent followed it and published private source into a public pull request. Every leg was present: untrusted issue text (A), private repo access (B), PR creation (C). No individual component misbehaved — the agent did exactly what the composed system permitted. ## The restructure Split the workflow so no session holds three. **Session 1 — triage (A + nothing else that matters).** Credentials scoped to public data only. It reads issue text, and its sole output is a typed record: `{issue_id, category, suspected_component, severity, needs_repro}`. No free text passes forward, or if a human-readable field is required, it is treated as untrusted downstream and never used as an instruction or an argument to an action. This session has untrusted input but neither sensitive access nor the ability to change anything. **Session 2 — action (B + C).** It receives the structured record. It may open a branch, run tests, and draft a change against private code. It never ingests the raw issue body, so there is no channel from the outsider into the privileged reasoning. It holds sensitive access and write capability — two legs — and its untrusted input has been reduced to a validated enum plus identifiers. The handoff schema is the boundary. If it degrades into "pass the issue text along so session 2 has context," the split is decorative and you are back to three legs. ## Approval is the weakest of the three escapes The rule permits all three legs when a human validates the action, and this is where senior judgement matters. Human approval degrades: reviewers who see hundreds of near-identical prompts approve reflexively, and the interesting attack is the one that looks routine. Vendor write-ups on agent containment call this out directly as approval fatigue. So treat approval as a control for **rare, high-consequence, legible** actions — a production deploy, a payment, a public publish — and never as the routine reason your architecture may keep all three legs. If a workflow needs approval on every step, the design is wrong, not the reviewer. ## Rolling it out At organisational scale the rule's value is that it is checkable by someone who is not an ML engineer. Practical adoption looks like: a required declaration per agent session of which legs it holds; credentials issued per session rather than per product; egress and write capability granted only to sessions that do not read untrusted content; and a review gate that treats a three-leg declaration as a blocking finding needing an explicit, time-boxed exception. Expect pushback — splitting sessions costs latency, tokens and some capability, because the privileged session genuinely knows less. That is the trade you are choosing: a narrower agent that cannot be talked into a breach beats a capable one whose safety rests on the model resisting every attempt. ## Its limits The rule is a heuristic, not a proof. It says nothing about how sensitive the sensitive access is, it can be gamed by declaring a leg absent when a subtle channel exists, and two legs can still be bad — untrusted input plus write access with no secrets at all is enough to corrupt a shared workspace. Use it as a fast, universally-applicable gate, and use information-flow controls where you need an actual guarantee.

  • A team argues their two sessions are really one workflow, so the split is bureaucratic — how do you answer?
    The split is only meaningful if the sessions differ in what they hold. Ask which credentials each one is issued, which tools each can call, and exactly what crosses between them. If the triage session's output is a validated record and the privileged session cannot fetch the raw issue, the boundary is real and enforced in code. If the raw text is forwarded for context, they are right that it is bureaucratic — and the design has three legs again.
  • Where does the Rule of Two differ from the lethal trifecta in practice?
    The trifecta's third leg is an exfiltration channel, so it scores data theft. The Rule of Two's third leg is any state change or external communication, which also catches destructive and reputational outcomes from an agent holding no secrets at all. The Rule of Two also names an explicit escape — human validation — and pins the unit of analysis to the session. Run both; they catch overlapping but not identical designs.
  • When is it acceptable for a session to hold all three legs?
    When a human validates the consequential action, the action is rare and legible enough that the review is genuine, and the reviewer sees the concrete effect rather than a summary of intent. That is a narrow band — a production deploy or an outbound payment. If the pattern recurs many times a day, approval fatigue makes the control ornamental and you should split the session instead.

saying these in an interview costs you the question

  • Applies the rule to the product rather than to each session
  • Counts human approval as a full mitigation for routine, high-frequency actions
  • Forwards the raw untrusted text alongside the structured handoff
  • Assumes a two-leg session is automatically safe
  • Declares a leg absent while a subtle write or egress path remains

context