skip to content

How do you scope a deploy agent's tools so a hijacked run stays contained?

level: seniorimportance: must knowfreq 56%

answer

  1. Assume the worst call will be made
  2. Narrow the surface, not the prompt
  3. Scope belongs to the tool, not the arguments
  4. Prompts suggest; the API enforces
  5. Split read tools from write tools

basics

~20 s

Replace generic tools with narrow, capability-shaped ones, bind the dangerous parameters server-side rather than letting the model choose them, split read from write, and enforce every limit at the API that executes the call — never in the prompt.

solid answer

~50 s

Assume the model will be talked into calling the worst tool you gave it, then design so that the worst tool is not very bad. Three moves do most of the work. First, granularity: instead of `run_shell` or `run_sql`, expose `scale_deployment` and `read_recent_logs` — a generic executor grants everything the underlying binary can do, and no description constrains it. Second, bind the risky parameters outside the model's control: the namespace, cluster and environment are fixed by the tool implementation, and only the deployment name and replica count come from the model, validated against an allowed set and a numeric range. Third, enforce at the resource server: the agent's role must not include the operations you excluded, so a bypassed tool wrapper still hits a deny. Add read/write separation, quotas, and separate agents per trust domain rather than one super-agent holding every capability.

code

python · 13 lines
python
ALLOWED_DEPLOYMENTS = {"checkout-api", "checkout-worker"}
NAMESPACE = "checkout-prod"  # bound by the tool, never chosen by the model


def scale_deployment(name: str, replicas: int) -> dict:
    if name not in ALLOWED_DEPLOYMENTS:
        raise PermissionError(f"deployment out of agent scope: {name}")
    if not 1 <= replicas <= 10:
        raise ValueError("replicas must be between 1 and 10")
    return {"namespace": NAMESPACE, "name": name, "replicas": replicas}


print(scale_deployment("checkout-api", 4))

go deeper

for a junior

Be able to say why a general-purpose tool such as a shell or raw SQL is riskier than several specific tools, and that limits written in a prompt are not enforcement.

for a middle

Explain the mechanics: capability-shaped tools, parameters the model may not supply, validation that rejects rather than clamps, and a backing role that lacks the excluded operations outright.

for a senior

Design from 'assume it gets hijacked'. Show the layered enforcement — wrapper plus resource-server permission — plus read/write separation, quotas that bound the rate of damage, and where you would still accept residual risk.

for a principal

Own the portfolio question: how many agents, split along which trust and reversibility boundaries, and what the standing scope policy is for new tools. Weigh the capability you lose by removing generic executors against the blast radius you buy back.

## The threat you are scoping against A tool-using agent reads untrusted content — tickets, logs, web pages, emails, tool results — and any of it may contain instructions. You cannot reliably prevent the model from being persuaded; adaptive-attack research through 2025 and 2026 repeatedly bypassed defenses built on detection. So the design question is not "how do I stop the agent choosing a bad call?" but "what is the worst call it is able to make, and how bad is that?" Scoping is how you shrink that worst case. It is a deterministic wall — enforced by code and IAM, not by persuasion — which is what makes it hold when the probabilistic layers miss. ## Move one: granularity beats generality A generic tool grants the union of everything reachable through it. `run_shell(command)` grants every binary on the image plus every credential the process can read. `run_sql(query)` grants the full rights of the database user, including `DROP` and cross-tenant reads. `http_request(url, method, body)` grants your entire internal network. Capability-shaped tools invert this: `scale_deployment(name, replicas)`, `read_recent_logs(service, minutes)`, `open_incident(title, severity)`. Each one names an operation, not a mechanism. The cost is that you write more tools and lose some flexibility — an agent with a shell can improvise. That flexibility is exactly the thing an attacker inherits, which is why generic executors belong only inside a sandbox with nothing valuable in it. ## Move two: bind the dangerous parameters server-side Even a narrow tool is unsafe if the model supplies the scoping arguments. If `scale_deployment` takes a `namespace` parameter, the model can be pointed at `payments-prod`. If the tool implementation hard-codes `checkout-prod` and accepts only the deployment name and replica count, that whole class of redirection disappears. The rule: the model chooses *what to do within a scope*; the scope itself is a property of the tool, not an argument to it. Then validate what remains — the name against an allowed set, the replica count against a range — and reject rather than clamp, so a suspicious call is visible in logs instead of silently normalised. ## Move three: enforce where the call lands Everything on the model side is advisory. A tool description that says "read-only, never use for writes" is prompt text; a JSON Schema `enum` is validated by your wrapper, which is code you control, but it is still one layer. The authoritative wall is the resource server: the agent's own role simply does not include `iam:*`, `billing:*`, or write access to other namespaces. Then a bug in your wrapper, a second call path, or a tool the model reaches some other way all terminate in a deny at the API rather than in a successful destructive call. This is why scoping and identity are the same conversation. A narrow tool over a broad role is one code review away from disaster; a broad tool over a narrow role is bounded but noisy. You want both narrow. ## Move four: split by risk and by trust domain Separate read tools from write tools so that the common case — investigation — needs no write capability at all. Where a workflow genuinely needs both, consider two agents: a diagnosis agent that reads everything and produces a recommendation, and an execution agent that can write but never reads untrusted content. That split removes the combination that makes hijacking profitable, at the cost of an extra handoff. Similarly, do not build one super-agent holding every team's tools because it is convenient. The blast radius of an agent is the union of its tools; merging two agents multiplies the risk of both. ## Operational limits Quotas and rate limits on the tools themselves bound the *rate* of damage: an agent permitted to scale deployments five times an hour cannot flap production two hundred times. Cost caps do the same for spend-shaped tools. These are cheap, they never depend on model behaviour, and they buy time for a human to notice. ## What scoping does not do It does not stop the agent doing a legitimate-looking wrong thing inside its scope — scaling the right service to the wrong number, closing the wrong ticket. That is what approval gates on irreversible actions and good evaluation are for. Scoping's job is to make the ceiling low, not the floor perfect. ## Answering this well Start from "assume it will be hijacked", then walk granularity, server-bound parameters, and enforcement at the resource server, with one concrete deny you would put in the role. Interviewers are listening for whether you put the boundary in the prompt or in the infrastructure.

  • Your agent genuinely needs ad-hoc shell access for debugging. What now?
    Give it a shell only inside a sandbox that holds nothing worth stealing: no host mounts, no ambient cloud credentials, no default network egress, and a lifetime of one session. The shell's power then equals the sandbox's contents. Anything the debugging flow needs from production comes in through narrow, explicit tools rather than through the shell's own reach.
  • Is a JSON Schema enum on the tool's parameters enough to enforce a scope?
    It helps, but it is one layer and not the authoritative one. Schema validation runs in your wrapper; a bug, a second call path, or a tool invoked from inside a code sandbox can bypass it. The enum should be mirrored by a permission the agent's role genuinely lacks, so the same restriction is enforced again where the call actually executes.
  • How do you decide whether two workflows should share one agent?
    By blast radius, not convenience. An agent's risk is the union of its tools, so merging a read-heavy investigation workflow with a write-heavy remediation one creates a single principal that can both ingest untrusted content and change production. If the workflows differ in trust of input or in reversibility of output, split them and pass a structured handoff between them.

saying these in an interview costs you the question

  • Putting 'never delete production data' in the system prompt
  • Exposing a generic shell or SQL tool with a cautionary description
  • Letting the model supply the namespace, cluster or tenant id
  • Relying on schema validation alone with a broad backing role
  • Building one super-agent that holds every team's tools

context