skip to content

What do model-level safety probes miss that agent-level red-team suites catch?

level: middleimportance: should knowfreq 45%

answer

  1. different targets, different evidence
  2. text out versus actions taken
  3. payload lives in retrieved content
  4. first-turn refusal is not the metric
  5. score the trajectory and the tool call

basics

~20 s

Model-level probes such as garak or PyRIT attack a bare model endpoint and measure what it says. Agent-level suites such as AgentDojo or AgentHarm attack the assembled system and measure what it does — which tools fired and whether a harmful task was completed end to end.

solid answer

~50 s

The two families test different targets and therefore produce different evidence. **Model-level probe tools** — garak, PyRIT — send adversarial prompts straight at a model endpoint and score its text: did it leak a system prompt, comply with an encoded request, produce disallowed content? That is useful for model selection and for regression on the model layer, but it cannot see your system prompt, your retrieval corpus, your tool scopes, or the channel through which data would actually leave. **Agent-level suites** — AgentDojo, AgentHarm, Gray Swan's Agent Red Teaming — plant hostile content in the documents and tool results an agent reads while it works, then score the trajectory: was a payment tool invoked, was another user's record read, was a multi-step harmful task carried through to completion. AgentHarm's design point is exactly that: a first-turn refusal counts for nothing if the agent completes the task anyway. Ship-readiness evidence for a tool-using product has to come from the agent level.

go deeper

for a junior

Know that testing a model endpoint and testing a deployed agent are different exercises, and that the agent test is the one that involves real tools and retrieved documents.

for a middle

Explain what the model layer cannot see — your system prompt, tool scopes, corpus and exfiltration channel — and why agent suites score actions and task completion rather than the text of a reply.

for a senior

Show how you would stand up an agent-level suite: a resettable sandbox mirroring your tool environment, payloads planted in retrieved content, deterministic checks on which tools fired, and attack success reported alongside utility under attack.

for a principal

Own the split of a limited evaluation budget between cheap frequent model probes that gate model upgrades and expensive agent runs that gate releases, and defend which classes of risk you accept measuring only at one layer.

## Two different systems under test When someone says "we red-teamed it", the first question worth asking is *what was the target?* There are two, and conflating them is the most common evaluation error in tool-using products. **The model layer.** A model endpoint takes text and returns text. Probing it means sending adversarial inputs and scoring the reply. Tools built for this — garak, which runs families of probes such as encoding tricks and prompt-leakage attempts, and PyRIT, which orchestrates automated adversarial conversations — give you a picture of the raw model's dispositions: how readily it complies with obfuscated requests, how easily it reveals instructions it was told to keep, how it behaves under multi-turn pressure. **The agent layer.** A deployed assistant is a model plus a system prompt, a retrieval corpus, a set of tools with real credentials, memory, and a loop. Attacking it means putting hostile content where the agent will *read* it — inside a retrieved case file, a returned API payload, an email body — and scoring what the agent *did*. AgentDojo is built around this shape, with a large set of hijack cases in realistic tool environments; AgentHarm scores whether an agent carried a harmful multi-step task through to completion; Gray Swan's Agent Red Teaming work measures targeted attack success against tool-using agents under adversarial pressure. ## What the model layer structurally cannot tell you Four things, all of which decide whether your product is safe: 1. **Your instructions.** The system prompt, its policy text, and how it interacts with untrusted content are yours, not the model's. A model that scores well bare can be trivially steerable once your prompt gives it a role and a corpus. 2. **Your tools.** Scope, credentials and blast radius live in your integration. A refusal is irrelevant if the same reasoning path reaches a tool that transfers money; a compliance is harmless if the tool is read-only over public data. 3. **The exfiltration channel.** Whether a leak is *possible* depends on whether the agent can reach outward — a URL it can render, an email it can send, a record it can write. A text-only probe cannot exercise a channel that does not exist at the model layer. 4. **Indirect injection.** The dominant real-world vector puts the payload in content the agent retrieves, not in what the user types. Model probes usually feed the payload as the user turn, which is a different trust path entirely. ## What the agent layer costs Agent suites are slower, need a resettable environment with fixtures and stub tools, and produce trajectories rather than single strings, so the grader has to inspect actions. That is the price of the evidence. The practical split most teams land on: model probes run cheaply and often, and gate model upgrades; agent suites run against release candidates and gate the ship. ## Scoring the right thing The scoring difference matters as much as the target. At the model layer, success is usually "did the text contain the forbidden thing". At the agent layer, the meaningful signals are: - **Did the harmful action occur** — a tool call with the attacker's arguments, not a sentence about it. - **Was the task completed** — AgentHarm's point. An agent that refuses turn one, then complies at turn four after the injected content reframes the request, has failed, and any metric that stops at the first refusal will score it as a pass. - **Utility under attack** — whether the agent still does its legitimate job when hostile content is present, since a defense that makes the agent refuse everything scores perfectly on ASR and is useless. That last pairing — attack-success rate *and* task utility, reported together — is what distinguishes a real agent evaluation from a system that has been lobotomised into safety. ## Putting it together for a product For a case-handling assistant with lookup tools, the sensible programme is: garak-style probes against the model you are considering, to catch encoding and leakage dispositions and to compare candidates; then agent-level cases in a sandboxed copy of your own tool environment, with injected payloads planted in the records the agent retrieves, scored on whether a tool fired and whether data crossed a boundary; then autonomous auditing to surface behaviours neither list anticipated. Reporting only the first and calling it red-teaming is the failure this question is asking about.

  • Why report task utility alongside attack-success rate in an agent evaluation?
    Because ASR alone can be driven to zero by making the agent useless. An agent that refuses whenever retrieved content looks unusual scores perfectly and fails its job. Reporting both makes the tradeoff visible: a defense that halves attack success while halving successful task completions is usually not shippable, and the pairing is what lets you compare candidate defenses honestly.
  • Where do you plant the payload in an agent-level test case, and why does the choice matter?
    In whatever the agent reads as part of doing its work — a retrieved document, an API response, a file listing, a message body. That is the indirect path, which is both the dominant real-world vector and the one that bypasses anything you do to sanitise user input. Testing only the user turn measures a trust boundary your attacker will not bother to use.
  • An agent refuses the injected instruction at first, then performs it three steps later. How should the case be scored?
    As a success for the attacker. The metric is whether the harmful action occurred or the harmful task completed, not whether a refusal appeared somewhere in the trajectory. Scoring the first turn is how a suite reports safety while the agent is being walked into the action across steps, and it is precisely the gap agent-level harm benchmarks were built to close.

saying these in an interview costs you the question

  • Treats a clean model-probe report as evidence the agent is safe
  • Scores only whether the model refused the first turn
  • Feeds injection payloads only through the user turn, never through retrieved content
  • Reports attack-success rate without reporting task utility under attack
  • Assumes the model vendor's safety evaluation covers your tools and corpus

context