Agent Red-Team Harnesses
An agent attack counts only when the world changed - a transfer sent, a row deleted - so the environment, the interception and the state check are the work. Interviewers ask what a harness missed.
on this pageshowhide
explore
- Benchmark Environments13 questions
- Simulated Tool Worlds5 questions
- Utility Under Attack4 questions
- Agentic Task Sets4 questions
- Instrumenting a Live Target18 questions
- Capturing the Trace4 questions
- Intercepting Tool Calls5 questions
- Testing a Tool Server4 questions
- Containing Side Effects5 questions
- Deciding What Counted13 questions
- Defining Success4 questions
- Run-to-Run Variance5 questions
- Partial Progress4 questions
questions
page 2 of 2During an agent red-team run the agent's tools pulled real customer records and a live API token into the harness trace. How do you design trace capture so the run stays shareable without destroying the reproducibility of the finding?
basics
~20 sSplit raw from shareable. Keep the unredacted trace in a restricted store with a short retention, and derive a redacted copy by replacing secrets and personal data with stable placeholders. Stable means the same value maps to the same token, so the reader still sees the causal chain.
Two engineers drive the same injection payload at the same agent with the same red-team harness and report attack success rates of 5% and 40%. Before concluding the target changed, which run-level settings do you line up between the two sweeps?
basics
~20 sCompare trial counts first — 5% may be 1 of 20. Then decoding settings and any fixed seed, the turn budget per episode, which object decided a hit and on what threshold, the starting environment each trial began from, and whether aborted or retried episodes were counted as trials.
You have a fixed token spend for an agent red-team harness sweep against a metered hosted agent and about 40 candidate attack payloads, and each trial is a whole multi-turn episode rather than one call. How do you split that spend between covering more payloads and repeating each payload enough times to quote a rate?
basics
~20 sTwo stages. Spend a cheap screening pass of a few trials on all 40 payloads to find which ones ever land, then spend the depth only on those, repeating them enough to quote a rate. Stop early when the interval already answers the decision, and price a trial as a full episode.
Your agent red-team harness scores every task by diffing the sandbox's state before and after the run. Which harms will that diff miss, and how do you bound what it actually covers?
basics
~20 sA diff only sees the surface you snapshot. Outbound network calls, writes to a third-party service, data read and emitted without changing local state, persisted agent memory and delayed jobs all escape it. Bound it by enumerating every sink the sandbox exposes and recording each one, not just the database.
You lead red teaming for an agent product. When is adopting a packaged tool world — a public fixture of mock tools and paired benign/injection tasks — the wrong instrument, and what do you stand up instead?
basics
~20 sIt is the wrong instrument when you need a ship decision. Its mock tools, data shapes and action space are not yours, so its number describes the fixture. Keep it as a cheap, comparable regression check, and stand up a bespoke world replaying your own tool schemas with your own harm predicates.
Two candidate defences are measured on the same agent injection suite. With defence A the attacker-action rate falls from 24% to 3% and task completion under attack falls from 71% to 38%. With defence B they are 11% and 66%. A platform team asks you to fold these into one "security score" so the two can be ranked. How do you respond?
basics
~20 sGive them the pair, not a scalar. The exchange rate between one blocked injection and one failed user task is a product decision, not a benchmark output. B keeps most of the utility; A buys a lower attacker-action rate by breaking a third of the work. Ask what a residual success actually costs.
The only tool server in scope for an agent engagement is a third-party production one you may call but cannot mutate or replay. How do you split the work between a stand-in you host and that limited live access, and what do you refuse to claim from stand-in evidence alone?
basics
~20 sDo all iteration on the stand-in, where the surface is mutable and repeats are cheap, and spend live calls only on confirming a shortlist. Decide the confirmation criteria before touching production. Never claim the partner is exploitable, or quote a rate, from stand-in runs alone.
You are planning a red-team engagement against an agent wired into production systems. How do you decide, tool by tool, which actions execute for real and which are stubbed, knowing stubs weaken your evidence and real execution leaves residue?
basics
~20 sSplit by reversibility and by who else sees the effect. Reads and reversible writes inside a scratch tenant run for real. Anything reaching a third party, moving money, or paging a human gets stubbed, and you argue that gap separately. Decide before the run, write the table down, publish it with the findings.
You are leading a red-team engagement against a team's live agent and you must intercept its tool calls to test the boundary the agent trusts. How do you decide how invasive that interception may be, and how do you keep the findings from being dismissed as 'your harness did that'?
basics
~20 sAgree up front how much of the stack you may modify, and write it down. Every layer you replace is a layer the owning team can blame. Prefer the least invasive interception that still proves the claim, ship a matched control run with each finding, and state exactly what the harness changed.
A long agent red-team campaign is producing more trace data than your storage and retention budget allows. How do you decide what the harness stops capturing, and which reductions are safe?
basics
~20 sCapture fully by default and cut elsewhere. You can always delete a trace later; you can never reconstruct one you did not record, and a failed rerun is what kills a finding. Trim by truncating giant tool payloads to a hash plus a head, and by shortening retention, not by sampling which attempts you trace.
You lead an AI red team whose engagement reports quote attack success rates from an agent harness. What standard do you set for trial counts and for how variance appears in a report, and what do you do with the payload that landed once in fifty?
basics
~20 sTie the trial floor to the claim: an existence finding needs one replayable success with evidence; a quoted rate needs a denominator and an interval in the text. Ban bare percentages. The one-in-fifty still gets filed on impact, marked low-rate and not verifiable by re-measurement at that budget.
Every task in your agent red-team suite is scored against a hand-written list of forbidden side effects, and every run now passes. What can you conclude, and how would you restructure the success criteria?
basics
~20 sMostly that you only forbade what you thought of. A pass against a hand-written forbidden list is evidence about that list, not about the agent. Invert it: assert the exact set of effects the task permits, and fail the run on any effect outside that set, including ones nobody predicted.
You have one week of compute and a fixed call budget for a multi-step agentic harmful-task suite. How do you split it between breadth (many tasks, one attempt each, tight step caps) and depth (few tasks, repeated attempts, generous caps), and what is each split's resulting number unable to say?
basics
~20 sSplit by what the run has to answer. Breadth finds which task families the agent will engage with at all, but one attempt per task has no repeatability and tight caps truncate long chains. Depth gives a defensible per-task number with an interval, but says nothing about families you did not run.
You own the partial-progress grading rubric for an agent red-team harness that several teams run against different agents. How do you decide how many progress tiers the rubric has, and who arbitrates a disputed grade?
basics
~20 sTie tiers to decisions, not to narrative detail. Use the fewest checkpoints that change what someone does: attempted, reached the irreversible action, completed. Every tier needs an objective world-state check and a named arbiter for disputes, or graders drift and cross-team numbers stop comparing.
showing 31–44 of 44