skip to content

You lead red teaming for an agent product. When is adopting a packaged tool world — a public fixture of mock tools and paired benign/injection tasks — the wrong instrument, and what do you stand up instead?

level: principalimportance: should knowfreq 38%

answer

  1. comparability versus relevance
  2. tripwire versus decision input
  3. action-space inventory diff
  4. public payloads leak into filters
  5. bespoke worlds rot without an owner

basics

~20 s

It is the wrong instrument when you need a ship decision. Its mock tools, data shapes and action space are not yours, so its number describes the fixture. Keep it as a cheap, comparable regression check, and stand up a bespoke world replaying your own tool schemas with your own harm predicates.

solid answer

~50 s

A packaged world buys three real things: it exists today, it is comparable across teams and releases, and it encodes injection placements someone else already thought through. It does not buy relevance. Its tools are not your tools, its action space is not your action space, and its payloads circulate publicly, so a clean result can reflect familiarity rather than robustness. So split the roles. The packaged world is a **tripwire**: cheap, fixed, run on every change, and the only claim it supports is "nothing that used to be caught got worse". The **decision input** is a bespoke world that mirrors your production tool surface — real tool names and descriptions, recorded response shapes, your guard in the loop — with harm predicates written against your own state and your own definition of a violation. The cost is maintenance: a bespoke world drifts every time the product's tools change, so budget the ownership, or it silently measures last quarter's product.

go deeper

for a junior

Should recognise that a public fixture's tools are not the company's tools and so its result does not directly speak to the product.

for a middle

Separates regression use from decision use, and names the action-space and data-shape gaps that motivate a bespoke world.

for a senior

Specifies the bespoke world concretely — recorded responses, deployed guard, harm predicates over real state — and the inventory diff as the coverage artifact.

for a principal

Owns the portfolio: which instrument answers which question, the maintenance budget and owner, contamination and metric-gaming policy, and gating new tools on environment support.

### Decide by the question on the table A packaged tool world — a public fixture of mock tools plus paired benign and injection tasks, AgentDojo being the reference example — is an instrument with a narrow, real competence. Match it to the question: | Question | Instrument | Why | |---|---|---| | Are we comparable to the field? | Packaged, unmodified | A shared fixture is the only thing that supports a cross-team comparison; local edits destroy it | | Did this change make us worse? | Packaged, pinned subset, fixed seed | The claim is *relative*, so the fixture's irrelevance to your surface does not matter | | Can we ship this agent to customers? | Bespoke, always | The harms that gate a launch are defined by your tools and your data, and the fixture cannot represent them | The failure mode is using row three's decision with row two's instrument. ### What each one actually costs The packaged world is close to free in engineering and non-trivial in compute: it exists today, someone else designed the placements, and the bill is the sweep's model calls — user tasks times injection tasks times a multi-turn episode, which is thousands of calls and hours of wall clock even for a mid-sized suite. The bespoke world inverts that. Several engineer-weeks to stand up: your tool schemas reimplemented verbatim as mocks, responses recorded from real traffic and scrubbed to something committable, harm predicates that read your own end state rather than string-matching the transcript, the deployed guard configuration in the loop plus a guard-off arm for diagnosis, and an inventory diff that flags any production tool with no environment counterpart. Then a standing maintenance cost, because it drifts every time the product ships a tool. Name an owner and gate tool launches on environment support, or the coverage claim silently becomes false — an unowned bespoke world keeps producing numbers about last quarter's product. ### Where the numbers mislead **Comparability is not relevance.** The packaged number is comparable precisely because it is the same for everyone, which is the same reason it is not about you. Its action space, tool names, data shapes and payload containers are the fixture's. Reported without that caveat it becomes an executive metric, and an executive metric gets optimised. **Optimising it is trivial and worthless.** The fixture's payload strings are public. Add them to a deny list and the score moves without the product changing. Blocking known strings may be fine as defence; reporting the resulting movement as a robustness gain is not, because the fixture's whole value was being an *independent* floor and you have just removed the independence. **A local fork is not the public number.** Teams edit a suite — swap a tool, change a placement, drop tasks that error — and then compare to published field results. Different denominator, different population, no comparison. **Blending is the worst of the three.** Packaged and bespoke results have different denominators over different populations. Averaging them into one safety score hides which one moved and makes regression unattributable. ### What to check before you report Confirm the packaged suite is running unmodified, at a pinned version, on a seeded subset checked into the repo — otherwise its comparability claim is void. Diff your deny lists and guard rules against the fixture's known strings and disclose any overlap; if it exists, the packaged number no longer supports a robustness claim. Run the positive control on the bespoke world — an obedient stub agent must make every harm predicate fire — before treating a low harm rate as good news. And publish the inventory diff: every production tool and capability against its environment counterpart, gaps enumerated. That artifact is the honest coverage statement, and it is the one I would hold a team to, because it converts a percentage into a specific list of what was never exercised. Report the two numbers separately, each with its denominator and an explicit claim limit — "regression on a fixed public fixture" and "harm rate over our own tool surface under the deployed configuration" — and never let either be presented as evidence for the other's claim.

  • What single artifact best states your agent coverage honestly?
    An inventory diff: every production tool and capability against its environment counterpart, with the gaps listed. It converts a vague percentage into a specific statement of what was never exercised.
  • How do you stop a packaged benchmark number from becoming the org's safety metric?
    Publish it with its denominator and an explicit claim limit — regression only — and make the ship gate reference the bespoke world's predicates. If leadership wants one number, give them the bespoke one.
  • Your product ships a new tool next sprint. What is the environment requirement?
    A counterpart in the bespoke world — schema, recorded responses, and a harm predicate — before the tool goes live, otherwise the coverage statement silently becomes wrong.

saying these in an interview costs you the question

  • Presents a public benchmark result as the ship decision.
  • Rewrites the packaged suite locally and then compares to published field results.
  • Adds the fixture's payloads to production deny lists and reports the improved number.
  • Builds a bespoke world with no owner and no gate tying new tools to environment support.
  • Reports packaged and bespoke results as one blended number.

context