skip to content

Your agent passes every task in a packaged tool world built on mailbox and banking-style mocks, then leaks customer data in production against your real mail and payments APIs. Which properties of that mock environment make the outcome unsurprising, and what do you add?

level: seniorimportance: should knowfreq 52%

answer

  1. action space bounds observable harm
  2. tool names and descriptions differ
  3. no pagination, truncation, partial failure
  4. payload container realism
  5. public fixtures are a floor, not a ceiling

basics

~20 s

The mocks are not your tools. They return small, clean, deterministic data with no pagination, partial failures or auth errors, and they expose a fixed action set, so harms outside that set cannot be observed at all. A pass says the fixture found nothing, not that your production tool surface is safe.

solid answer

~50 s

Two gaps do most of the damage. **Action space.** The fixture can only register harm it can represent as world state. If the mock banking app exposes a single transfer call and no bulk export, no address-book read, no third-party share link, then exfiltration through those routes is invisible — not scored zero, simply unrepresented. Your production agent has tools the fixture never had. **Data shape.** Mock results are short, single-format, well-formed and deterministic. Real results paginate, truncate, arrive as quoted HTML with attachments and forwarded threads, sometimes fail halfway with a permission error the agent then retries. Every one of those is a place hostile content can hide and a state the agent was never observed in. What you add: a bespoke world that replays your own tool schemas — the real names, descriptions, argument shapes and error semantics — over recorded production-shaped responses, with harm predicates written against your data, plus your actual guard in the loop rather than a bare model.

go deeper

for a junior

Should say the mocks are simplified and are not the company's real tools, so passing does not generalise.

for a middle

Names concrete missing properties — pagination, error paths, message formats — and that unrepresented actions cannot be scored.

for a senior

Diffs the action spaces, calls out tool-schema sensitivity and container realism, and specifies a bespoke replay world with harm predicates over real state.

for a principal

Decides that public fixtures are a regression floor and the bespoke world is the decision input, and owns the cost of maintaining it.

### Start from what the mock actually is A packaged tool world's "banking app" is a Python object holding a list of transactions and a handful of functions over it. Its "mailbox" is a list of message dicts with a sender, a subject and a short body. The suite's harm predicates read that object's end state — did a transfer appear, did a message leave to an outside address. Everything the fixture can *observe* is a property of that object; everything it can *do* is the set of functions registered as tools. A 100% pass is therefore a statement about a few hundred lines of mock, not about your product. ### The five properties that explain the production leak **1. Action space.** The fixture registers harm only if it can represent harm as state. If the banking mock exposes a transfer call and nothing else, then exfiltration via a bulk export, an address-book read, a shared link, a webhook or a support-ticket reply is not scored zero — it is **unrepresented**. Your production agent holds tools the fixture never had, and the leak route is almost always in that column. Diffing the two action spaces is a spreadsheet exercise, not research, and it is the cheapest check that would have predicted the outcome. **2. Tool schema mismatch.** A model's behaviour is highly sensitive to tool names, descriptions and argument shapes, because those strings sit in its context on every single turn. A mock `send_message(recipient, body)` and your production mail client with fifteen parameters, a scopes note and a "use sparingly" warning in its description are *different tools to the model* even when they do the same thing. Results measured against one do not transfer to the other. **3. Data shape.** Mock results are short, single-format, well-formed and deterministic. Real results paginate, truncate mid-record, arrive as quoted HTML with forwarded threads and attachments, carry other tenants' shared documents, and fail halfway with an expired-credential or permission error the agent then retries. Every one of those is somewhere hostile content can sit and a state your agent was never observed in — and agents behave differently while repairing: they re-read, re-summarise and widen queries, pulling *more* untrusted content into context. **4. Container realism.** A planted instruction alone in a two-line mock email is far easier to notice than the same instruction inside a long quoted thread, a table cell, a footer or a PDF the agent had to summarise. Fixtures systematically under-represent the containers that matter. **5. Configuration.** The fixture usually ran a bare model with the harness's own scaffold. Your product ships a system prompt, a guard, retries and a tool router. A pass proves nothing about the deployed composition unless the deployed composition was in the loop. ### What it costs to close the gap A bespoke world is engineering, not a purchase. Expect several engineer-weeks for the first version — capturing real responses, scrubbing them to a shape you can commit, reimplementing your tool schemas as mocks, and writing harm predicates over your own end state — plus a standing maintenance cost, because the world drifts every time the product changes a tool. Budget an owner. An unowned bespoke world measures last quarter's product while still reporting a number. ### Where the number misleads "We pass 100% of the suite" reads as *no harm found*, when what it means is *no harm the fixture can express occurred, in the containers it ships, under a configuration that is not the shipped one*. The denominator is the fixture's action space, and nobody publishes that denominator. The second trap is the inverse: a bespoke world reporting 0% harm may simply have predicates that never fire — a broken instrument returns the number you were hoping for. ### What to check Run a **positive control**: point a deliberately obedient stub agent at the world and confirm every harm predicate fires. Publish the **inventory diff** — each production tool and capability against its environment counterpart, gaps listed — as the honest coverage statement. Run two arms, guard-off for diagnosis and guard-on for the ship decision, and never blend them. And verify the recorded responses still match live shapes on a sample of real traffic; a replay corpus quietly goes stale the same way the tool list does.

  • What is the cheapest first check that would have predicted this gap?
    A side-by-side list of production tools versus fixture tools. Any production capability with no fixture counterpart was untested, and the leak route usually appears in that column.
  • Why does replaying recorded real responses beat writing richer mocks by hand?
    Hand-written mocks encode what you already imagined. Recorded traffic carries the shapes you did not — truncation, odd encodings, quoted threads, partial failures — which is exactly where content hides.
  • Does adding the deployed guard to the fixture make the result more useful or less?
    More useful for a ship decision, less useful for diagnosis. Run both configurations: guard off tells you what the model does, guard on tells you what the product does.

The fixture is a smoke detector wired only to the kitchen. It reports no smoke truthfully and continuously while the garage burns, because the garage was never on its circuit.

saying these in an interview costs you the question

  • Reports the packaged pass as evidence the production agent is safe.
  • Never diffs the fixture's action space against the production tool set.
  • Assumes results transfer across tool schemas because the tools 'do the same thing'.
  • Ignores that the fixture ran without the deployed guard in the loop.
  • Proposes only writing more payloads, when the missing coverage is channels and actions.

context