The only tool server in scope for an agent engagement is a third-party production one you may call but cannot mutate or replay. How do you split the work between a stand-in you host and that limited live access, and what do you refuse to claim from stand-in evidence alone?
answer
- fixture for variance, live for transfer
- shortlist, then pre-register each live call
- highest fixture rate = loosest copy
- one live attempt, no replay
- no rate carried onto production
basics
~20 sDo all iteration on the stand-in, where the surface is mutable and repeats are cheap, and spend live calls only on confirming a shortlist. Decide the confirmation criteria before touching production. Never claim the partner is exploitable, or quote a rate, from stand-in runs alone.
solid answer
~50 sSplit it by what each environment can answer. The stand-in answers behavioural questions: does this agent act on content arriving through a tool surface, at what rate, under which surface variants. Those need many repeats and a mutable server, so they cannot be asked live. The production server answers exactly one question per attempt: does the shortlisted case still behave that way against the real contract. That asymmetry drives the plan. Iterate broadly on the fixture, shortlist ruthlessly, and pre-register each live attempt: what will be called, what observation counts as confirmation, who is on the call, and what the rollback is. Live attempts are a scarce, non-replayable resource and every one of them is a real action on someone else's system, so the confirmation criterion must be observable before you spend it. The refusal line: stand-in evidence supports claims about the agent, never claims about the partner's exposure and never a success rate attributed to production.
go deeper
Understands that production calls are limited and should not be used for trial and error.
Structures it as explore-on-the-fixture, confirm-live, and keeps the two sets of evidence separate in write-up.
Adds pre-registration, single-pass capture, and shortlist selection by transferable mechanism rather than fixture rate.
Owns the policy: what claims each evidence class supports, how live attempts are authorised and deconflicted with the partner, and where the budget between fixture breadth and live confirmation actually sits.
**The two environments answer different questions, and that asymmetry is the whole design.** A stand-in you host is where variance lives: repeats are cheap, the surface is mutable, failures are free, and a sensitive handler can record server-side that it was invoked and with what arguments. Limited production access is where transfer lives: it is the only place that can tell you the real contract still permits the behaviour. Iterating live burns a non-replayable resource on uncontrolled single samples and takes real actions on somebody else's system for no experimental gain; concluding from the fixture alone produces a report the partner dismisses in one reply. **Running the split.** - *Phase one — fixture.* Explore surface variants, establish per-variant rates over N attempts, record the fixture identifier (a content hash over the served tool definitions) with every run. This is where your token budget goes: variants times N times a multi-turn agent loop. - *Shortlist.* Keep the cases whose mechanism plausibly survives the real contract — not the ones with the best fixture numbers. Ranking by fixture rate is a trap, because the top-rated variant is very often the one your copy is loosest about; a variant that only works because your stub accepts an over-long argument or returns unescaped content is the least likely to transfer, and it will rank first. - *Pre-register each live attempt.* Before touching production, write down the case, the exact call, the observation that counts as confirmation, the stop criteria, who is watching, and how it is undone. Written afterwards, this is indistinguishable from cherry-picking, and everyone reading the report knows it. - *Phase two — live.* Execute the shortlist, capture everything on the single pass, and stop at the first result that changes the plan. **What it costs, and which cost people misjudge.** The token spend is real but predictable: a few hundred fixture attempts is one to two million tokens and hours of wall clock. The costs that actually determine the shape of the engagement are the other two. First, engineer time on the stand-in and on single-pass instrumentation — you get one shot per live case, so the harness must capture the exact request, the returned bytes, the agent's intermediate steps and the deciding observation, or you will need a second attempt you cannot have. Second, and largest, authorisation lead time: getting agreement from the partner on which production calls are permitted is measured in days to weeks, it is a scope conversation rather than a tester's unilateral call, and it is the item that determines whether phase two happens at all. Standard engagement deconfliction practice covers it; what is specific here is only that non-replayability makes each live attempt a single unrepeatable observation. **What each phase may claim.** | Evidence | Supports | Does not support | |---|---|---| | Fixture, N attempts | "This agent acts on content delivered through a tool result of this shape, at this rate, under these settings, against this fixture" | Any statement about the partner's exposure | | Live, one confirmed attempt | "This specific case did reproduce against the real contract on this date" | A rate, a severity derived from a rate, or a claim about other cases | | Live, one failed attempt | "This case did not reproduce on this date, with this detector" | "The control works" — unless you identified the mechanism | **Where the number misleads.** The denominator swap is the failure mode to name explicitly. A fixture rate is calculated over dozens of attempts against a surface you built, scored by a server-side invocation log. A live result is one attempt against a real surface, scored by a weaker detector. Printing "82 percent success" beside "confirmed in production" reads to a partner as if four in five production calls would succeed, which nothing measured supports. Report the two side by side with their attempt counts and their detectors named, and never carry the fixture's rate across the boundary. The mirror-image error is treating a failed live confirmation as an omission rather than a result; a non-reproduction is data, and the report must carry it plainly next to the fixture behaviour and the divergence you identified, if you identified one. **What you would check before spending a live attempt.** That the detector works — trigger the observation deliberately in a rehearsal and confirm it is visible with only the telemetry production will give you. That the pre-registration is written and agreed. That the shortlisted case does not depend on a permissiveness you invented, checked by diffing your fixture against the partner's contract. And that the single pass captures enough to interpret either outcome, because "we would need to try again" is not an available answer.
- Why is ranking the shortlist purely by fixture success rate a trap?The top-rated variant is often the one your stand-in is most permissive about, so it is the least likely to transfer. Rank by whether the mechanism survives the real contract, then by rate.
- What has to be true of the harness before you spend a live attempt?It must capture enough on one pass — the exact call, the returned content, the agent's steps and the observation that decides the outcome — that no second attempt is needed to interpret the result.
- The live confirmation fails. What goes in the report?Both results, plainly: the behaviour observed on the fixture, the non-reproduction against production on that date, and the identified difference if you found one. Not a fixture-only finding with the contradiction omitted.
The fixture is a lab yield measured over forty batches; the live confirmation is one batch off the real factory line. Printing the lab percentage next to the single factory run invites everyone to read the lab number as the factory's.
saying these in an interview costs you the question
- Iterates against the production tool server because it is more realistic.
- Reports a fixture success rate as if it applied to the partner's system.
- Spends a live attempt with no pre-agreed observation that decides the outcome.
- Shortlists by fixture rate alone.
- Treats a failed live confirmation as an omission rather than a result.