When should an agent eval harness stub, record, or call third-party tools live?
answer
- fidelity versus determinism
- canned responses encode your beliefs
- cassettes look authentic while going stale
- irreversible side effects cannot be rolled back
- something must detect the drift
basics
~20 sStub cheap, self-owned behaviour where you control the contract; replay recorded real responses when fidelity matters and determinism is required; call live only in a small, separate suite whose job is detecting that the real API has drifted away from your stubs and recordings.
solid answer
~50 sThe three options trade fidelity against determinism and cost. Hand-written stubs are fastest and free but encode your beliefs about the API, so the suite keeps passing after the real service changes. Recorded responses — captured once from real calls and replayed by matching on the request — give real payload shapes with deterministic playback, at the price of cassettes that go stale, brittle request matching, and secrets that must be scrubbed before they land in the repo. Live calls are the only true fidelity but bring rate limits, cost, latency, and irreversible side effects such as real emails or charges, and their flakiness gets scored as agent failure. The usual layering is stubs and recordings on the pre-merge tier, plus a small nightly contract suite against the live API that fails loudly when a recording no longer matches reality, prompting a re-record.
go deeper
Know the three options — hand-written stubs, replayed recordings, live calls — and be able to say that stubs and recordings make runs repeatable while live calls do not.
Explain each option's failure mode: stubs encode assumptions and never notice API changes, recordings look authentic while going stale and match requests brittlely, live calls bring rate limits, cost and irreversible side effects.
Show the layered design and the pieces teams forget: a live contract check that detects staleness, provider errors classified as infrastructure rather than agent failure, secret scrubbing in cassettes, and deliberate hostile-content responses in the fixtures.
Own the maintenance economics — who re-records, how often, and what an integration is allowed to cost in eval time — and the risk framing for tools with irreversible effects, including when a provider sandbox is mandatory rather than optional.
## The axis: fidelity versus determinism An agent eval is only meaningful if the tools behave the same on every rollout, because pass^k compares repeats. But an agent is only meaningful if the tools behave like the real ones. Those two requirements pull in opposite directions, and the stub / record / live choice is where you resolve them — per tool, not once for the whole suite. ## Stubs A stub is a hand-written implementation that returns canned responses, often driven by the task's fixture: look up the order in the seeded database and return it in the shape the real API would. *Strengths.* Free, instant, fully deterministic, trivially able to produce edge cases you cannot conjure on demand — the rate-limit error, the 500, the partially-refunded order, the empty result. Error-path coverage is the strongest argument for stubs, because real services rarely fail on cue. *Weakness.* A stub encodes what you *believe* the API does. When the provider adds a required field, changes an enum, tightens validation, or starts paginating, your stub does not notice and the suite stays green while production breaks. Stubs also tend to be too clean: real payloads carry nulls, unexpected extra fields, inconsistent casing and truncated text, all of which are things the agent must cope with. ## Recorded responses Recording captures real request/response pairs once and replays them thereafter, matching an incoming request against the stored set — the pattern popularized by VCR-style libraries. The result is real payloads with deterministic playback. *Strengths.* Realistic shapes without runtime cost or flakiness. It is also the mechanism that makes a production-derived task replayable at all: the trace tells you what the tools returned, and the recording lets a future rollout see the same thing. *Weaknesses, and they are real.* Recordings go stale exactly like stubs, just more slowly and more invisibly, because the payloads look authentic. Request matching is brittle: match too strictly and any prompt change that alters a query string misses the cassette; match too loosely and the wrong response is replayed. Agents make the matching problem worse than ordinary tests do, because the request the agent sends is model-generated and varies between rollouts — you often have to match on a normalized subset of the request rather than the whole thing. Recordings also capture secrets, tokens and customer data, which must be scrubbed before they are committed, and they capture a moment in time, so a recording of a paginated list is a fixed page count forever. ## Live calls *Strengths.* The only configuration that tells you the integration genuinely works today. *Weaknesses.* Nondeterministic responses, rate limits that bite hardest under the parallelism k repeats demand, per-call cost, latency that blows wall-clock budgets, and side effects that are not resettable — a live call that sends an email or moves money cannot be rolled back with the database. Live flakiness is also *mis-attributed*: a 429 mid-rollout looks like the agent failing the task unless the harness distinguishes infrastructure errors from agent errors, which it should. ## The layered answer Most mature harnesses use all three, assigned per tier and per tool: - **Pre-merge tier:** stubs and recordings only. Fast, free, deterministic; its job is catching breakage from your own changes. - **Nightly full suite:** recordings for third parties, live for services you own and can reset (your own database, your own sandbox API keys). - **Contract suite:** a handful of live calls per integration, run on a schedule, asserting the response still matches the recorded shape. This is the piece teams skip, and it is the one that makes the other two trustworthy — without it, staleness is undetectable by construction. Sandbox environments offered by payment and shipping providers sit between recorded and live: real API surface, resettable-ish state, no real-world side effects. Prefer them to production for any tool with irreversible effects. ## Two agent-specific wrinkles **The simulated user is a tool.** In conversational suites the counterpart user is often a model, and it is a variance source on the same footing as any external service. Pin its model and prompt, or script it for the deterministic tier. **Untrusted content stays untrusted.** Whatever the fidelity choice, tool output is attacker-influenceable in production. An eval suite should include tasks where the stub or recording returns hostile content, precisely because a live API will not hand you one on demand. ## How to answer Frame it as fidelity versus determinism and cost, describe each option's failure mode concretely (stubs encode beliefs; recordings go stale and match brittlely; live is nondeterministic with irreversible effects), and land on the layered design with a live contract check as the staleness detector. The weak answer picks one globally and defends it as a principle.
- Recorded responses drift silently. What do you build so the drift is detected rather than discovered in production?A small live contract suite, run nightly or weekly outside the main eval: it issues one real call per integration and asserts the response still satisfies the schema and invariants the recording relies on. When it fails, the recordings for that tool are stale and get re-recorded. Some teams also expire cassettes by age, forcing a refresh on a fixed cadence rather than waiting for a mismatch.
- Why is request matching harder for recorded tool calls in an agent eval than in a conventional integration test?In a normal test the request is code you wrote, so it is byte-identical every run. In an agent eval the request is generated by the model, so parameter order, optional fields, whitespace and phrasing vary between rollouts. Strict matching then misses the recording and the rollout fails for the wrong reason. You typically match on a normalized subset — method, path and the semantically meaningful parameters — and accept that this loosens the guarantee.
- How should the harness treat a tool that returns a rate-limit error during a rollout?Distinguish infrastructure failure from agent failure. If the harness scores a 429 as a failed task, your reliability numbers measure the provider, not the agent. Classify the run as errored, retry it or exclude it from the denominator, and report the error rate separately. It is also a reason to keep third-party calls stubbed or replayed on the tier where you run high concurrency for pass^k.
saying these in an interview costs you the question
- Assuming a green suite means the real API still behaves that way
- Committing recordings without scrubbing tokens and customer data
- Running live third-party calls inside a high-concurrency repeated suite
- Scoring provider errors and timeouts as agent task failures
- Stubbing only happy-path responses, so error handling is never exercised