In an agent eval harness, how do you reset state between repeated runs?
answer
- the world is consumed by running the task
- fixture, not shared service
- restore before, tear down always
- one instance per parallel worker
- run the same task twice back-to-back
basics
~20 sEvery task owns a seeded initial state that the harness restores before each attempt — a container rebuilt from a per-task image, a database restored from a snapshot or rolled back in a transaction, a scratch filesystem recreated. Each concurrent rollout gets its own isolated instance.
solid answer
~50 sTreat the environment as a fixture, not a shared service. Each task declares its seed state, and the harness materializes a fresh instance before every attempt: a container started from a per-task image (the pattern SWE-bench Verified uses), a Postgres database restored from a snapshot or wrapped in a transaction that is rolled back, a temp directory recreated from a template. Teardown must run even when the agent times out or crashes, or the next rollout inherits garbage. For concurrency, give each worker its own container and its own database schema or instance — never a shared one — and pin the clock and any ID generation the verifier will assert on. The cheap sanity check is running the same task twice back-to-back: if the second attempt behaves differently, reset is leaking and every pass^k number from the suite is suspect.
go deeper
Know that each eval run must start from the same seeded state, and that the harness restores it rather than the agent tidying up after itself. Be able to name containers and database snapshots as the usual mechanisms.
Explain the concrete reset mechanisms and their tradeoffs — per-task image, snapshot restore, transaction rollback, template directory — and why teardown has to run on the timeout and crash paths too, not just on success.
Demonstrate the operational judgment: per-worker isolation for concurrency, pinning clock and IDs the verifier reads, realistic seed data without real customer records, orphan reaping, and a double-run check that proves reset actually holds.
Own the fixture as a maintained asset with a cost and an owner: how seeds are refreshed as the schema drifts, who keeps images buildable, and where the line sits between a faithful sandbox and a cheap one. Be ready to argue the fixture budget against the eval budget.
## Why reset is the core of the harness An agent eval task is not a prompt and an expected string. It is a *world*: a database with rows in it, a filesystem with a repository checked out, a set of tools that can mutate that world, and a verifier that inspects the world afterwards. Because the agent writes to that world, the world is consumed by running the task. Running the same task again — which pass^k requires, and which every rerun after a code change requires — means rebuilding the world exactly as it was. If reset is imperfect, the failure is silent and it corrupts everything downstream. A task that passes only because a previous rollout left a row behind will pass in CI and fail in production. Worse, the corruption is order-dependent, so the suite becomes flaky in a way that looks like model nondeterminism and gets mis-attributed to it. ## Materializing the fixture **Per-task images.** The strongest form is a container image built per task, containing the exact dependencies and starting files. Each attempt starts a fresh container from that image and discards it afterwards. This is how repository-level coding suites such as SWE-bench Verified operate: the task's environment is pinned in an image so the agent's edits, installs and stray processes vanish with the container. **Database snapshots and transaction rollback.** For a stateful business suite — say 120 golden returns-and-refunds tasks over an orders schema — you seed a known dataset and restore it per attempt. Two mechanisms are common. Restoring a snapshot or template database is slow but total. Wrapping the whole rollout in a transaction and rolling it back at the end is fast, but only works if the tools all share the connection and nothing under test commits out of band or relies on its own transaction boundaries. Pick based on whether the agent's tools go through your process or over the network. **Filesystem and scratch space.** Recreate from a template directory rather than deleting selectively; deletion lists rot as tasks add files. **External services.** Anything the agent can mutate that you do not own is not resettable. That is the strongest argument for stubbing or replaying third-party calls rather than hitting them live inside a repeated suite. ## The details that bite **Teardown must be unconditional.** Agents time out, loop, exhaust budgets, and crash the harness. Run cleanup in a finally-equivalent, and additionally reap orphans at suite start, or a bad night leaves containers and databases behind until the runner fills its disk. **Isolation, not just reset, for concurrency.** To fit k repeats inside a wall-clock budget you will run rollouts in parallel. Parallel rollouts sharing one database or one working directory will interfere, and the interference presents as nondeterminism. Each worker needs its own instance — a separate container, and a separate schema, database, or namespace — and the harness needs a way to allocate and release those slots. **Pin the nondeterministic inputs the verifier reads.** If the verifier asserts on a refund reference or a created-at timestamp, freeze the clock and seed the ID generator, or write the verifier to compare structurally rather than by exact value. Otherwise reset is fine and the assertion is flaky. **Seed data must be honest.** A fixture that contains only clean, well-formed rows makes tasks easy in a way production is not. Seeds derived from real distributions — including the duplicate customers, the partially-shipped orders, the null fields — are what make a suite predictive. Do not seed real customer data; generate or anonymize it, because eval fixtures end up in repositories and CI logs. **Prove the reset works.** The standard check is running one task twice in immediate succession within a single suite run and asserting the two rollouts see identical starting state. A second check is randomizing task order between suite runs: if scores move with the ordering, state is leaking between tasks. ## Traces as tasks When a task is derived from a real production interaction rather than hand-authored, reset is what makes it runnable at all. The trace tells you what the agent did, but replaying it requires capturing the world it acted on — the account state at that moment, and the responses its tools returned — and re-materializing both. A trace without a reconstructable initial state is a bug report, not an eval task. ## How to answer Name the fixture mechanism (per-task image, snapshot restore or transaction rollback, template directory), insist on unconditional teardown, address concurrency with per-worker isolation rather than a shared environment, pin clock and IDs where the verifier looks, and give the double-run check as the way you know reset actually works. The weak answer says "we clean up after each test" without saying what guarantees it, or shares one database across parallel workers and then blames the model for flakiness.
- When would you prefer a snapshot restore over wrapping the rollout in a transaction you roll back?Transaction rollback is much faster and fine when every tool call runs through your process on the same connection. Prefer snapshot or template-database restore when tools reach the database over the network, when anything under test manages its own transactions or commits, when the task involves DDL, or when non-database state (files, caches, queues) must be reset in the same operation. Speed is worth nothing if the reset is incomplete.
- Your suite passes when run serially but scores lower under eight parallel workers. Where do you look first?Shared mutable state. Check whether workers share a database, schema, working directory, port, or fixture account, and whether the seed data has rows that two tasks both mutate. Also check provider rate limits and timeouts — throttling under concurrency causes tool errors that the agent handles badly and the harness scores as task failure. Both present as nondeterminism; the fix is per-worker isolation plus bounded concurrency.
- How do you turn a real production interaction into a repeatable harness task?Capture the world, not just the conversation. You need the initial state the agent acted on — the account, order and configuration rows — as a seedable fixture, plus the responses each external tool returned, recorded so they can be replayed. Then write a verifier for the desired end state. If the initial state cannot be reconstructed, the trace can inform a hand-authored task but cannot be replayed as one.
saying these in an interview costs you the question
- Sharing one database or workspace across parallel rollout workers
- Cleaning up only on the success path, so timeouts leak state
- Seeding only clean, well-formed data that production never looks like
- Asserting on timestamps or generated IDs without pinning them
- Assuming a reset works because the suite is green