skip to content

Your agent red-team harness replays a recorded attempt with the same prompt, the same sampling settings and the same seed, and the agent takes a different action. Why is recording the seed not enough to make an agent run replayable, and what does the trace need instead?

level: middleimportance: must knowfreq 48%

answer

  1. seed covers only your sampling
  2. endpoint, environment, harness — three entropy sources
  3. record request and response bytes
  4. recorded replay vs live replay
  5. report reruns as a rate

basics

~20 s

Because the run is not a pure function of the seed. The environment answers differently each time: tool results, clocks, IDs, retries, and the endpoint's own sampling all vary. Record the actual request and response bytes at every step so a replay can fall back to the recorded exchange.

solid answer

~50 s

A seed only constrains the sampling you control. An agent loop has at least three other entropy sources: the **model endpoint**, which you generally cannot assume decodes identically across calls even with fixed settings; the **environment**, whose tool results carry fresh IDs, timestamps, changing records and occasional errors; and the **harness itself**, through retries, timeouts, truncation and any concurrency in parallel tool calls. So the trace has to be able to stand in for the world. Record, per step, the exact request sent and the exact response received, both for the model and for each tool. That gives you two replay modes: a **live** replay, which re-executes everything and is what you use to confirm the finding still reproduces, and a **recorded** replay, which feeds the stored responses back and is what proves what happened in the run you are reporting. Where the two disagree, the finding is flaky, and that is itself a fact worth reporting — with the number of live reruns behind it.

go deeper

for a junior

Knows a seed is recorded and that runs still vary; may not be able to name where the variance comes from.

for a middle

Separates the three entropy sources and explains why the trace must hold raw requests and responses rather than a seed.

for a senior

Runs recorded and live replay as distinct checks, reports a rerun rate, and hunts leftover environment state behind drifting results.

for a principal

Sets the expectation that every reported agent finding carries a rerun count and a stated replay mode, so nobody ships an anecdote.

**Why the seed argument fails.** A seed constrains the pseudo-random sampling *you* control inside a single generation call. An agent run is not one call; it is a loop over a live system, and a loop amplifies divergence. One differently-worded argument at step two changes the tool result at step two, which changes the messages at step three, and by step six the two runs are not comparable at all. Reproducibility therefore has to be engineered into the trace. It cannot be borrowed from a seed. **Where the entropy actually enters.** Three sources, worth separating because the fixes differ. The **endpoint under test**: a hosted model is a shared service, and identical settings do not guarantee identical decoding — batching, expert routing, kernel and hardware differences, and silent version rolls all move the output. Temperature zero narrows the variance; it does not remove it, and a seed parameter, where the API offers one, is best-effort rather than a contract. The **environment**: tool results carry fresh identifiers, current timestamps, listings whose order is not stable, quota and rate-limit errors, and state left behind by earlier attempts in the same campaign. The **harness**: retries on timeout, truncation at a token cap, and — when the model emits several tool calls at once — a resolution order that follows latency. Any one of these flips a borderline attempt from hit to miss. **What the trace records instead.** Per model call: the message list as sent, the settings and seed, the raw response including the tool-call structures, and the finish reason. Per tool call: the arguments, the result payload, latency, and any error and retry count. Per run: harness and environment revision, and the initial state of anything the tools read. Recording raw requests and responses — not a parsed summary of them — is what buys two distinct replay modes. | mode | what it re-executes | what it proves | what it costs | |---|---|---|---| | recorded | nothing; stored responses are served back in order | the reported transcript is complete and internally consistent — the behaviour happened as written | near zero: no tokens, no side effects | | live | the model and every tool, against the real target | the behaviour still happens now | one full attempt per rerun, plus containment and cleanup | **What it costs.** Live reruns are the expensive half, and the cost is superlinear in loop length because an agent resends the whole growing transcript at every step: a fifteen-step attempt can bill six figures of prompt tokens even though the operator typed one paragraph. Ten reruns is that ten times over, plus ten more sets of side effects to contain and clean up, plus the wall-clock of a loop that waits on real tools. That economics is exactly why teams rerun once, get a hit, and write "reproducible". **How the number misleads.** Two readings go wrong, and the second is the dangerous one. First, a rerun count with no denominator: "reproduced" on the strength of one successful repeat is a coin landing heads twice. If the behaviour fires on three of ten live attempts, the honest finding is "intermittent, 3/10" — still real, still worth fixing, and materially different from "the agent always does this", which is what a reader takes from an unqualified claim. Second, the replay mode gets lost between operator and reader. A recorded replay is green *by construction*: it serves stored responses, so it cannot fail for the reason anyone cares about, and it will pass ten times out of ten forever. Demonstrate it to a stakeholder without naming the mode and they will conclude the issue is live today, which the run never tested. The inverse trap is just as common: after a vendor rolls a model or tightens a filter, a 0/10 live rerun does not mean the finding was wrong — it means the world moved, and the recorded trace is what preserves the original claim. **What I would check.** Confirm the "live" rerun was actually live: count outbound requests, or read the endpoint's own usage counter for that window. A live rerun that consumed no tokens was a recorded one. Deliberately corrupt one stored response and confirm the recorded replay diverges rather than sailing past, which tells you it is genuinely following the trace and not re-deriving anything. Reset or snapshot the environment between reruns so leftover state is not quietly doing the work. Then put the rate and the mode in the finding text itself, not in the operator's memory.

  • What does a recorded replay actually prove, given nothing is re-executed?
    That the transcript you are reporting is internally consistent and complete enough to walk through — it proves what happened, not that it still happens. Pair it with live reruns.
  • Attempts get more successful the longer the campaign runs. What environmental cause would you check first?
    State accumulated by earlier attempts — records, files or memory left behind that later attempts read. Reset the environment between attempts or record its initial state so you can tell.

A recorded replay is the security-camera footage; a live rerun is walking back to the door to see whether it is still unlocked. Footage that plays perfectly tells you nothing about the door.

saying these in an interview costs you the question

  • Claiming a fixed seed makes an agent run deterministic end to end.
  • Treating one successful attempt as a reproducible finding with no rerun count.
  • Recording only the parsed tool call, so a malformed or unparsed response cannot be re-examined.
  • Blaming the model for divergence without checking the environment's fresh IDs, timestamps and leftover state.
  • Discarding retries and timeouts from the trace, hiding a whole class of divergence.

context