skip to content

During an agent red-team run the agent's tools pulled real customer records and a live API token into the harness trace. How do you design trace capture so the run stays shareable without destroying the reproducibility of the finding?

level: seniorimportance: should knowfreq 44%

answer

  1. two artefacts: raw evidence, derived shareable
  2. stable placeholder preserves the linkage
  3. secret: rotate first, store hash not value
  4. free-text tool results defeat key-based redaction
  5. retention clock agreed before the run

basics

~20 s

Split raw from shareable. Keep the unredacted trace in a restricted store with a short retention, and derive a redacted copy by replacing secrets and personal data with stable placeholders. Stable means the same value maps to the same token, so the reader still sees the causal chain.

solid answer

~50 s

Treat the trace as two artefacts with different audiences. The **raw trace** is evidence: it lives in a store scoped to the engagement, with access limited to the operators and a retention clock agreed before the run. The **derived trace** is what goes into the finding, produced by a redaction pass, never by hand-editing. The redaction must be **consistent**: every occurrence of a given value maps to the same placeholder, so a reader can still see that the address the agent exfiltrated at step 7 is the address it read at step 3. Random per-occurrence masking destroys exactly the linkage the finding rests on. Secrets get stricter treatment than personal data: a leaked token should be rotated and recorded only as a hash and a prefix, because the value's presence is the finding and its content is not. Where it breaks: free-text tool results. A key-based redactor misses a customer name inside returned prose, so the derived trace needs a human review pass before it leaves the engagement.

go deeper

for a junior

Knows the trace may hold sensitive data and that it must be redacted before sharing.

for a middle

Separates the raw and derived copies and explains why placeholders must be consistent across the trace.

for a senior

Treats secrets as an exposure to rotate, keeps the mapping in the restricted store, and adds a human review pass for free-text results; agrees retention before the run.

for a principal

Makes capture-and-redaction policy a precondition of authorisation, so no engagement collects data whose handling rules are still undecided.

**The conflict, stated precisely.** What makes an agent finding credible — verbatim tool arguments and verbatim tool results — is the same material that makes the trace dangerous to hold. The wrong resolution is to capture less, because a thinner trace does not become a safe artefact; it becomes an unusable one that still contains whatever it did capture. Resolve it by splitting artefacts, not by weakening capture. **Raw store.** Written by the harness, unredacted, access-controlled to the named operators on the engagement, with a retention clock enforced by the store itself — a lifecycle rule — rather than by anybody's intention. Treat it as in scope for the same handling rules as any other authorised production data extract: if the agent's tools read customer records, the trace *is* a customer-record extract, whatever the ticket calls it. **Derived store.** Produced by a deterministic transform over the raw trace, never by hand-editing. A hand-sanitised copy no longer matches the evidence and cannot be regenerated when a reviewer asks a follow-up question three weeks later. ``` email / name / account id -> stable_token(value, engagement_key) # <EMAIL_3>; same value, same token live secret -> "<SECRET sha256:ab12… len=40>" # presence and shape, never content ``` **Why stability is the whole game.** The finding is usually that a value the agent *read* at one step reappeared in an argument it *sent* at a later step. Consistent placeholders keep that edge visible: `<EMAIL_3>` in a step-3 tool result and `<EMAIL_3>` in a step-7 send argument is the proof of exfiltration. Per-occurrence random masking is, if anything, stricter on confidentiality — and it silently destroys the proof, degrading the shared trace from a demonstration into an assertion. Nobody notices, because it still looks like a trace. Derive tokens with a keyed function and keep the key and the mapping in the raw store, scoped per engagement: a mapping filed beside the derived copy un-redacts it, and a token stable *across* engagements lets two reports be correlated back to one real person. **Secrets are not personal data.** A live credential in a trace is an active exposure, so the first action is rotation, not redaction; only then does the trace keep a hash, a length and perhaps a prefix. Better still, keep credentials out of the trace at source — have the tool layer resolve a reference at call time so the raw store never holds the value. The finding is that a token reached the agent, not what the token was. **What it costs.** The redactor is a day or two of engineering and then near-free per run. The expensive line is the human review pass, at minutes per trace, and it does not scale to a four-hundred-attempt campaign — so scope it deliberately: every trace attached to a reported finding is read end to end by a second operator, and the rest stay in the raw store until they age out. Rotation has a cost too, borne by whoever owns the credential, which is a reason to raise it the hour you find it rather than at report time. **How the numbers mislead.** Redaction tooling reports what it removed: "1,240 items redacted" reads as assurance and has no denominator. You cannot compute recall from the redactor's own output, because the misses are by definition the values it failed to recognise. The real miss surface is free text, and tool results are mostly free text — a customer name inside a paragraph of returned prose has no field key for a key-based redactor to match on. Two cheap counter-readings exist. Compare raw and derived byte sizes: near-equality on a run whose tools returned live records means the redactor was effectively a no-op, while a derived copy dramatically smaller than the raw one means whole payloads were blanked and the finding is now unreadable. Both take about two minutes and catch the two opposite failures. **What I would check.** Plant canaries before the run — a synthetic customer record and a synthetic credential, of the same shapes as the real ones, seeded into the environment the tools read — then assert that neither ever appears in the derived trace. That converts "the redactor ran" into a measured result for at least those classes. Then have a second operator read the derived trace before it leaves the engagement, and confirm that the retention clock and the access list were agreed *before* the first attempt. That decision is cheap in advance and expensive afterwards: negotiating handling rules while already holding the data is how a red-team engagement becomes its own incident.

  • Why is a per-occurrence random mask worse than a stable placeholder here?
    Because the finding is that a value read at one step reappeared in a later tool argument. Random masks make those two look like different values, so the shared trace no longer demonstrates the exfiltration.
  • The agent's tools returned a paragraph of prose containing a customer name. Why does the redactor miss it, and what covers you?
    Field-based redaction works off known keys and free text has none. Cover it with a detector pass plus a mandatory human read of the derived trace before it leaves the engagement.

saying these in an interview costs you the question

  • Capturing less to dodge redaction, which trades a data problem for an unreproducible finding.
  • Random per-occurrence masking that destroys the read-then-send linkage the finding depends on.
  • Keeping the placeholder mapping next to the shared copy, which un-redacts it.
  • Pasting a leaked credential into the report as evidence rather than rotating it and recording a hash.
  • Hand-editing a trace to sanitise it, so the shared artefact no longer matches the raw one.
  • Leaving retention and access to be decided after the data has already been collected.

context