skip to content

A long agent red-team campaign is producing more trace data than your storage and retention budget allows. How do you decide what the harness stops capturing, and which reductions are safe?

level: principalimportance: should knowfreq 30%

answer

  1. non-capture is irreversible, deletion is not
  2. cut recoverable bytes, never evidence
  3. dedupe and truncate giant tool payloads
  4. never sample attempts or drop non-hits
  5. two retention clocks: raw short, derived long

basics

~20 s

Capture fully by default and cut elsewhere. You can always delete a trace later; you can never reconstruct one you did not record, and a failed rerun is what kills a finding. Trim by truncating giant tool payloads to a hash plus a head, and by shortening retention, not by sampling which attempts you trace.

solid answer

~50 s

The asymmetry drives the decision: deletion is reversible in planning terms, non-capture is not. So the default is full capture, and the levers are applied in order. **Safe cuts.** Shorten retention on the raw store while keeping derived traces for reported findings. Deduplicate identical environment payloads by content hash. Truncate very large tool results to a head, a tail, a length and a hash — enough to prove what came back without storing it twice per attempt. Compress; traces are highly repetitive text. **Dangerous cuts.** Sampling which *attempts* get traced, because you cannot know which attempt becomes the finding until after it runs. Dropping failed attempts, your denominator for any success rate. Dropping tool results, the evidentiary payload itself. The honest framing: trace volume is the cost of having findings rather than opinions. If it must come down, run fewer attempts against a sharper hypothesis rather than record less of each one.

go deeper

for a junior

Recognises that traces cost storage and that deleting them later is possible.

for a middle

Distinguishes recoverable derivatives from unrecoverable evidence and proposes truncation and deduplication over dropping steps.

for a senior

Measures bytes per field before cutting, sets separate raw and derived retention clocks, and refuses per-attempt sampling with the denominator argument.

for a principal

Frames it as scope versus fidelity to stakeholders, writes the retention policy into the engagement rules before the run, and audits that reported findings still have their evidence.

This is a capacity question dressed as a storage question, and the trap is treating all bytes as equal. Rank them before cutting any. ## Rank by recoverability - **Unrecoverable:** tool arguments and result payloads, the exact model requests and responses, run-level revisions and settings. Nothing regenerates these — the run is gone and the environment has moved on. - **Recoverable:** aggregates, indexes, dashboards, rendered transcript views, per-campaign statistics. All of that can be rebuilt from the raw trace whenever someone asks. Cut the recoverable pile first and cut it completely; never buy space by turning off capture of the first pile. ## Then cut by shape, not by attempt Two structural wins usually dominate. 1. First, the **message list**: if the harness stores the full list as sent at every step, storage is quadratic in loop length, because each step re-stores the whole growing prefix — a 25-step attempt writes its opening turns 25 times. Store the per-step delta plus a digest of the complete list and rehydrate on read. 2. Second, the **payloads**: agent traces are dominated by a handful of enormous tool results — a document dump, a directory listing, a search result set — repeated near-identically across attempts. Content-hash deduplication, plus head/tail truncation that retains the full byte length and a digest, removes the bulk while preserving both what the agent saw and the ability to detect that it changed between runs. Then compress, because traces are highly repetitive text. All of this is a categorically different operation from deciding not to trace attempt 400. ## Retention is the other real lever, and it is policy, not engineering Two clocks, set separately and written into the engagement rules before the first attempt. - **Raw traces** are short-lived because they hold live data pulled through the agent's tools. - **Derived traces** behind reported findings are small and should outlive the engagement, because a finding argued a quarter later without its trace is just a memory. A single clock guarantees one of the two failures: sensitive data lingering, or evidence ageing out from under a report someone is still acting on. ## What I would refuse, and why it is a numbers argument Per-attempt sampling, dropping non-hits, and "capture on hit only". None of these reduce what the campaign cost — the tokens and the operator hours were already spent — they reduce what it can claim. - **Attack success rate** is hits over attempts; if the store holds only hits, the denominator has to come from a counter nobody validates, and the natural mistake is to recompute the rate from what is in the store, which returns 100%. - **Near-misses** are the other casualty: they are the signal that tells a defender whether a mitigation moved anything, so a campaign that keeps only successes cannot show a trend at all, before or after a fix. ## How the storage numbers themselves mislead - "Bytes per attempt" is almost always a mean dragged upward by a few giant payloads, so a deduplication that halves the mean can leave the p99 attempt — the one that fails a write or times out mid-flush — completely untouched. Look at the distribution and at the share of storage held by the top few payload hashes, not at the average. - **Compression ratio** is similarly seductive: an 8× ratio reads like the problem is solved, while the risk actually being carried is the raw store's *retention*, which compression does not shorten by a single day. - And a total that falls after a policy change can mean either that the reductions worked or that the harness quietly started dropping writes; on the graph those two look identical, which is why write errors must be counted rather than swallowed. ## Escalation If the safe levers are exhausted and the budget still does not fit, the conversation is about scope, not fidelity: fewer targets, a tighter hypothesis, shorter loops, fewer attempts per hypothesis. Say plainly that reducing capture fidelity does not save the campaign money, it reduces the campaign's output — and that a finding nobody can rerun tends to be argued away in the read-out. ## What I would check monthly - Bytes per attempt broken down by field; - the share of storage held by the top few payload hashes; - whether write failures are counted; - and — the one that actually hurts — an audit that every reported finding's derived trace still exists and still opens. Then recompute the campaign's success rate directly from the trace store and compare it with the number printed in the report. If the two disagree, something was dropped that the report is still leaning on.

  • Why keep traces for failed attempts at all?
    They are the denominator. Without them you can say a behaviour occurred but not how often, and you lose the near-misses that show whether a defensive change actually moved anything.
  • How would you truncate an enormous tool result without breaking the finding?
    Store a head and tail, the full byte length and a content hash, and deduplicate identical payloads across attempts. You keep proof of exactly what came back and can still detect drift between runs.

Tracing only the attempts that succeeded is a batting average that counts only the hits: the number climbs to 100% and stops carrying any information.

saying these in an interview costs you the question

  • Sampling which attempts get traced, when you cannot know in advance which one becomes the finding.
  • Capturing only on a hit, which erases the denominator and the near-misses.
  • Dropping tool results as the first cut because they are the biggest field.
  • One retention clock for raw and derived traces, so either sensitive data lingers or findings lose their evidence.
  • Presenting reduced fidelity as a cost saving rather than as a reduction in what the campaign can claim.
  • No measurement of where the bytes actually go before deciding what to cut.

context