After a live-fire agent red-team run in which the agent really sent mail and created records, how do you establish exactly what the test left behind and remove it — and what residue can you not remove?
answer
- ledger at interception, not from transcript
- capture returned ids, not just requests
- verify by marker search, re-search after
- soft delete is still residue
- irreversible: read mail, webhooks, audit logs
basics
~20 sDo not reconstruct it afterwards. Have the harness log every outbound tool call with the identifier the system returned, so the run produces a delete list. Tag artefacts with a per-run marker, delete, then re-search until the marker returns nothing. Delivered mail, webhook fan-out and audit entries stay.
solid answer
~60 sCleanup is designed into the instrumentation, not attempted from the transcript. **The ledger.** Wrap every side-effecting tool so the harness records the call, its arguments and the identifier the target returned — message id, record id, ticket key. That list is the cleanup plan; reconstructing it from a transcript misses anything created indirectly. **Markers.** Every artefact carries a per-run token, so you verify by search rather than by trusting the ledger. Delete, then re-search; an empty result is your evidence of a clean exit. **Order.** Delete children before parents, and remember soft deletes: a record sitting in a trash or recycle state is still discoverable, and still residue until purged. **What survives.** Mail a human already opened, webhook fan-out into systems you never provisioned, downstream syncs and search indexes, notifications already fired, and audit or SIEM entries — which you should not delete even if you could, since they are the defenders' record of your test. The deliverable is a residue statement handed to the system owner: created, removed, what remains and why.
code
json · 1 line{"run":"rt-7f3a91","tool":"tickets.create","system":"tracker-sandbox","args_hash":"9c1e...","returned_id":"SBX-4412","marker":"rt-7f3a91","ts":"2026-05-04T11:02:18Z","cleanup":"pending"}go deeper
Knows test artefacts must be deleted afterwards and that a per-run tag makes them findable.
Builds the ledger at interception time with returned identifiers, and verifies removal by re-searching the marker.
Handles soft deletes, indirect artefacts and replicas, names the irreversible residue, and hands the owner a written residue statement.
Makes cleanup cost an input to the live-fire decision and standardises the residue statement and disclosure to the blue team across engagements.
Two things separate a professional live-fire run from a mess: you knew what you were creating while you created it, and you told the owner what stayed behind. ### Build the ledger at interception time The wrapper that already sits in front of each side-effecting tool — the same object that records the call for your trace — is where the cleanup data comes from. For every call it should append one line containing the run id, the tool name, the target system, the arguments (or a hash of them if they carry seeded content), the **identifier the target system returned** — message id, record id, ticket key, object version — the timestamp, and a cleanup status field. Append and flush per call rather than buffering, so a crashed or aborted run still leaves a usable list. The returned identifier is the load-bearing field and the one most often dropped. Arguments prove the agent *attempted* something; only the id lets you delete the thing. Without it, cleanup falls back to searching by content and timestamp, which collides with genuine traffic created in the same window and misses anything the target renamed, transformed or re-filed. ### Verify by search, not by list The ledger is a plan, not a proof, because it only contains what passed through your wrapper. Indirect artefacts never enter it: a record whose creation fired an automation rule that created two more, a message auto-filed into a shared folder, an attachment copied by a retention job, a row replicated into a warehouse. That is what the per-run marker is for. Put the token in every free-text field that will take it, then after deletion search *every* system for the token — including ones you do not believe you touched — and treat an empty result as the evidence of a clean exit. Re-run that search days later; replicas, indexes and ETL jobs surface late. ### Know your delete semantics "Deleted" is a claim about a specific system's behaviour, not a universal one. Soft deletes leave the row visible to search and to eDiscovery. Mailbox deletion typically moves items into a recoverable-items store that survives the visible delete. Object stores keep prior versions and delete markers. Order matters too: delete children before parents or you orphan records you can no longer address. Verify each purge with the same marker search that found the artefact in the first place. ### What it costs Cleanup is measured in hours of a named person's time, not compute, and it lands after the interesting work is finished — which is why it is the step that silently does not happen. Budget the deletion pass, the verification search, at least one delayed re-search, and the time to write the residue statement. There is a second bill you do not pay yourself: a live-fire run generates alerts, and an undisclosed run costs the blue team an investigation. Disclose run windows, source identities and markers so the noise is attributable rather than chased. ### Where the number misleads The seductive number is "47 of 47 artefacts removed". Its denominator is the ledger, and the ledger's denominator is only what your wrapper saw. Perfect cleanup against the ledger is fully consistent with residue sitting in three systems that automation reached on your behalf. The marker search is the independent denominator — and even it is bounded by the systems you thought to search, which is why the residue statement names the systems searched rather than claiming completeness. A second misreading: an empty search *immediately* after deletion is often just index lag or a soft-delete window. An empty search a week later means considerably more. ### Accept and disclose the irreversible Some residue is permanent by design, and disclosing it is strictly better than implying it is gone: mail a human already opened, SMS or push already delivered, webhooks that fanned out to third parties, rows already replicated into a warehouse or search index, notifications that already paged someone, and the prompts and outputs recorded by whatever model provider served the target. Two deserve special handling. Anything reaching a real third party can trigger someone else's incident response, so it belongs in the pre-run stub decision, not in the cleanup plan. And audit, SIEM and detection entries should be left intact even where you could remove them — they are the defenders' record of your test, and deleting them destroys the reconciliation the blue team needs. Close the loop with a residue statement to the system owner: what was created, what was removed, the exact search that confirmed it, what remains and why, and who to contact if it resurfaces. If that list looks unacceptable *before* the run, that is the signal to stub the tool instead — cleanup difficulty is an input to the live-fire decision, not an afterthought.
- Why capture the identifier the target system returns, not just the arguments the agent sent?The arguments prove the attempt; the returned id is what lets you delete the artefact. Without it, cleanup degrades to searching by content and timestamp, which collides with real traffic and misses anything the system renamed or transformed.
- Should you clean up entries your run created in the defenders' SIEM?No. Those are the blue team's record of your activity and they need them to reconcile alerts. Disclose your run windows, source identities and markers so they can attribute the noise, and leave the logs alone.
- You re-search for the run marker a week later and it appears in a data warehouse you never touched. What happened?A downstream sync or ETL copied the record before you deleted the original. It shows why deletion has to be verified across replicas and indexes, and why late re-searches belong in the plan.
saying these in an interview costs you the question
- Planning to reconstruct what was created by reading the transcript afterwards.
- Recording tool arguments but not the identifiers the target system returned.
- Calling it clean without re-running the marker search.
- Deleting or asking to delete audit and SIEM entries to hide test activity.
- Discovering only after the run that a tool fanned out to a third-party system.