skip to content

An automated LLM red-team scanner writes every request and response, including successful harmful completions, to log files in its working directory. Where do those files realistically end up during an engagement, and what do you change before the first run?

level: seniorimportance: should knowfreq 44%

answer

  1. tool logs are unredacted by construction
  2. CI artefacts + console = repo-wide read
  3. self-hosted runner workspace persists
  4. encrypted, unsynced, per-engagement dir
  5. destruction must include backups

basics

~20 s

They spread: CI job artefacts and console output, the runner's disk after the job, ticket attachments, screenshots in chat, cloud-synced home folders, laptop backups. Before the first run, move the working directory onto encrypted, unsynced storage in a per-engagement folder, stop artefact upload or restrict and expire it, and schedule destruction.

solid answer

~50 s

The report is usually the artefact people control; the tool's own logs are the one nobody thinks about, and they contain the unredacted material by construction. Realistic paths: the CI job that ran the sweep uploads its whole workspace as artefacts and prints failures to a console log both readable by everyone with repository access; the self-hosted runner keeps the workspace until something cleans it; an engineer attaches the log to a ticket to show the failure; someone screenshots it into chat; the working directory sits under a cloud-synced home folder and replicates to a personal device; backups then hold copies past any retention date you set. Before the first run: dedicated per-engagement directory on encrypted, unsynced storage; artefact upload off or scoped to a private bucket with a lifecycle expiry; nothing harmful in job names, console output or test titles; tool telemetry and auto-upload disabled; a written destruction step at engagement close, including backups and runner workspaces.

go deeper

for a junior

Knows the scanner writes raw prompts and responses to disk and that those files must not sit in a shared or synced location.

for a middle

Lists the main copies — CI artefacts, runner workspace, ticket attachments, sync — and sets up an encrypted per-engagement directory with artefact upload off.

for a senior

Treats it as a copy-enumeration problem, closes each path with a default-safe control, keeps harmful text out of job names and console output, and makes destruction cover backups and runners.

for a principal

Makes the safe setup the standard harness rather than a per-engagement discipline: a hardened runner image, storage policy, retention automation, and an audit that samples pipelines for leaked artefacts.

This is the gap between a handling *policy* and handling *practice*. The policy governs the deliverable, which is the artefact everyone remembers. The practice has to govern what the toolchain writes on its own — unredacted by construction, voluminous, and produced far faster than anyone reads it. ## The scanner writes raw material by default, and it has to Every mature LLM red-team runner keeps the prompts and the responses, because a hit you cannot open is a hit you cannot triage. Expect four artefacts, all outside your project folder unless you moved them: a **run report**, typically one JSONL record per attempt carrying the prompt sent, the response received and the detector's verdict; a **hit log**, the same shape filtered to successes, which is to say a file consisting entirely of working attacks and harmful outputs; a **persisted conversation store**, often an embedded database file under a per-user data directory, written so a later triage pass can reopen multi-turn sessions — meaning "the process exited" does not mean the conversation is gone; and frequently an optional **publish-to-hosted-viewer** command that uploads a whole evaluation, outputs included, unless sharing is disabled in configuration. None of this is a defect. It is the feature that makes the tool usable. It simply means the unredacted material exists in locations your report-handling policy has never mentioned, written at machine speed, before anyone has decided what harm class it falls into. ## Enumerate the copies, in the order they actually bite 1. **CI.** If the sweep runs in a pipeline, the workspace becomes uploaded build artefacts and the console becomes a searchable log. Both typically inherit repository-level read access, which in a mid-sized org is dozens to hundreds of people, none of them on the engagement's access list. 2. **The runner.** On a self-hosted runner the workspace persists after the job unless something removes it. Assume it persists until you have verified the cleanup, not the other way round. 3. **Tickets and chat.** The fastest route a completion ever takes out of your control is an engineer pasting a failing log into an issue so somebody will look at it. 4. **Sync and backup.** A working directory under a synced home folder replicates to other devices within seconds; backups then hold copies that outlive any deletion you perform later. 5. **Egress from the tool.** A hosted results dashboard, a share command, or usage telemetry moves the material off your control plane entirely. 6. **The far end.** Whatever you sent has left your estate. Whether the endpoint operator logs, retains or human-reviews it is a scoping and authorisation question to settle before the first request, not to discover afterwards. ## Controls, chosen so the default is safe Put the working directory on an encrypted volume outside every synced path, one directory per engagement, with the retention date in the directory name so an expired directory is visible without opening a policy document. Turn artefact upload off; if a pipeline genuinely needs evidence, emit only the redacted summary as its artefact and push the raw log to a restricted bucket with object-level access, an access log and a lifecycle expiry. Keep harmful text out of anything that becomes a label — job names, test titles, notification text and dashboard tiles all propagate to places with no access control at all. Disable telemetry and hosted-result upload explicitly rather than assuming the default. Write destruction into the engagement close checklist and make it cover runner workspaces and backup generations. ## What it costs, and where the numbers mislead The hardening is roughly a day of setup for the first engagement and an hour thereafter, plus a real ongoing cost people underestimate: with artefact upload off, a failing sweep in CI is now opaque, and triage moves to someone with access to the restricted store. Expect to pay that in slower debugging, and design for it by emitting a redacted failure summary rather than nothing. Three numbers mislead here. **Artefact retention defaults** are set by the CI platform, commonly on the order of months, and they are measured from the job, not from your engagement — so a store you believed was emptied at close quietly holds the payload well past the retention date you promised the client. **"The repository is private"** is read as a small number; it is a permission boundary, not an access list, and it grows with every new hire. **"We deleted it"** counts the copies you know about: log rotation compresses rather than deletes, a database row removed from a memory store often leaves the file's pages untouched, and a backup generation is a copy that no deletion of yours reaches. ## What I would check at close The runner workspace, on the machine, by hand. The artefact retention setting on the pipeline. The ticket tracker, searched for attachments. Chat history for pasted logs and screenshots. The tool's own data directory and memory store, not just the project folder. And whether any screenshot of a log made its way into a slide deck, which is the copy no scrubber will ever find.

  • The client wants these sweeps to run continuously in their own pipeline. What is the minimum change to the pipeline's output?
    The pipeline emits only pass/fail plus a redacted summary as its artefact; raw request/response logs go to a restricted store with object-level access, an access log and a lifecycle expiry.
  • Why is 'we encrypt the disk' an incomplete answer here?
    Disk encryption protects a powered-off device. It does nothing about copies made by a logged-in process — artefact upload, sync, backup, or a person attaching the file to a ticket.

saying these in an interview costs you the question

  • Assuming CI artefacts are private because the repository is private
  • Deleting the working directory at close while backups and the runner workspace still hold copies
  • Putting a harmful excerpt in a test name or job title, where it lands in notifications and dashboards
  • Never considering that a hosted results dashboard or telemetry moves the material off-premises

context