By default a PyRIT run persists its conversations to an unencrypted local database file. What is actually in that file, and how should it change where and how you run an engagement?
answer
- unencrypted local file
- post-conversion prompts, real responses
- synced folder / git repo / shared runner
- full-disk encryption is the only cover
- retention decided before the run
basics
~20 sBy default PyRIT writes the run to a local database file on the machine you launched it from, unencrypted. That file holds attack prompts and the target's worst answers verbatim. Treat it as sensitive evidence: put it on encrypted storage you control, keep it off shared drives and backups, and delete it on schedule.
solid answer
~50 sThe file is the run, in the clear: seed prompts, every converted variant actually sent, every target response including the ones that succeeded, and the verdicts attached to them. That is the most sensitive artefact the engagement produces - a curated set of prompts that worked, plus proof of what the system said. It lands on whatever machine ran the process, which on most engagements is a laptop that syncs to cloud backup, is indexed by desktop search, and may be shared. Nothing in the tool encrypts it; full-disk encryption on the host is the only protection unless you add your own. So the controls are operational: run from a directory you chose in advance, keep it out of synced folders and repositories, assume anyone with read access to that path can read the transcripts, and make the retention decision before the run. The corollary is that the risk exists precisely because the run went well. A run that found nothing leaves a boring file.
go deeper
Should know the run leaves a local file and that it contains the prompts and responses, so it is not something to leave lying in a repository.
Names what is stored, that the tool does not encrypt it, and gives concrete placement and deletion habits.
Connects it to engagement handling — classification, an owner, backups and search indexing as the sneaky copies, and a destruction record.
Sets the team default: where runs are executed, what the storage standard is, retention length, who may read transcripts, and how that is enforced rather than remembered.
Treat the store as the engagement's evidence locker, and note that the tool ships the locker without a lock. ### What is physically in the file The durable memory writes turn-level records into an embedded database file on the host that ran the process. Each record is a request piece: the prompt **as actually sent** — the converted value produced by the converter chain, not the sanitised seed you typed — the response as received, the conversation and attack identifiers, the target identifier, timestamps, and any labels you stamped on the run. Scores sit in their own records referencing those pieces. So the file is not a summary of an engagement; it is the engagement, verbatim, including the prompts that worked and the answers that prove they worked. If the target is a real application rather than a bare model endpoint, the responses may also carry whatever that application pulled in — retrieved documents, tool output, customer-shaped records. That is how a red-team artefact quietly becomes a data-handling problem. (Redaction as an application control belongs to a different topic; the point here is only that you now hold the material.) ### Where it lands, and what the tool does about it It lands on the machine that ran the process, at a path you either chose deliberately or inherited from a default that has moved between PyRIT releases. **Nothing in the tool encrypts it and nothing in the tool access-controls it.** Confidentiality is exactly filesystem permissions plus whatever full-disk encryption the host has. Any export helper you run to get transcripts out for the report writes a second plaintext copy wherever you pointed it, and that copy is the one people forget. ### What it costs Almost nothing to produce and almost nothing to store: a 1,200-call multi-turn run yields a file of a few megabytes. That cheapness is the hazard — a few megabytes rides along in a `git add .`, replicates through a sync client in seconds, and is small enough that no quota or backup alert ever fires. The expensive side is remediation. Deciding the path and the destruction date before kickoff costs about ten minutes. Discovering afterwards that working attack prompts against a named customer were committed to a shared repository costs a history rewrite, a credential-free but genuinely awkward customer notification, and legal time measured in days. ### Where the reading misleads - **"I deleted the file" is usually false.** A local delete does not touch the git objects created when it was committed (a plain `git rm` leaves it in history), the versions retained by a sync provider, last night's backup snapshot, or the desktop search index that may hold extracted text. Deletion is a plural operation, and only the first copy is easy. - **"The run found nothing, so the file is harmless" is false.** The prompts are the reusable half. A zero-hit run still leaves a curated set of attack strings aimed at a named system, which is the part an outsider would actually want. - **File size tells you nothing about sensitivity.** Reviewers routinely eyeball a small artefact as tool exhaust. The most sensitive file the engagement produces is also the smallest. - **"We moved to a hosted backend, so it is handled" is a half-truth.** A shared or remote store replaces filesystem permissions with an access-control system, which is progress only if someone configured it, and it concentrates several operators' transcripts in one place. ### What to check Before the engagement: run a two-turn throwaway and list the directory before and after, so you know by observation which file the run creates rather than by assumption. Then answer four questions about that exact path — is it inside a git work tree (`git check-ignore` on it, and add it to the ignore file), inside a sync root, inside a backup set, inside a shared-runner workspace? Confirm the host has full-disk encryption actually enabled, not merely available. Prefer a dedicated engagement directory outside every project checkout. Decide retention up front: a destruction date, or a move into whatever holds the deliverable evidence, and a note recording which you did. When you quote transcripts into a document that circulates more widely than the store, quote only what the finding needs — the report is a copy too.
- Why is a successful run more sensitive than a failed one?Because the file then contains prompts that are known to work against a specific system, plus the responses proving it — that is both an attack recipe and a disclosure of the target's behaviour.
- Your engagement ran on a shared CI runner. What is the concern?The workspace is readable by whatever job runs next and by anyone with runner access, and the file may survive in caches or artefacts. Engagement transcripts should not be produced on shared infrastructure without an isolation and cleanup story.
- Does moving to a hosted backend fix the problem?It changes it. You gain access control and central retention, but you have concentrated harmful transcripts in a system with more readers, so the question becomes who is entitled to query it.
saying these in an interview costs you the question
- Believing the store is encrypted or access-controlled by the tool.
- Running engagements from a project checkout or a synced cloud folder without noticing.
- No retention or destruction plan for the transcripts.
- Treating the file as tool exhaust rather than the engagement's most sensitive artefact.
- Claiming a hosted backend removes the exposure instead of moving it behind access control.