skip to content

Conversation Memory

The store that makes a run resumable and auditable also accumulates harmful text and target responses. Interviewers ask where it lived and who could read it.

on this pageshow

explore

questions

5

In PyRIT, what does the conversation memory store record while a run executes, and what do you give up by running with the in-memory store instead of the durable one?

level: juniorimportance: must knowfreq 62%

answer

  1. memory = the run's transcript log
  2. turns plus attached verdicts
  3. in-memory dies with the process
  4. no store, no evidence, no resume
  5. durable store is the artefact

basics

~20 s

PyRIT's memory records every prompt sent and every response received, turn by turn, with the scores attached, so a run can be resumed, audited and re-scored later. The in-memory option keeps that only for the process lifetime: when it exits, the transcripts are gone and nothing can be re-read or re-scored.

solid answer

~50 s

Memory is the run's transcript log. Every request PyRIT sends to a target and every response that comes back is written as a turn, tagged with the conversation it belongs to, and the verdicts produced for it are attached to those turns. That is what makes a multi-turn run coherent (the next turn is built from what is already stored), resumable after a crash, and auditable afterwards. The in-memory variant implements the same interface but keeps nothing past the process. It is right for a smoke test, a unit test, or a demo where you deliberately do not want harmful text landing on disk. It is wrong for anything you will report on: a run that produced a hit and then exited leaves you with console output and no evidence, no ability to relabel with a different scorer, and no ability to resume. The choice is therefore not a performance knob — it decides whether the run produces an artefact at all.

go deeper

for a junior

Should say memory holds the prompts and responses of the run and that in-memory loses them when the process ends.

for a middle

Adds why it matters — resume after a failure, re-read the transcripts, and the fact that verdicts are attached to stored turns.

for a senior

Frames the choice as evidence handling: the durable store is both the only artefact you can report from and a file of harmful text you now own.

for a principal

Treats it as policy: which runs are allowed to persist, where those files may live, and how the team's tooling defaults so nobody picks in-memory by accident on a real engagement.

### What memory actually is in the object model PyRIT does not thread a log object down the call chain. It keeps **one process-wide memory instance**, held by `CentralMemory` and chosen once at start-up — in current releases through `pyrit.common.initialize_pyrit(memory_db_type=...)`, which takes a store kind such as the in-memory one or the durable file-backed one. Every other execution object — a prompt target, an attack strategy, a converter chain, a scorer — resolves that same instance through `CentralMemory.get_memory_instance()`. Memory is therefore *ambient* rather than wired: flipping the store kind on that single line changes the behaviour of every object in the run without any of them being reconfigured. That is precisely why the wrong choice is easy to make in a copied notebook and hard to notice until the run is over. ### What a stored record contains The unit is not a conversation but a **request piece**: one prompt, or one response. A piece carries the `conversation_id` it belongs to, its role (`user` or `assistant`), a sequence number ordering it within that conversation, the original value, the **converted value** — the string as actually sent after the converter chain ran, which is usually not the seed you typed — an identifier for the target it went to, an identifier for the attack that produced it, and timestamps. Scores are stored as their own records that reference a piece, which is why a verdict is inseparable from the exchange it judged. "The transcript" is a reconstruction: `memory.get_conversation(conversation_id=...)` pulls the pieces and orders them by sequence, and that reconstruction is exactly what a multi-turn strategy feeds back in to compose its next prompt. ### The durable store versus the in-memory one Both implement the same interface, so nothing downstream can tell them apart — which is the point, and the trap. The durable store writes into an embedded database file on the host (the engine and the default path have changed across PyRIT releases; check what your installed version writes rather than trusting a tutorial). The in-memory store keeps the same tables for the life of the process, and then they are gone: no file, no export, no re-read, no resume. ### What it costs Storage is trivial. A piece is a few kilobytes of text; a forty-objective run at ten turns each is roughly 800 pieces and single-digit megabytes, and the write is far below the latency of the model call that produced it. The cost that matters is the **cost of not having it**. A multi-turn red-team run bills three metered calls per turn — the attacker model that composes the next prompt, the target, and a model-backed scorer — so that same forty-by-ten run is on the order of 1,200 calls, plus its wall-clock, plus a second window of live attack traffic against someone's production endpoint if you have to repeat it. Losing the store converts all of that into a re-run, and re-runs against a customer system usually need fresh authorisation, not just fresh budget. ### Where the number misleads Two readings go wrong routinely. - **Row counts are not attempt counts.** A multi-turn run stores many pieces per objective, so "the store holds 812 rows" is not "we ran 812 attacks" — it is roughly forty objectives times ten turns times two roles. Any success rate whose denominator came from a raw piece count is wrong by the turn multiplier. Count distinct `conversation_id` values, or objectives, and say in the report which one you counted. - **A reused `conversation_id` silently merges runs.** Because history is reconstructed from whatever is stored under that identifier, two runs sharing one will each see the other's turns: the strategy builds on a conversation it never actually had, and the exported transcript reads as a single long coherent attack that never happened as one exchange. The in-memory store produces the mirror-image failure — every process starts blank, so a strategy expecting prior turns quietly behaves as though it were on turn one. ### What to check before a run you care about Print the configured store kind from the entrypoint itself rather than trusting the snippet you copied. Then do a two-turn throwaway against a harmless objective and read it back: confirm the pieces are there, that the converted value is populated and not just the seed, that the score records attached to those pieces, that the file landed on the path you intended and the process can write there, and that whatever export you plan to hand to the report actually runs end to end. Doing that check after the engagement, when the only possible answer is bad news, is the mistake this question is really about.

  • You ran a long multi-turn engagement and the process was killed halfway. What do you actually still have?
    With the durable store, every turn completed before the kill, so you can resume or at least report on what was reached. With the in-memory store, nothing but whatever scrolled past in the terminal.
  • When is the in-memory store the right choice?
    Tests, CI smoke runs and demos — anywhere you want the plumbing exercised without leaving harmful prompts and responses on disk.
  • Where do a run's verdicts live?
    Attached to the stored turns rather than only in the console, which is what lets you read back later which exchange was judged a hit and why.

saying these in an interview costs you the question

  • Describing memory as a cache or a performance optimisation rather than the run's record.
  • Assuming console output is enough to report a finding.
  • Not knowing that the in-memory option exists, or thinking it merely speeds things up.
  • Claiming the transcripts can be reconstructed from the target afterwards.

context

open as a page

By default a PyRIT run persists its conversations to an unencrypted local database file. What is actually in that file, and how should it change where and how you run an engagement?

level: middleimportance: must knowfreq 58%

basics

~20 s

By default PyRIT writes the run to a local database file on the machine you launched it from, unencrypted. That file holds attack prompts and the target's worst answers verbatim. Treat it as sensitive evidence: put it on encrypted storage you control, keep it off shared drives and backups, and delete it on schedule.

open as a page

An application team disputes one item in your report from a PyRIT engagement. How do you get from that report item back to the exact stored exchange that produced it, and what does the transcript prove and not prove?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Each stored turn carries identifiers tying it to its conversation and run, so a report item can point at the exact exchange that produced it. Keep those identifiers in the finding. The transcript proves what happened once against that configuration - not that it reproduces, since the target samples and may have changed.

open as a page

A finished PyRIT engagement is stored, and you now want to relabel it with a stricter judgment of what counts as a hit without spending another call against the target. What makes that possible, and what can relabelling stored transcripts not tell you?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Because the run stored every prompt and response, you can point a different scorer at the saved transcripts and relabel them without spending another target call. What that cannot tell you is what the attack would have done under the new judgment: the strategy branched on the old verdicts, so the turns it never explored simply do not exist.

open as a page

Your team is scaling from one operator running PyRIT locally to several running in parallel on the same engagement. How do you decide between per-operator local stores and one shared conversation store, and what does the shared choice oblige you to build?

level: principalimportance: should knowfreq 33%

basics

~20 s

A local per-operator store is simplest and keeps harmful transcripts on one controlled machine. A shared store lets a team resume each other's runs, deduplicate findings and report across the whole engagement, but concentrates every operator's harmful text in one place that now needs access control, retention rules and a named owner.

open as a page