skip to content

Your team is scaling from one operator running PyRIT locally to several running in parallel on the same engagement. How do you decide between per-operator local stores and one shared conversation store, and what does the shared choice oblige you to build?

level: principalimportance: should knowfreq 33%

answer

  1. decide by the reporting unit
  2. join transcripts or join documents
  3. concentration trades many copies for one
  4. labels, named access, scheduled deletion
  5. per-engagement store, not perpetual

basics

~20 s

A local per-operator store is simplest and keeps harmful transcripts on one controlled machine. A shared store lets a team resume each other's runs, deduplicate findings and report across the whole engagement, but concentrates every operator's harmful text in one place that now needs access control, retention rules and a named owner.

solid answer

~60 s

Decide on the reporting unit. If each operator delivers findings separately, local stores are fine and the aggregation happens in a document. If the deliverable is one engagement-level picture — coverage across operators, deduplicated findings, a defensible attempts-and-hits count — someone will have to join the transcripts anyway, and joining after the fact from five laptops is worse than sharing from the start. What sharing obliges you to build: authentication and per-person access, because the store is now a curated corpus of working attack prompts against a named customer; a naming or labelling scheme so runs, operators and target deployments stay distinguishable; retention and destruction as a scheduled action rather than someone's memory; and an owner accountable for both. You also inherit operational concerns — concurrent writers, a single point of failure mid-engagement, and network dependence during runs. The honest framing is that centralising does not reduce exposure; it trades many small uncontrolled copies for one large controlled one. That is usually the better trade, but only if the controls are actually built.

go deeper

for a junior

Should recognise that a shared store means other people can read the transcripts, and that this needs permission rather than convenience.

for a middle

Names the practical benefits — resume a colleague's run, one place to export from — and that the file is sensitive wherever it lives.

for a senior

Weighs concentration against aggregation, and lists the controls a shared store needs: labels, named access, encryption, scheduled destruction.

for a principal

Decides from the deliverable and the client agreement, scopes the store per engagement, assigns an owner, and refuses the shared option until the controls exist.

### The mechanism is a one-line swap, which is what makes the decision easy to under-think Both the local file-backed store and a network-backed shared one implement PyRIT's memory interface, and both are selected the same way at start-up — the store kind handed to `initialize_pyrit`, resolved thereafter by every object through `CentralMemory`. No strategy, target, converter or scorer changes. Because the swap is invisible to the code, teams tend to evaluate it as a configuration choice when it is actually a decision about the engagement's reporting unit and about who may read a corpus of working attacks. ### Decide from the deliverable If each operator delivers findings separately, local stores are correct and aggregation happens in a document. If the deliverable is one engagement-level picture — coverage across operators, deduplicated findings, a defensible attempts-and-hits count with one denominator — then somebody has to join the transcripts anyway, and joining five laptops at the end is worse in every way than sharing from the first run. Multi-operator work fails at the report, not at the run: two operators unknowingly hammering the same objective while a third surface goes untested is a coverage failure that becomes visible only when the records are joined. ### What sharing costs, concretely - **Latency and failure mode.** A multi-turn strategy reconstructs conversation history from the store *on every turn*. A local file read is microseconds; a network round trip is tens of milliseconds, several times a turn. Over a campaign of thousands of turns that is minutes to tens of minutes of added wall-clock, and worse, a network blip now fails runs that a local file would have survived — and it fails every operator at once, which a set of independent laptops does not. - **Money and engineer time.** The hosted database itself is small change next to model spend. The real spend is human: initial setup, an access review, a destruction date someone owns, and a repeat of that per engagement — days, not hours, and recurring. - **Concentration.** One store holding every operator's successful attacks against a named customer is a high-value artefact. Over-broad access or compromise becomes a single event with engagement-wide blast radius, where before it was one laptop. - **Access and retention.** A local file is governed by whoever holds the machine. A shared store needs a real answer to who may read transcripts — the whole team, the leads, the client's engineers? — usually differing by role, plus removal when someone rotates off, plus scheduled deletion, because accumulation across engagements is exactly what makes the store attractive and what a client contract is least likely to permit. ### Where the number misleads The seductive metric after consolidation is a big total: "the shared store shows 4,000 exchanges across the team." That is a row count, and row counts are not coverage. Coverage's denominator is objectives or surfaces in scope, not stored pieces, and without labels distinguishing operator, target and deployment, a pooled store cannot even produce attempts-per-objective — it produces one large undifferentiated number that reads as thoroughness. The second misleading claim is qualitative: presenting the move as a security improvement. It is not. It relocates exposure from many small uncontrolled copies to one large controlled one, and it is an improvement only once the controls exist. ### A workable middle, and what to insist on Prefer a **per-engagement** shared store over one perpetual team store: created at kickoff, access scoped to that engagement's team, destroyed or archived at close under the client agreement. That keeps the join benefits inside the engagement while capping accumulation and making the retention decision a natural event rather than a chore someone remembers. Before switching, insist on: a labelling scheme applied from the very first run, since labels cannot be added retroactively; named access with removal on rotation; encryption in transit and at rest; a scheduled destruction date with an owner; and a rehearsal that the export you will need for the report actually runs — discovering at the end of a large engagement that it does not is expensive. And measure the per-turn latency delta on a throwaway multi-turn run before committing a campaign to it, because that number is cheap to obtain in advance and painful to discover halfway through.

  • What is the single strongest argument for sharing on a multi-operator engagement?
    The report needs one honest denominator — what was attempted across the whole team — and reconstructing that from separate local files at the end is error-prone and often quietly abandoned.
  • What would make you keep operators local anyway?
    Separately delivered findings, a client contract that forbids pooling their transcripts with anyone else's, or no owner willing to be accountable for the shared store's access and retention.
  • How do you cap accumulation without losing the benefits?
    Scope the store to one engagement, create it at kickoff and destroy or archive it at close under the client agreement, rather than running one perpetual store for the team.

Consolidating everyone's transcripts into one shared store is like moving cash out of several unlocked drawers into a single safe. It is an improvement only once the safe has a lock, a list of who holds keys, and someone who notices when it is opened.

saying these in an interview costs you the question

  • Calling a shared backend more secure without naming the access controls it requires.
  • Adopting a perpetual team-wide store with no retention schedule.
  • No labelling scheme, so runs, operators and deployments blur together.
  • Ignoring availability and concurrent-writer effects on runs in flight.
  • Choosing local stores for an engagement whose deliverable requires an aggregated coverage claim.

context