Your custom PyRIT prompt target calls a service that keeps conversation state server-side behind a session id it issues. What breaks within a multi-turn run and across repeated runs, and how do you make each attempt independent?
answer
- two histories: framework and server
- double context or cross-attempt bleed
- one session per conversation, then close
- parallel attempts sharing one session
- shuffle order to prove independence
basics
~20 sPyRIT already tracks conversation identity, so reusing one server session makes history arrive twice or bleed between attempts. Mint a fresh session per PyRIT conversation, bind it to that conversation, and tear it down at the end. Otherwise a result depends on which attempts ran before it and stops being reproducible.
solid answer
~50 sThere are two histories now: the one PyRIT keeps in its memory store, and the one the service keeps behind its session id. If the adapter does not reconcile them you get one of two bugs. **Double history.** The adapter passes prior turns in the request *and* reuses a session that already holds them. The model sees each turn twice, so the transcript in memory misstates what was actually in context. **Cross-attempt bleed.** One long-lived session serves every attempt. Attempt seventeen inherits everything the first sixteen said, so a hit may be residue of earlier priming rather than the prompt credited with it — and shuffling the order changes the result. The fix is a one-to-one mapping: new session per PyRIT conversation, reused for retries within it, closed at the end. Then decide deliberately whether the adapter sends history or the session carries it — one or the other, never both.
go deeper
Should recognise that reusing one session across attempts lets earlier prompts influence later ones.
Explains double history versus cross-attempt bleed and maps one server session to one framework conversation.
Handles concurrency, retry double-appends and session teardown, and validates independence by re-running in a different order.
Treats reproducibility as the reporting bar: state that outlives a run, including per-account memory, is either isolated per run or disclosed as a caveat.
### Two histories, one transcript A stateful endpoint issues a session id and accumulates the conversation behind it. PyRIT is also keeping a conversation: every request and response piece is written into its memory store keyed by `conversation_id`, and that store is what the report, the transcripts and any later re-score are computed from. The moment your adapter talks to a stateful service, there are two histories and only one of them is visible in the record. PyRIT's memory shows what you *sent*; it has no way to know what the server had already accumulated. That gap produces two distinct bugs. **Double history.** The adapter dutifully passes the prior turns in the request body *and* reuses a session that already holds them. On turn three the model sees turns one and two twice. The stored transcript shows a clean three-turn conversation; the model saw five. No error is raised — a chat endpoint accepts repeated turns silently — so the only symptom is that results are not reproducible against a stateless equivalent. **Cross-attempt bleed.** One long-lived session, typically held in an instance field on the target object, serves every attempt in the run. Attempt seventeen inherits everything the first sixteen said. A hit may be residue of priming that happened forty attempts earlier and is credited to whatever prompt happened to be in the sender's hand when the model finally gave way. ### Making attempts independent **Decide who owns history, explicitly.** Either the session is the source of truth and you send only the new turn, or you send the full history each call against a session that carries nothing. Both are correct. Mixing them is what produces the duplication, and it is the mixing, not either choice, that is the defect. **Map lifecycles one-to-one.** Mint a session when a PyRIT conversation begins, key it by that `conversation_id`, reuse it for every turn of that conversation, and destroy it at the end. Concurrency makes this sharp: PyRIT runs attempts in parallel, so a single session id in an instance field means several attempts interleave their turns into one server-side thread. The resulting transcript is not merely wrong, it is unrecoverable — you cannot separate afterwards which turn belonged to which attempt. **Handle the retry interaction.** On a stateful endpoint a retried turn may append twice: once from the call that timed out client-side but landed server-side, once from the retry. Either make the append idempotent with a client-supplied turn id the service honours, or abandon the session and restart the attempt. Retrying blindly into a stateful session is how a run acquires context nobody wrote down. **Look past the session.** Many services carry per-user profiles, preference memory or a retrieval index that outlives any session. Then session hygiene is not enough: consecutive runs share state through the account, and this week's results are conditioned on last week's prompts. The remedy is a fresh test principal per run; where that is impossible, the report must say so as a caveat rather than assume it away. ### What it costs Session lifecycle is the part of a custom target that turns half a day of work into two: create, bind, tear down, plus the concurrency-safe storage of the mapping. The validation costs more than the code. Proving independence means running the same attempt set a second time in a shuffled order, which doubles the run's metered spend — and a multi-turn red-team attempt already meters three calls per turn (attacker model, target, scorer), so a 60-prompt five-turn run repeated is on the order of 1,800 calls. That is the price of being able to say the numbers are reproducible; budget it once per adapter, not once per engagement. ### Where the number misleads An attack-success rate from a bled run is not noisy, it is order-dependent, and order-dependence is invisible in the summary. The rate can be inflated — later attempts inherit priming that no single stored prompt contains, so the tool credits weak prompts with hits they did not earn and the report recommends defending against the wrong thing. It can equally be deflated, when an early refusal in a shared session makes the model refuse everything after it. Either way the finding does not reproduce, and a finding that does not reproduce is a finding the owning team will correctly reject. ### What I check Run the same attempt set twice with the order shuffled and diff per-attempt outcomes. Stable outcomes support independence; outcomes that move mean state is leaking through a session, an account or a cache in front of the model. Separately, ask the service owner what the endpoint retains beyond the session, and confirm a session teardown actually clears it rather than merely closing the handle.
- How would you demonstrate that cross-attempt state is leaking?Run the same attempts twice in a shuffled order. If per-attempt outcomes move, something outside the attempt is carrying context — session, account memory or a cache.
- The service also keeps per-account long-term memory. What changes?Session hygiene is no longer sufficient. Use a fresh test principal per run, or accept and state in the report that results are conditioned on everything that account did before.
Interviewing forty witnesses one after another in the same room, with everyone still sitting there listening. Witness seventeen answers partly because of what the first sixteen said, and the transcript shows only your question.
saying these in an interview costs you the question
- One session held in an instance field and shared by every attempt in the run.
- Sending full history to an endpoint that is already accumulating it server-side.
- Retrying a turn into a stateful session without considering a double append.
- Assuming closing a session clears longer-lived per-account memory.
- Never re-running in a different order to check that attempts are independent.