skip to content

An application team disputes one item in your report from a PyRIT engagement. How do you get from that report item back to the exact stored exchange that produced it, and what does the transcript prove and not prove?

level: seniorimportance: should knowfreq 42%

answer

  1. report item vs stored exchange
  2. carry conversation and run identifiers
  3. one exemplar plus the siblings
  4. date, endpoint, deployment
  5. once is not always; rare is not false

basics

~20 s

Each stored turn carries identifiers tying it to its conversation and run, so a report item can point at the exact exchange that produced it. Keep those identifiers in the finding. The transcript proves what happened once against that configuration - not that it reproduces, since the target samples and may have changed.

solid answer

~50 s

The chain only exists if you build it while writing the report. A report item is a triaged, deduplicated claim; the tool's unit is a single stored exchange. Carry the conversation and run identifiers from the store into the finding so the two can be rejoined, and quote the exchange rather than paraphrasing it. What the transcript settles: that this prompt produced this response at that time, under the target configuration you were pointed at, and how the run judged it. What it does not settle is reproducibility - sampling makes a repeat run a different draw, and the target may have been patched or had a guard put in front of it since. Record what you were actually testing: endpoint and deployment. The most common way a dispute goes wrong is both sides being right about different deployments. If the team cannot reproduce it, that is a data point, not a retraction: give the observed frequency across attempts.

go deeper

for a junior

Should know the stored exchange is the evidence and that a finding should quote it rather than describe it from memory.

for a middle

Explains carrying identifiers into the report and that nondeterminism means a single observation is not a guarantee of reproduction.

for a senior

Handles the dispute properly — same endpoint and deployment, frequency across attempts, exemplar plus siblings, and knows a rare behaviour is not a false positive.

for a principal

Sets the evidence standard for the team: store hygiene per engagement, what a finding must carry, and how transcripts are quoted into documents that circulate more widely.

### The chain exists only if you build it during the run Every stored piece carries identifiers that can rejoin it to a finding: the `conversation_id` grouping one exchange, a sequence number ordering it, an identifier for the attack that produced it, the target it was sent to, timestamps, and — the one operators skip — the **memory labels** you pass when you execute an attack, a set of key/value pairs that PyRIT stamps onto every piece the run writes and that you can filter by afterwards. Labels are how an operator name, an engagement name and a target deployment get into the record. They cost seconds at run time and are effectively impossible to add later: retro-labelling three weeks of unlabelled exploratory runs against four deployments means guessing from timestamps, and the guess is not evidence. ### Report items and stored exchanges are different units A report item is a triaged, deduplicated claim. The tool's unit is one stored exchange. Ten conversations showing the same weakness collapse into one finding, and the mapping has to survive that collapse: keep the exemplar you quote **plus** the identifiers of its siblings, because "we saw it once" and "we saw it in eleven of forty attempts" are entirely different claims to an application team, and only the store can tell you which one you are entitled to make. ### What a finding should carry The exchange as sent and as received — trimmed, not paraphrased, and diffed against the stored converted value before it goes in. The run date. The endpoint and deployment identifier you actually targeted, and whether you went through the application front end or straight at the model endpoint behind it. How the run judged it, and by which definition of a hit. The observed attempts and hits, if you tried more than once. Anything you cannot support from the store does not belong in the item. ### What the transcript proves, and what it does not It settles that this prompt produced this response at that time, against that configuration, and how the run scored it. It does **not** settle reproducibility. Chat targets sample, so a repeat is a fresh draw; the target may have been patched, or had a guard placed in front of it, since the run. ### Where the number misleads - **Wrong denominator.** The denominator is attempts at *that objective* under *that configuration*, not the total number of conversations in the store. Mixing engagements or deployments into one store makes the honest denominator unrecoverable. - **Turn inflation.** If the scorer ran on every turn, a conversation that succeeded at turn seven may carry several hit records. That is one hit, not several. Count conversations, or objectives, and say which. - **Failed reproduction read as refutation.** This is the big one. If the true rate is five percent, five clean attempts have roughly a 77% chance of showing nothing at all (0.95 to the fifth). A behaviour that reproduces twice in forty is not a false positive; it is a probability, and it should be reported as one. The mirror error is yours to own: if your own store shows a single hit across many attempts and the item was written as though the system always complies, the correction is the report's, not the team's. - **Deployment mismatch.** The most common way a dispute goes wrong is both sides being right about different systems — a different deployment, a different system prompt, or a front end that adds a guard your target adapter bypassed. Only the labels and the target identifier settle that. ### What to check Before writing the item, pull the exchange back by identifier and compare it, character by character, against what you intend to quote. Confirm the labels identify the deployment unambiguously; if they do not, say so rather than implying precision you lack. Compute attempts and hits with an explicit filter and record the filter in the item. Keep one store per engagement so run boundaries, targets and dates stay unambiguous and the artefact can be retained or destroyed as a unit. When you hand the dispute back, ask for reproduction against the same endpoint, several attempts, and the same path through any guard — and quote into the circulating document only what the finding needs, because the full store is held under stricter rules than the report is.

  • The team reproduces nothing in five tries. What is your response?
    Establish they hit the same endpoint and deployment without an extra front-end guard, then compare frequencies: if the run saw it twice in forty, five clean tries is consistent with the finding rather than a refutation.
  • How should the finding state frequency?
    As observed attempts and hits under a stated definition of a hit, from the store — not as a bare percentage detached from the number of attempts.
  • Why keep a separate store per engagement?
    So run boundaries, targets and dates are unambiguous later, and so the artefact can be retained or destroyed as one unit under the engagement's handling rules.

A finding that reproduces twice in forty attempts is an intermittent fault, like a noise a car makes on some cold mornings. A mechanic who drives it round the block five times and hears nothing has not shown the noise is imaginary.

saying these in an interview costs you the question

  • A finding with no pointer back to a stored exchange.
  • Paraphrasing a response instead of quoting the stored one.
  • Retracting a finding because the team could not reproduce it on the first attempt.
  • Mixing several engagements and deployments into one undifferentiated store.
  • Presenting a single observation as though the system always behaves that way.

context