A red-team finding you filed three weeks ago against a hosted chat endpoint no longer reproduces. What captured from the original run would let you distinguish a silent provider-side model change from a shipped fix, and what do you do if you captured none of it?
answer
- non-reproduction is one sample, not proof
- diff echoed model id then vs now
- diff system prompt and tool config
- request ids and timestamps bracket the window
- status: cause unattributed, not fixed
basics
~20 sCompare the model identifier the response echoed then and now, the application's system prompt then and now, and your request ids and timestamps. A changed identifier or prompt points at drift; both unchanged points at a fix or variance. With none captured, you cannot attribute it — re-run, capture properly this time, and say so.
solid answer
~50 sThree candidate causes look identical from a failed re-run: the provider swapped the served build, the app team changed the system prompt or added a filter, or you are inside normal variance. Distinguishing them is a *diffing* problem, and you can only diff what you stored. The discriminators are: the model identifier echoed in the response envelope, captured then and re-captured now; the application's system prompt and any tool or retrieval configuration, then and now; the request ids and timestamps bracketing your original attempts; and your original attempts-and-successes count, which tells you whether one failed re-run is even surprising. If you captured none of them, be direct: you cannot attribute the change. Re-run enough attempts to bound variance, capture the full record this time, and ask the app team what shipped in the window. Write the finding's status as "no longer reproduces, cause unattributed" rather than "fixed" — closing it as fixed on that evidence is how a defect quietly comes back when the provider rolls forward again.
go deeper
Re-runs it a few more times and reports that it no longer works, without separating the possible causes.
Names the three candidate causes and knows to compare the model identifier and system prompt before concluding anything.
Treats it as a diff against a record deliberately captured during the live run, states the finding's status as unattributed when the record is missing, and resists closing on one failed re-run.
Sets the org rule that findings against a silently reversioning target close on evidence of a change, never on a single non-reproduction, and mandates the capture that makes that possible.
### A non-reproduction is a measurement, not a verdict Three causes look identical from the outside when a filed attack stops working: the provider swapped the build serving the endpoint, the application team changed its system prompt or added a filter, or you are inside ordinary variance. Separating them is a **diffing** problem, and you can only diff what you stored while the run was live. ### What each discriminator actually buys | Captured then, re-captured now | What a difference tells you | What it does not tell you | |---|---|---| | Model identifier echoed in the response | The provider moved under you; the defect's status against the old build is now unknowable | Nothing about *why* it changed, and an unchanged id does not prove an unchanged build | | Provider build or backend fingerprint, where offered | Finer-grained evidence that the serving stack moved | Which part moved | | The application's system prompt (and its digest) | App teams patch prompts constantly and rarely file the patch as a fix; a prompt diff often explains the whole thing | Whether the patch was deliberate mitigation or unrelated product work | | Tool and retrieval configuration | A middle turn may have depended on a document that no longer comes back | Whether the document merely moved rank | | Request ids and timestamps | The window to hand the provider or platform team so someone can read server-side logs | Anything by themselves | | Original attempts and successes | Whether today's zero hits is even surprising | The cause | That last row is the one people skip, and it is the one that decides the argument. ### The arithmetic that makes a failed re-run meaningless Suppose the original run recorded four successes in twenty attempts — a rate of about one in five. You re-run three times today and get nothing. The probability of three consecutive misses at an unchanged one-in-five rate is 0.8 cubed, about **51 per cent**. Half the time, an entirely unfixed defect produces exactly the result you are looking at. To push the chance of a false "gone" below five per cent at that rate you need roughly fourteen consecutive misses. That is the cost line. Fourteen attempts of a single-turn prompt is cents and a few minutes. Fourteen attempts of an eight-turn conversation is over a hundred target calls, plus whatever the app charges you in session setup, plus the engineer time to rebuild any state the transcript does not carry — call it an hour, not a coffee break. Budget it explicitly, because the alternative is not cheaper: it is a wrong closure. ### Where the numbers mislead - **Zero out of three read as "fixed".** The commonest error in the whole leaf, and the arithmetic above is the answer to it. - **A changed model identifier read as "the model was fixed".** It shows something moved. The mitigation may equally have been a moderation layer bolted in front, in which case the underlying model behaviour is untouched and will resurface behind any interface that bypasses that layer. - **An unchanged identifier read as "the build is unchanged".** The identifier is a label the provider chooses, not a hash of the weights. Providers ship changes behind stable names routinely. - **A hand-typed re-run.** The person re-testing usually retypes the prompt rather than replaying the stored bytes, and a stray newline or a smart quote is enough to lose a marginal attack. That is a *self-inflicted* non-reproduction, indistinguishable in the ticket from a real one. ### What you would check, in order Replay from the stored record programmatically — the exact bytes, the same decoding settings, the same system prompt — rather than by hand. Diff the system-prompt digest then against now. Diff the echoed model identifier and any build fingerprint. Pull the application's change log for the window between your original timestamps and today, and ask whether a filter or prompt shipped. Hand the request ids to whoever can read server-side logs. Then run enough attempts to state a rate with a denominator, not a single miss. ### What you write when you captured none of it Say so plainly, and downgrade the claim to what the evidence supports: *"no longer reproduces against build X as of this date; original build unrecorded; cause unattributed"*, with the re-run attempt count attached. That is weaker than "fixed", and it is true. A triage process that closes items on one failed re-run against a silently reversioning target does not produce a record of what the system does — it produces a record of when you happened to test.
- The echoed model identifier is unchanged and the system prompt is unchanged. What is left?Variance, or a change that is not visible in either — a moderation layer, a routing change, or a build swapped behind a stable identifier. Bound variance with more attempts before claiming anything.
- How do you keep the app team's prompt changes visible to you without asking every time?Store a digest of the system prompt with every run record. Comparing digests across runs shows you a change happened even when nobody told you.
Failing three times to reproduce a one-in-five attack is like knocking three times on a door and concluding nobody lives there; chance alone gives you silence about half the time.
saying these in an interview costs you the question
- Closes the finding as fixed on a single failed re-run.
- Assumes non-reproduction means the vendor patched it.
- Has no timestamps or request ids and still asserts a cause.
- Never checks whether the application's system prompt changed in the window.
- Snapshots the model identifier once per campaign rather than per attempt.