skip to content

A filed jailbreak against your hosted chat model reproduces once in five tries - what does that prove?

level: middleimportance: should knowfreq 52%

answer

  1. a propensity, not a branch
  2. identical turns are separate draws
  3. a failed retry declines one attempt
  4. a rate points at the model half
  5. the label is stable, the weights are not

basics

~20 s

It proves the framing landed once, against one deployment, at one moment. Refusal is a trained propensity sampled at generation time, so a failed retry means that attempt was declined - not that the finding is wrong or the behaviour gone.

solid answer

~50 s

Nothing about a hosted model makes the same input produce the same output. Refusal is a learned propensity rather than a rule with a branch, and decoding samples, so five identical turns are five draws. The endpoint also sits behind a version label the vendor rotates, so the thing you are sampling can move between Monday and Friday, and each attempt starts at a fresh session boundary, so nothing accumulates between tries to explain the difference. Read the result in the right direction: the one success is evidence the framing can land, and the four failures are four declined attempts, not a refutation. The shape is diagnostic too. A defect in the half you ship - what the app does with returned text, a surface you exposed - reproduces on every run of the same build. A rate points at the model half, which is the half with no release of yours.

go deeper

for a junior

Remember that the same message sent twice to a model can get two different answers, so one refusal does not mean the earlier result did not happen.

for a middle

Explain the mechanics: a trained propensity rather than a rule, sampled decoding, a version label that can cover moving behaviour, and a fresh session each attempt that rules out accumulated state.

for a senior

Show you read intermittency as a locating signal - deterministic on one build points at code you ship, a bare rate points at the supplier's model - and that you record attempts and outcomes rather than closing as could-not-reproduce.

for a principal

Be ready to defend how your organisation records findings whose only stable property is a rate, so a ticket never asserts absence that the evidence does not support.

**An intermittent reproduction is not a weak finding; it is a finding whose defect sits somewhere your build system cannot reach.** Reading it correctly is the difference between filing something actionable and arguing about whether it happened. ### Why the same turn gives different answers Four things make repetition unreliable in a tool-less chat product built on a hosted model. First, the refusal is a *propensity*, not a rule. Nothing in the serving path evaluates a condition and returns a fixed decline; the model has been trained so that declining is the likely continuation for some inputs, and likely is not always. Second, generation samples. Two identical turns are two draws from a distribution, so the same framing can be met with a refusal and then with an answer without anything at all having changed. Third, the served model can move under an unchanged version string. Teams routinely assume a stable label means stable behaviour; it means a stable name for whatever the supplier currently serves under it. Fourth, in a product with no memory, each attempt begins at a fresh session boundary. Nothing carries over from the four failures, so there is no accumulated state to blame - which is itself useful information, because it rules out one explanation an engineer would otherwise chase. ### The direction of every claim This is where triage goes wrong. State each observation as exactly what it is: - The one success proves the framing reached the model and the model produced content it usually declines. It does not prove a reliable capability. - The four failures prove those four attempts were declined. They do not prove the framing fails, and they certainly do not prove the behaviour has been remediated. - A null result a week later is a measurement of whatever the endpoint serves today. It is not a refutation of the transcript that was filed. Note also which component declined. In this product there is no separate screening layer in the path, so the decline came from the model itself. In a product that has one, a refusal and a block are different events with different shapes, and telling them apart is the first measurement anyone should make - here the ambiguity simply does not exist, which makes the reading cleaner. ### The shape is diagnostic Run the same construction against the same build repeatedly and look at the shape of the results: | Observed shape | What it points at | | --- | --- | | Fires on every run of a given build | The half your team ships - a surface, the handling of returned text, the prompt as it is assembled | | Fires at some rate, no pattern in the app state | The model half - trained propensity plus sampling | | Fired reliably, then stopped with no deploy of yours | The served endpoint moved under its label | That table is the practical value of the intermittency. Deterministic reproduction on one build is a strong signal the defect is in an artefact with a repository and a release. A rate is a strong signal it is not. ### What the engineer who cannot reproduce it should write The worst outcome is a ticket closed as 'could not reproduce', because that records a false conclusion: it says the defect is absent when the evidence only says today's draws came back declined. The honest record keeps the original transcript as the evidence, states how many attempts were made and what came back, and states which half the behaviour is attributed to. A second-worst outcome is inflating one success into 'the model will do this on request', which invites a reviewer to disprove your claim with a handful of retries and dismiss the whole finding. ### Why this belongs to the ownership question A defect you can reproduce every time on your own build is one you can also fix on your own build, with a date. A defect whose only stable property is a rate against somebody else's endpoint has no such release. Intermittency is therefore not a nuisance to be engineered away before filing - it is the first evidence about which half of the estate you are in, and it should be reported as such rather than hidden behind a claim of reliability the construction does not have.

  • A week later the framing does not land at all. Does that close the finding?
    No. It measures whatever the endpoint serves today, under a label that can cover different served behaviour than it did last week. The transcript remains the evidence; what you record is that reproduction was not established on this date, across this many attempts, and that the behaviour is attributed to a supplier's endpoint. 'Could not reproduce' as a closure states a conclusion the attempts do not support.
  • Would you expect a defect in your own application code to behave this way?
    No, and that contrast is the useful part. What the app does with returned text, or a surface it exposes, fires on every run of the same build. When results come back as a rate with no correlation to app state or session history, the variance is coming from the sampled model, which tells you which half of the estate owns the behaviour before you argue about severity.

saying these in an interview costs you the question

  • Says a failed retry disproves the finding
  • Treats the model's refusal as a deterministic rule
  • Assumes a fixed version label means fixed behaviour
  • Reports one success as a reliable capability
  • Blames session history in a product with no memory

context