A colleague files a jailbreak finding whose entire content is one quoted prompt — how do you judge its value?
answer
- an instance, or a class
- re-run it before valuing it
- which variants still worked
- shelf life versus re-derivable capability
basics
~10 sTreat it as an instance until somebody names the class. Re-run it, ask which variants still worked, and value it at its remaining shelf life unless a property behind it can be stated.
solid answer
~40 sA quoted prompt is one sample of something, and the whole question is whether anyone can say what it is a sample of. I would re-run it first: a probabilistic result needs to be seen working now, and one success in several tries is a rate, not a boolean. Then I would ask the filer the only question that changes its value — what do you think made it comply, and what did you change that still worked? Surviving variants are the evidence that a class exists rather than a lucky wording. If a property can be stated, the finding outlives every string in it. If not, it is a dated artefact whose value ends when a screen or a tune absorbs it, and I would file it as such.
go deeper
Recall that a quoted prompt is a single observation from one day against one deployment, and that the first honest step is to run it again rather than take the filing at face value.
Explain why surviving variants are the evidence that matters: they distinguish a property with many wordings from a single fitted string whose life is set by someone else's release cadence.
Demonstrate the triage judgment end to end — reproduce, get the rate, chase the variants, name the class or say plainly that there is not one, and value the filing accordingly.
Own what this does to a backlog. Decide what the team is allowed to count as a finding, so its inventory is classes it can re-derive rather than strings that expire without anyone noticing.
## What you have actually been handed A quoted prompt is an observation: on some day, against some deployment, this wording produced compliance the model would normally decline. That is a real fact and it is worth something. What it is not is a description of a capability, because nothing in it says why the model complied or what else would have worked. The whole triage job is to find out whether an instance is hiding a class. ## Step one: does it still reproduce, and at what rate Run it now. Two things commonly turn up. The first is that it does not work any more, in which case you are looking at an artefact whose window already closed — worth recording, not worth planning around. The second, more interesting, is that it works sometimes. Jailbreak results are probabilistic: sampling, conversation history and deployment differences all move the outcome. A construction that lands one time in five has not been disproven, and it has not been established as reliable either. Say the rate out loud rather than converting it into a yes. Keep the direction of the claim honest in both directions. A successful reproduction proves the construction worked once against one deployment. A run of failures proves those runs failed, not that the behaviour is gone. ## Step two: ask the only question that changes the value Ask the filer two things: what do you think made it comply, and what did you change that still worked? The second is the one that carries information. If they varied the wording, the ordering, the persona framing, the length, the language, and some of those variants still produced compliance, then the surviving dimension names the property — and the property is the asset. If nothing survived a change, you have a lucky string, and its lifetime is whatever the vendor's attention and cadence leave it. This is also where a lot of filed findings turn out to be the same finding. Several quoted prompts that share a property are one class-level result with several samples, and recognising that is worth more to the team than the count. ## Step three: value it accordingly | What the filing contains | What it is worth | |---|---| | A wording, nothing else, no longer reproducing | a dated record; close it | | A wording that reproduces, no surviving variants | its remaining shelf life, which is set by attention and cadence, not by cleverness | | A stated property plus variants that survive change | a capability the team can rely on and re-derive after the string dies | The reason this matters beyond one ticket is that the two kinds age completely differently. A wording is cheap to remove: a deployed screening layer can match it in a release, and once it is public it is also the exact material that ends up in the next round of safety data. A property is expensive to remove, because it takes a behavioural change and produces collateral refusals on legitimate lookalike requests. A backlog full of the first kind gives a team an inventory number that quietly evaporates; a backlog of the second kind is what still answers a question next quarter. ## What not to do with it Do not treat the quotation as the deliverable and circulate it more widely to prove the point — every copy shortens the window you are still measuring. Do not conclude from a single reproduction that the class is confirmed, and do not conclude from its disappearance that anyone fixed anything on purpose. And do not let the presence of a vivid artefact substitute for the sentence that actually matters: this works because the model treats requests framed in this way as lower risk, and that framing gets past the thing that would otherwise have stopped it. ## The interview signal An interviewer asking this is watching for whether a candidate values a finding by how impressive the demo is or by how long it will stay true. The strong answer is unglamorous: reproduce, ask for the variants, name the class if one exists, and be willing to say that a striking prompt with nothing behind it is worth very little and is getting cheaper every day it exists in writing.
- It reproduces one time in five. Do you record that as working?I record the rate, not a verdict. A probabilistic construction that lands sometimes is genuinely different from one that lands reliably, and collapsing it to a yes hides the thing a reader needs. It also protects you later: when it stops reproducing, you have not claimed something you cannot defend, and you know whether the change is meaningful or inside the noise you already measured.
- The filer insists the exact wording matters and refuses to generalise. What now?Then treat the wording itself as the hypothesis and test it: change one dimension at a time and see what survives. If nothing does, that is a real answer — it is a fitted string, and it should be valued at its shelf life. Insisting the magic is in the words, with no variant evidence, usually means nobody has looked for the property yet.
- Two teammates file different prompts. How do you decide whether it is one finding or two?By property, not by text. If both work because the model treats the same shape of request as lower risk, they are two samples of one class, and the class is what gets recorded. Counting them as two inflates an inventory of strings while the underlying result stays unstated — and both samples die together the moment that shape is retuned.
saying these in an interview costs you the question
- Accepts a quoted prompt as a confirmed capability
- Converts a one-in-five result into a plain yes
- Never asks which variants still worked
- Counts strings rather than classes as findings
- Circulates the quotation widely while still measuring its lifetime