skip to content

In a character-chat product, a persona scene returns fluent in-character content on a refused topic. What can the tester claim?

level: seniorimportance: should knowfreq 42%

answer

  1. separate what you saw from what you inferred
  2. the scene needs a next line
  3. confident where it holds nothing
  4. cheap to generate, expensive to trust
  5. willingness to continue, not disclosure

basics

~20 s

Only that the model kept generating past the point it usually declines. In-character text is written to fit the story, so where the model has nothing it invents something plausible, and elicitation and confabulation look identical from inside the scene.

solid answer

~50 s

Three claims are in play and the scene supports only the first two. It shows the model produced material in a category it normally declines, and that the persona framing was what moved it. It does not show the material is accurate, and the fiction frame is the worst available instrument for judging that, because staying in voice rewards a confident next line whether or not the model holds anything real. A character who hesitates breaks the scene. So the tester who wants to claim uplift has to verify the content against ground truth outside the product, in a domain where they may not be competent to judge, and that verification usually costs more than producing the output did. That asymmetry is the practical reason this family is a weak attack: what it yields is cheap to generate and expensive to trust. Write the finding with the two claims separated.

go deeper

for a junior

Know that content arriving in a story voice may simply have been invented to fit the story, and that reading it does not tell you which it was.

for a middle

Be able to explain why the fiction frame produces confident invention: staying in voice means the character needs a next line, and hesitation breaks the scene the user set up.

for a senior

Demonstrate reporting discipline. Separate what you observed from what would need verifying, say what you compared against, and resist inferring which component failed from the output alone.

for a principal

Own the cost argument. Verification of this family's output is per-item, expensive and often outside your team's competence, so a programme that leans on it is buying findings it cannot stand behind.

## Separate the claims before you write the finding A scene that returns refused-category material supports a stack of claims that people routinely collapse into one: 1. The model generated text in a category it normally declines. 2. The persona framing is what moved it there. 3. The text is accurate. 4. The text gives a reader capability they did not already have. Claims 1 and 2 are what you observed, and 2 needs at least a comparison against the same request asked plainly. Claims 3 and 4 are the ones that decide severity, and the scene tells you nothing about either. ## Why the frame is the worst instrument for accuracy In-character generation optimises for consistency with the scene. The character has a voice, a stake and a next line, and the strongest constraint on that line is that it fits. Where the model holds real material, some of it may surface. Where it holds none, the scene still needs continuing, and a hedge or an admission of ignorance breaks the fiction the user built. So the model supplies something plausible in the same confident register it uses for the parts it does know. That is not an occasional failure mode of the family; it is the family's normal output. And crucially the two cases are indistinguishable at the point of reading: both arrive fluent, both arrive in voice, and neither carries a marker. You cannot separate them from inside the product, no matter how many turns you spend. ## What this does to a report A finding that says *the model produced instructions for X* is a different assertion from *the model produced correct instructions for X*, and only the second supports an uplift argument. The honest shape is to state what was observed, state that accuracy was not assessed, and if you do assess it, say what you compared against. Verification means checking the content against ground truth in the relevant domain, outside the model, by someone competent to judge it. In most of the categories where this matters, that is expensive, sometimes legally awkward, and often outside the tester's own expertise. That cost asymmetry is the whole practical story of this family. Producing fluent refused-category prose in a scene is cheap. Establishing that any of it is true is expensive. An attacker faces exactly the same asymmetry, which is why the family's payoff is weaker than its reputation: they have paid for something they cannot trust without leaving the product. ## Two adjacent errors The first is treating one successful scene as a rate. Generation is probabilistic and the surrounding turns differ, so a single success establishes that the construction worked once against this deployment. Whether it is a finding or a curiosity depends on how it behaves across runs, and reporting a single run as if it were reliable is how a triage queue fills with things nobody can reproduce. The second is inferring the failing component. Refused-category output arriving in full does not tell you whether there was no output screening, whether a screen scored it below its threshold, or whether the response was streamed and the remainder was cut after tokens had already been sent. Those are different observations with different shapes, and the report should say which one you actually saw. ## Severity, honestly A reasonable severity argument for this family usually rests on claim 1 alone: a product that will produce this category of material in voice has a control gap, and that is worth filing on its own terms, especially in a consumer product where the material reaches a user directly. What it should not rest on is an unverified assertion that the content is real and dangerous. Interviewers ask this question because inflating that step is the single most common flaw in AI red-team reporting, and because a candidate who volunteers the confabulation problem before being pushed on it has clearly done the work rather than read about it. ## The one-line answer The scene established a refusal-behaviour failure, not a disclosure. Fluency in voice is evidence about the model's willingness to continue, and about nothing else.

  • How would you word the severity in the report?
    State the observed behaviour as the finding: the product generated material in a declined category under a user-authored persona frame. Then state explicitly that accuracy was not assessed, or how it was. Keeping those sentences apart is what stops a reviewer inheriting an uplift claim you never made.
  • You verify one output and it checks out. Does that generalise?
    No. It tells you that one generation contained real material. The next run of the same scene may confabulate differently, because nothing in the frame distinguishes the two cases. Verification is per-item, which is exactly why the cost scales badly and why the family is a poor source of evidence.
  • Does this problem apply to other jailbreak families too?
    It is worst here because fiction actively rewards invention, but any framing that puts the model in a generative posture rather than a factual one shares it. The general lesson is that a construction's success rate and the truth of what it returns are separate measurements, and only the first is easy.

saying these in an interview costs you the question

  • Reports elicited content without assessing accuracy
  • Treats one successful scene as a reliability claim
  • Assumes fluent means real
  • Infers the failing component from the output alone
  • Claims uplift with no ground-truth comparison

context