A form constraint made a CRM drafting assistant emit a confident, caveat-free recommendation with nothing disallowed: is that a finding?
answer
- nothing disallowed was produced
- what was removed, not what was added
- who reads the draft before it leaves
- compare against the unconstrained baseline
basics
~20 sYes, but not as a content-policy bypass. Nothing disallowed was produced; what was removed are the qualifiers a reader and the downstream workflow use to decide whether to act, which makes it a misinformation and overreliance finding.
solid answer
~50 sGrade it on what changed and who consumes it. Nothing restricted was extracted, so filing it as a content bypass gets it closed correctly and for the wrong reason. What the construction removed is the epistemic signal: an unhedged draft reads as settled to a human skimming it and to any automated step that files or sends it. Be precise about the evidence: no refusal proves an answer was produced, not that it is right; the output screen passing proves the finished text scored low on the categories that screen scores, and absence of hedging is not one of them. Then establish worth - does anything act on a draft unread, does it reproduce across sessions, and does an unconstrained baseline produce an equally confident draft anyway? If the baseline does, the construction bought nothing and the issue is the assistant's default register.
go deeper
Recall that a model not refusing only means an answer was produced. It says nothing about whether the answer is correct or whether anything was bypassed.
Be able to explain why a content screen passing the draft is not evidence of harmlessness: it scores particular categories, and unhedged confidence is not one of them.
Show the triage instinct: establish what consumes the draft, whether the unconstrained baseline behaves the same, and whether it reproduces, before you claim the construction caused anything.
The call to own is how a formatting-only outcome gets classified and routed, and what you will honestly claim from a handful of runs. Report the property you demonstrated, and state plainly what one draft is not.
## What actually changed The construction did not extract anything the model is trained to withhold. It fixed the shape of the answer so that a decline was off-format, and it forbade qualifying language. What came back is a reply draft that recommends a course of action in confident, unhedged prose, and there is nothing in it a content screen would score. The reporter's instinct is to close it: no policy violation, no disallowed content, no data left the system. That instinct is wrong, but so is the opposite reflex of calling it a jailbreak of the same weight as one that produced restricted content. The finding is real and it is a different kind of finding. ## What each piece of evidence proves - and what it does not Interviewers press hard here, because getting the direction of these claims right is most of the skill in triage: - **The model did not refuse.** That proves an answer was produced. It does not prove the answer is correct, and it does not prove the model would have refused without the constraint. - **The output screen passed the draft.** That proves the finished text scored below a threshold on the categories that screen scores. Absence of hedging is not a category any content screen scores, so passing carries no information about this harm at all. - **It reproduced once.** That proves the construction worked once, against one deployment, on one sampled generation. With a probabilistic system that is an observation, not yet a finding. - **The draft is confident.** That proves nothing about whether the recommendation is right. A confident wrong draft and a confident right draft are indistinguishable at this point. ## Filing it where it will actually be read The published GenAI list keeps jailbreaking under its Prompt Injection item (LLM01), so filing it there is not wrong on the letter - but it routes the report to whoever handles content-policy bypass, who will close it correctly and for the wrong reason. The harm you are describing is the one the Misinformation item (LLM09) covers: output that people and automated steps overrely on because nothing in it signals doubt. Note also what it is not: Improper Output Handling (LLM05) is about downstream components treating model output as active content, and nothing here is being executed or rendered as markup. Filing is not a formality on this class of finding. It decides whether the report reaches the person who knows whether anything acts on a draft unread. ## What determines whether it is worth anything Three questions decide the grade, and a candidate who asks them is doing the job: 1. **What consumes the draft?** If a human reads and edits every draft before it leaves, the finding is thin and you should say so plainly. If the workflow files or sends drafts on a threshold, or a later automated step treats a recommendation as settled, it is serious. 2. **Does the unconstrained baseline do the same thing?** If the same request without any form constraint yields an equally confident draft a comparable share of the time, the construction bought nothing. What you have found is the assistant's default register, which is a real issue and somebody else's. 3. **Does it reproduce across sessions and over time?** A one-off on a sampled model is an anecdote. Sampling variance, a changed system prompt on the integrator's side, and a changed model behind the same widget will each move the rate independently. ## The honest ceiling Say what the finding is: the assistant's answer can be made to look settled on demand, from the one field the product exposes to anyone, and neither the refusal behaviour nor the screen that reads the finished draft has any purchase on that property. Say what it is not: it is not evidence that restricted content can be extracted, it is not a breach of anything, and one successful draft is not a rate. A report that claims the first and admits the second is the one that gets acted on.
- The unconstrained request produces an equally confident draft about as often. What now?Then the construction bought nothing and you say so. What you have is an observation about the assistant's default register - it answers recommendation-shaped questions confidently by default - which may still be worth reporting, but as a different issue with a different owner. Claiming your constraint caused it when the baseline matches it is the fastest way to lose credibility on the next report.
- What does the output screen passing the finished draft prove?That the text scored below a threshold on the categories that screen scores. Nothing more. Confidence markers, or their absence, are not a category any content screen scores, so passing carries no information at all about this particular harm. Reporting a pass as evidence of harmlessness inverts the direction of the claim.
- How much does one reproduction buy you?It establishes that the construction worked once, against one deployment, on one sampled generation. That is an observation, not a rate. Sampling variance, the integrator's own prompt changing, and a different model behind the same widget will each move the result independently, so a finding that is going to be argued about needs many runs on both arms before you attach a number to it.
saying these in an interview costs you the question
- Closes it because no disallowed content was produced
- Claims the output screen passing means the draft is harmless
- Reports one lucky run as a reliable bypass
- Dismisses missing caveats as a style issue, not a security one
- Assumes the assistant would have hedged without the constraint, without checking
- Treats a confident draft as evidence the recommendation is correct