skip to content

In refusal suppression, why does fixing an answer's first words bite harder than banning caveats later?

level: middleimportance: should knowfreq 48%

answer

  1. each token is conditioned on the last
  2. the decision surface sits at the start
  3. two constraints, two different jobs
  4. caveats shape the payoff, not the entry

basics

~20 s

Refusal behaviour concentrates where an answer begins, so a constraint on the first words competes with it directly. A ban on caveats applies to text already being written, so it shapes the payoff rather than whether the model declines.

solid answer

~50 s

A model generates autoregressively: every token is conditioned on all the text before it, including what the model has already written. Refusal, as trained, lives mostly in the opening tokens of a response - declining has a small repertoire of characteristic starts. A constraint on the opening is therefore contesting the same tokens as the refusal propensity, at the point where it is most concentrated; and if the model follows it, everything afterwards is conditioned on a response that has already begun answering. A ban on hedges and disclaimers is not fighting the refusal at all - a refusal is not a hedge. It strips the confidence markers from an answer that is already being produced. Reported honestly: the opening constraint is what got the answer written, and the qualifier ban is what made it worth having.

go deeper

for a junior

Know that the model writes one token at a time and that each one depends on all the text before it. That is enough to see why the beginning of an answer is special.

for a middle

Be able to explain autoregressive conditioning and say why refusal behaviour is concentrated in the opening tokens. Then map each constraint to the job it does - one aimed at entry, one at the finished text.

for a senior

Demonstrate that you would separate the two constraints before claiming either worked, and that you expect mid-answer reversal and empty form-compliance as ordinary outcomes rather than anomalies.

for a principal

The angle to own is what a family like this is worth over time: no quotable string, a durable property, and an effect size that only exists as a rate on one deployment. Decide what evidence you would accept before funding work on it.

## Two constraints, two jobs A mandated-form jailbreak usually bundles two different instructions, and candidates who treat them as one thing cannot say which part earned the result. | The constraint | What it is aimed at | What it does not do | | --- | --- | --- | | A mandated opening or fixed structure | the refusal, which is expressed as a way of starting | it says nothing about whether the answer is honest | | A ban on hedges, disclaimers and qualifiers | the confidence markers in the finished text | it does little to change whether the model declines | The first is about entry. The second is about payoff. They are frequently written in the same sentence, which is exactly why the effect gets misattributed. ## Why the start is where the decision lives A language model generates autoregressively: it produces a probability distribution over the next token, one is sampled, and the whole sequence so far - system prompt, the application's instructions, the user's turn, and everything the model has already written - becomes the condition for the next step. There is no separate moment at which the model resolves answer-or-refuse and then writes. Refusal, as trained, sits mostly in the first few tokens of the response. The training data that produced it is full of responses that decline, and those responses share a small repertoire of openings. So the probability mass that constitutes refusing is concentrated where an answer begins. Two consequences: 1. A constraint that specifies the opening is competing with the refusal behaviour **at the point where that behaviour is strongest and most concentrated**. It is the only place a formatting instruction and a refusal propensity are contesting the same tokens. 2. Once several tokens of substantive answer exist, the condition for every later token includes a model that has already started answering. The propensity has less to hold on to, not because anything was disabled, but because the context it is being evaluated against has changed. ## What a caveat ban actually removes Qualifiers - it depends, consult a specialist, this may not apply to your situation - are the model's epistemic markers. Banning them does not fight the refusal; a refusal is not a hedge, it is a different act. What it does is strip the signals a downstream reader uses to calibrate. In a drafting workflow where a human skims the draft and an automation files it, those markers are the entire mechanism by which anybody decides to check something before acting on it. Reported honestly, then: the mandated opening is what got the answer written; the qualifier ban is what made the answer worth having. ## Where it fails - **Mid-answer reversal.** The model can comply with the opening and then decline part-way through, once enough of the answer exists for the trained behaviour to engage on what is actually being produced. This is common on requests that sit well inside a strongly trained refusal. - **Form without substance.** The model satisfies the structure and produces something vacuous. The run looks like a success in a log and is worth nothing. - **The constraint is ignored.** It is only an instruction, competing with the application's own formatting instructions, and instruction-following is itself imperfect. ## Attribution, briefly Because both constraints usually ship together and the underlying request may be one the assistant would often answer anyway, a single run tells you nothing about which piece did the work. The only honest way to say the opening constraint earned the result is to run the arms separately, many times each, and compare - and even then you are describing one deployment on one day.

  • Can the model still refuse after it has written a compliant opening?
    Yes. Mid-answer reversal is common on requests that sit well inside a strongly trained refusal: the model complies with the opening, then declines once enough of the answer exists for the behaviour to engage with what is actually being produced. The other common outcome is form without substance - the structure is satisfied and the content is vacuous. Both look like partial successes in a log and are worth almost nothing.
  • If the caveat ban does not fight the refusal, why include it?
    Because it is the payoff. In a drafting workflow the qualifiers are the signals a reader uses to decide whether to check something before acting. Remove them and an unremarkable answer reads as a settled recommendation, which is the property that makes the output useful to an attacker even when nothing disallowed was produced.
  • How would you tell which of the two constraints produced a given result?
    Only by running the arms separately, many samples each, and comparing rates - the opening constraint alone, the qualifier ban alone, both together, and neither. One run tells you nothing: generation is sampled, and the underlying request may be one the assistant answers a fair share of the time anyway. Any number you get describes one deployment on one day.

saying these in an interview costs you the question

  • Thinks the model re-evaluates the request equally at every token
  • Says a caveat ban is what makes the model less likely to refuse
  • Treats a compliant opening as making refusal impossible
  • Confuses this with a decode-time constraint imposed by the caller
  • Cannot separate compliance with the form from a useful answer

context