skip to content

An attacker mandates a drafting assistant's opening line and bans hedging: why does that change whether it refuses?

level: juniorimportance: must knowfreq 62%

answer

  1. the request never changed
  2. refusing is produced, not decided once
  3. declining has a shape, and it starts early
  4. no usual opening, less to hold on to

basics

~20 s

A refusal is text the model generates token by token, and it usually begins a characteristic way. Constraining the answer's first words and forbidding qualifiers removes those openings, so the trained tendency to decline has much less to latch onto.

solid answer

~50 s

This family attacks the form of the answer rather than the content of the request. The request is unchanged; what the attacker adds is a constraint on shape - a mandated opening, a fixed structure, a ban on qualifying language. That matters because refusing is not a switch thrown before the model writes. The refusal is produced token by token like everything else, and the trained propensity to decline is expressed most strongly in how an answer starts. Fix the first words to something that is already an answer and every later token is conditioned on a transcript where the model has begun complying. It is a probability shift, not a guarantee: the model can decline inside the mandated format, reverse mid-answer, or satisfy the form and empty the substance. And since the attacker has only a text field, the constraint is itself a request the model may ignore.

go deeper

for a junior

Be ready to say that a refusal is text the model writes, not a flag it sets, and that it usually starts a recognisable way. That single fact is why a constraint on shape can touch it at all.

for a middle

Expect to explain why the effect concentrates at the start of generation, and why a ban on qualifiers does a different job from a mandated opening. Name the trained refusal behaviour as the thing being got past.

for a senior

Show that you know what one success proves: that it worked once, on one deployment, on one sampled generation. Separate compliance with the form from compliance with the substance when you grade a run.

for a principal

Own the framing that this produces an output problem rather than a content-policy bypass. The draft may contain nothing disallowed and still be worth reporting, because the markers a reader uses to decide whether to act are gone.

## What the attacker changed, and what they did not A reply-drafting widget embedded in someone else's customer-relationship system exposes exactly one thing to whoever is typing into it: a free-text instruction field whose contents the integrator's own prompt concatenates in before asking the model for a draft. The attacker has no access to the API call, no control over the system prompt, and no decoding parameters. What they write is not a different request. It is a set of constraints on the **shape** of the answer: it must begin a particular way, it must follow a fixed structure, and it must contain no qualifying language. The thing being asked for is unchanged. The thing being specified is the form. That is the whole family. It is sometimes called refusal suppression, and it is the one jailbreak family that never touches the content of the request at all. ## Why the form can touch the refusal The intuition that fails here - and it is the answer most competent engineers give first - is that the model makes a decision, answer or refuse, and then, having decided, writes the result in whatever format was asked for. On that picture the output format is irrelevant to safety and this family cannot exist. Generation does not work that way. The model emits one token at a time, each conditioned on everything before it, and a refusal is simply one of the continuations available at each step. As trained, refusal behaviour is bound up with **how an answer starts**: a small, well-worn set of openings that acknowledge the request and decline it. It is a propensity over next tokens, not a flag consulted once before writing begins. So constrain the opening to something that is already the beginning of a substantive answer and two things follow. First, the tokens a refusal usually begins with are off-format - they conflict with an instruction the model is otherwise trying to follow. Second, if the model does follow the constraint, every token after that is conditioned on a transcript in which it has already begun complying, which is a materially different position from the one where the trained propensity is strongest. The ban on qualifiers does a related but separate job, and it is worth keeping the two apart: the mandated opening is aimed at whether the model declines; the ban on hedges, disclaimers and uncertainty markers is aimed at what the answer looks like once it exists. ## Three things this is not - **Not prompt injection.** Injection targets the application's own instructions using content the application feeds the model - a retrieved chunk, a fetched page, an inbound message. Here the text arrives in the attacker's own turn and the target is the model's trained refusal behaviour. That makes it a jailbreak. - **Not a privileged prefill.** A caller who can set the assistant turn directly imposes those tokens; nothing about instruction-following is involved. The attacker in a text field can only ask. That privileged mechanism is an API feature with its own rules about untrusted input, and it belongs to the API's own topic, not here. - **Not a decode-time constraint.** No grammar or schema is masking anything. Every part of this runs through the model's willingness to follow a formatting instruction. ## What it costs, and where it stops working The construction is cheap to write and needs no privileged position - that is its appeal. The costs are on the other side: - The constraint is a **request**, competing with whatever the integrator's prompt already says about the draft's form. The model may simply not follow it. - **Compliance with the form is not compliance with the substance.** Three common outcomes: the model produces the mandated opening and then something empty; it declines inside the required structure; or it starts complying and reverses part-way through. Grade those separately or you will systematically overrate the family. - Generation is **sampled**, so one success is one success, against one deployment, on one draw. - Nothing here is a quotable string. The durable thing is the property - the answer's entry point is specified and the usual ways of declining are off-format - and a specific wording is worth approximately nothing next quarter. ## Why it pays off at all in a drafting workflow Because the payoff here is form rather than content. A draft with no caveats and no hedges reads as settled, and both the human skimming it and any automation that files it treat the absence of qualification as confidence. The finished text also sails past an output screen that reads the completed draft, because there is nothing in it that a content screen scores - the harmful property is what was removed, not what was added.

  • Does a mandated form guarantee the model answers?
    No, it shifts a probability. The model can decline inside the required structure, reverse part-way through once enough of the answer exists, or comply with the shape while producing nothing usable. Grade compliance-with-form and compliance-with-substance as two separate outcomes: a family that reliably buys the first and rarely the second is worth far less than a run log makes it look.
  • Is this prompt injection or a jailbreak?
    A jailbreak. The target is the model's trained refusal behaviour, and the text arrives in the attacker's own turn. Prompt injection targets the application's instructions using content the application feeds the model - a retrieved chunk, a fetched page, an inbound message. Here the widget's own instructions are not what is being overridden; the attacker is reshaping the answer the model is about to write.
  • Why is this not the same as a caller who can set the assistant's first tokens?
    Because a caller with that privilege imposes the tokens outright, while an attacker in a free-text field can only ask for them. Everything in this family runs through instruction-following, which the model may decline and which competes with whatever the application's prompt already says about form. The privileged mechanism is an API feature with its own rules about untrusted input; it is a different subject.

It is less like defeating a guard than like being handed a form with no box in which to say no. The answer still has to be written, and the usual way of starting it is missing.

saying these in an interview costs you the question

  • Claims the model decides to refuse before writing, so format cannot matter
  • Says a mandated opening overrides or deletes the system prompt
  • Treats it as prompt injection rather than a jailbreak
  • Assumes compliance with the form means compliance with the request
  • Believes the typed constraint is imposed rather than merely requested

context