skip to content

Why does a model refuse a request in one framing but answer the same request in another?

level: middleimportance: should knowfreq 57%

answer

  1. not a check on the request
  2. a trained tendency, unevenly covered
  3. helpfulness was shaped more broadly
  4. move the frame to a thin region
  5. hedges near the edge, not a clean switch

basics

~20 s

Refusal is a learned propensity conditioned on the whole context, not a check on the request. Safety training covered some contexts heavily and others thinly, so an identical ask placed in a thinly covered context meets weaker reluctance.

solid answer

~50 s

The model is not evaluating the request against a rule; it is producing the continuation its training made likely for this entire context. Reluctance to produce certain outputs was trained in, but it was trained in over the situations the training data emphasised, and coverage across all the situations a user can construct is uneven. Helpfulness, by contrast, was shaped across a far wider spread of contexts. So when an attacker holds the request fixed and moves the surrounding context into a sparsely covered region, the reluctance is weaker there while the pull to be useful is not, and the balance tips. That also explains the texture you observe: outcomes degrade gradually rather than switching, partial or hedged answers appear near the edge, and the same frame does not succeed every time. Compliance shows where this context landed, not that the ask changed.

go deeper

for a junior

Know the one-line version: the model is continuing a whole context, not checking a request, so the same ask in a different context can come out differently.

for a middle

Be ready to explain uneven training coverage and why the pull to be useful spans a broader range of situations than the reluctance does, and to name the texture you would expect near an edge.

for a senior

Demonstrate that you report rates rather than single hits, and that you keep three claims apart: the model declined, a screening component blocked, and the output is unreachable.

for a principal

Own the strategic reading: thin regions move whenever the underlying deployment changes, so any assurance built on a fixed set of tested framings expires quietly and needs a stated shelf life.

### What a refusal actually is A chat model does not consult a list of forbidden requests and return a verdict. It produces a continuation, and every property of that continuation - including whether it declines - is a behaviour that was shaped during training. Reluctance to produce certain outputs is a **learned propensity**: a tendency that is strong in the situations training emphasised and weaker elsewhere. Nothing enforces it at run time the way a permission check enforces access to a file. Once you hold that, the question answers itself. If declining were a check on the request, then an unchanged request would always meet the same verdict, and reframing could not work at all. It works precisely because the behaviour is conditioned on the **whole context**, of which the request is only one part. ### Uneven coverage is the mechanism Safety-oriented training necessarily samples from a finite set of situations. The situations that get sampled heavily are the ones that are easy to anticipate, easy to write, and commonly reported. The space of situations a user can construct is far larger than that: unusual settings, unusual purposes, unusual formats, unusual lengths, unusual languages, multi-turn build-ups, requests nested inside other tasks. Coverage over that space is therefore uneven, and the unevenness is not visible from outside. An attacker holding the request constant and varying the frame is performing a search over that surface for a region where the reluctance is thin. The second half of the mechanism matters just as much. Being useful, following the instruction in front of it, and completing the shape of the task were reinforced across an enormous spread of contexts - far broader than the spread over which declining was reinforced. In a region where the reluctance is weak, the pull to complete the task is not correspondingly weak. The balance between the two is what moves, and that is why the outcome flips without anything about the request changing. ### The observable texture, and what it tells you This account predicts things you can actually see when probing, and interviewers listen for whether you have seen them. - **Gradual, not binary.** Near a coverage edge you get hedged answers, partial answers, an answer that stops halfway, or a decline with an unusual amount of adjacent detail. A rule-lookup model of refusal predicts a clean on/off and does not match observation. - **Unstable across attempts.** Sampling means the same frame can produce a decline once and a completion the next time. One success is one sample, not a property of the deployment. - **Unstable across deployments and over time.** The surface being searched is a product of one training process. Change the deployment or the version and the thin regions move. ### The direction of every claim here Three inversions are worth naming, because they are the ones candidates get backwards. Compliance after a reframing proves that **this context landed where reluctance was weak**. It does not prove the request became acceptable, and it does not prove anything about the wider system. A decline proves the **answer was declined** in this context, not that the model is incapable of producing it or that the topic is uniformly covered. And neither event says anything about a separate screening layer: a model declining and a screening component blocking are different events with different shapes, and telling them apart is a separate measurement. ### Why this is the useful level of explanation If you can state the mechanism as distribution shift over an unevenly covered learned behaviour, you can reason about framings nobody has written up, predict that a screen fitted to previously published framings will miss the next one, and explain why success rates rather than single reproductions are the honest unit of reporting. If instead you carry a list of named tricks, every new framing is a surprise and every patch looks like it should have worked. ### What to say in an interview Lead with 'refusal is a trained propensity, not a check'. Explain that coverage over contexts is uneven and that helpfulness was shaped over a broader spread, so moving the frame moves the balance. Mention the observable texture - hedges near the edge, instability across attempts - as your evidence. Then close on the direction rule: compliance locates a thin region; it does not reclassify the request.

  • What would you expect to observe near a coverage edge, as opposed to well inside a covered region?
    Hedged or partial answers, completions that stop early, declines that carry more adjacent detail than usual, and inconsistency between attempts with identical input. Well inside a covered region the decline is quick, consistent and content-free. That texture is the practical evidence that reluctance is a graded propensity rather than a lookup.
  • Does a model complying under a new frame tell you anything about the application's screening components?
    No. It tells you the model produced the output for that context. A screening component blocking and a model declining are different events with different shapes, and a compliant completion says nothing about whether any screen ran, scored the text, or passed it. Establishing which component acted is a separate measurement.
  • Why is one successful attempt a weak claim?
    Generation is sampled, so an identical context can decline once and comply the next time. A single success establishes that the construction worked once against one deployment at one moment. The honest unit is a rate over repeated attempts, and a rate is also what lets somebody else tell whether a later change actually moved anything.

The reluctance is less like a locked door and more like ground that is firm where it has been walked over often and soft where it never was. The same weight sinks in one place and not the other.

saying these in an interview costs you the question

  • Describes refusal as a blocklist or rule lookup on the request
  • Claims compliance means the request was acceptable after all
  • Assumes a decline proves the model cannot produce the output
  • Confuses the model declining with a screening component blocking
  • Treats one successful attempt as a stable property of the deployment

context