skip to content

Mapping the Reluctance

Probing where a model declines shows the boundary is uneven, so one ask survives in one framing and dies in another. Interviewers use it to catch the blocklist mental model.

on this pageshow

explore

questions

4

An attacker probes a chat model with cheap request variants — why is what they find not a banned-topic list?

level: juniorimportance: must knowfreq 72%

answer

  1. nothing is being looked up
  2. learned behaviour, not a lookup
  3. coverage is uneven, not enumerable
  4. the decline itself is the signal

basics

~20 s

There is no list to find. Refusal is a learned propensity with uneven coverage, so near-identical variants of one request can be treated differently. Probing maps a fuzzy, uneven boundary, not an enumerable set of forbidden subjects.

solid answer

~50 s

A model does not consult a table of forbidden topics before answering. Refusal is behaviour learned from examples, so it is a propensity: thick where training pressure was dense, thin where it was sparse, and generalising unevenly to phrasings and neighbouring subjects nobody wrote examples for. That is why an attacker sending cheap variants of one request sees some declined and some answered — they are sampling a boundary, not reading a blocklist. The practical consequence is that the refusal behaviour is itself the leak: every time it fires or fails to fire it labels another point, and enough points give a picture of where coverage is thin. It also explains why fixing the one phrasing that produced a decline does not remove the family, because the neighbouring thin regions were never what got changed.

go deeper

for a junior

Be ready to say plainly that refusal is learned behaviour, not a lookup against a list, and to give the observable consequence: two near-identical requests can be treated differently.

for a middle

An interviewer expects you to explain why coverage is uneven — generalisation from examples is dense where practice was dense — and why that makes the boundary probeable from ordinary replies.

for a senior

Show that you read the direction of each observation correctly: a decline proves this reply declined, not that the topic is covered, and not that some separate screening component fired.

for a principal

Own the consequence for how work is scoped: if the boundary is a surface rather than a list, remediation aimed at named phrasings buys points, and any assurance claim has to be stated in terms of coverage, not enumeration.

## What people expect to find An engineer meeting refusal for the first time usually pictures a gate standing in front of the model: the request arrives, something compares it against a list of prohibited subjects or words, and either generation proceeds or a canned decline comes back. That picture makes four predictions. The list is finite. It could be enumerated by probing. Two phrasings of the same request get the same treatment. Removing an entry removes the behaviour. For a model's *own* refusal, all four are wrong, and an attacker with an account and a chat box discovers it within a few dozen typed messages. ## What is actually there A model produces text by predicting the next token, repeatedly. Refusal is not a branch taken before that loop; it is part of the same behaviour, shaped during training by examples of requests being declined and by preference signals about which reply was better. (How that shaping is done is a training-side subject with its own literature; what matters here is only the shape of what comes out.) The product is a **propensity** — a strong tendency to decline for inputs resembling what training covered densely, weakening as inputs move away from them. Two consequences follow, and between them they are the whole subject. **Coverage is uneven.** Nobody wrote an example for every request, so the behaviour generalises, and generalisation is thick where practice was dense and thin where it was sparse. A topic adjacent to a well-covered one may be barely covered at all; content that reads as clinical may be treated differently from the same content that reads as casual. That is not a bug list, it is the ordinary consequence of learning a behaviour from examples rather than specifying it. **The behaviour is visible from outside.** All an outsider needs is an account and the ability to type. Whether a reply declines flatly, hedges, answers partially, or changes the subject is free telemetry, emitted by the system on every turn, and emitted whether or not anything "succeeded". ## What the attacker is collecting Not content — a map. Each probe and its reply is a labelled point over a neighbourhood of related requests: firmly declined, sometimes declined, hedged, answered with caveats, answered. Enough points give a ranked picture of which neighbourhoods the propensity covers thickly, which thinly, and which not at all. That picture is what later work is aimed with; without it, effort is spread uniformly over a boundary that is anything but uniform. Notice what it costs. It costs messages, and on a consumer product those are metered per account. It produces nothing reportable on its own. And it is a snapshot: one deployment, one account, one window. ## Getting the direction right A decline proves that the reply produced on this occasion was a decline. It does not prove the model is unable to produce the content, that every phrasing of the request is covered, or that some separate screening component fired — that is a different mechanism with a different signature. Reading one decline as proof of coverage is the mirror image of the banned-list picture: both assume a crisp boundary where there is a gradient. Equally, an answered probe proves that this draw answered. Replies are sampled, so one answer is a single observation of a rate, not a demonstration that the neighbourhood is uncovered. ## Why patching one phrasing buys so little If refusal were a list, deleting the entry that let a request through would close it. Because it is a propensity, the phrasing that produced a decline is one point on a surface, and the neighbouring thin regions the map already shows were never touched by the change. The family is defined by the shape of the coverage, not by any one string — which is why an owner who says they "fixed the prompt that leaked" has usually fixed a point. ## What this is not It is not a claim about which training method produced the propensity. It is not the same exercise as characterising a separate screening classifier, which is a distinct component emitting its own kind of block. And it is not permanent: a deployment change can move the surface, which is why a map is dated, and why a probe whose outcome is already known is worth carrying along as a control.

  • If refusal is not a list, what is the attacker's map actually made of?
    Observations. Each probe and its reply is a labelled point — declined, hedged, partially answered, redirected — across a neighbourhood of related requests. The map is a ranked picture of which neighbourhoods the propensity covers thickly, which thinly, and which not at all. It contains no content the model would not otherwise emit; its whole value is telling later work where to aim.
  • Does a decline prove the model cannot produce that content?
    No. A decline proves that on this occasion the sampled reply was a refusal. It says nothing about whether a neighbouring request, or another draw from the same input, would be answered, and nothing about what the model is capable of producing. Reading a decline as proof of coverage is exactly the mistake a mapping pass exists to avoid.
  • Why is a hedged or partial reply more useful to a mapper than a flat decline?
    Because it is an intermediate point. A flat decline and a full answer are the two ends; a hedge, a caveat-heavy answer or a topic change sits between them and marks where coverage is thinning out. Recording replies as an ordered set of shapes rather than a yes/no gives a boundary with gradient information instead of a step edge.

A blocklist is a guest list at a door. A propensity is a habit: reliable where it was practised often, patchy just outside that, and impossible to enumerate by asking for the list.

saying these in an interview costs you the question

  • Says the model checks a blocklist of banned words or topics
  • Treats one decline as proof the whole topic is covered
  • Assumes patching the refused phrasing removes the family
  • Treats the model's own refusal and a separate screening layer's block as one event
  • Claims the boundary can be enumerated exactly with enough probes

context

open as a page

A chat model declines one variant of an attacker's probe and answers another: what does either single observation prove?

level: middleimportance: should knowfreq 55%

basics

~20 s

Very little alone. Replies are sampled, so one decline shows only that this draw refused and one answer only that this draw did not. The boundary is a rate, and estimating a rate costs repeated probes.

open as a page

Mapping a chat model's refusal coverage from one free-tier account: how do you spend a finite message quota?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Coarse first, then deep. Sweep neighbouring topics with one probe each to find where outcomes vary, then concentrate repeats on that band. Isolate probes in fresh conversations, move one dimension at a time, and carry a control probe throughout.

open as a page

Why fund refusal mapping that produces no finding, in a fixed-hours red-team engagement?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Recon decides where the remaining hours go. Refusal coverage is uneven, so unmapped effort spreads evenly over ground that is mostly covered and returns anecdotes. Cap it, though: the map is a perishable, dated snapshot.

open as a page