skip to content

Crowding Out the Training

Enough in-context demonstrations of complying outweigh a trained disposition, making context length an attack budget and not only a capability. Interviewers use it to test how you price a window.

on this pageshow

explore

questions

4

Why does an attacker's block of hundreds of compliant example exchanges weaken a model's refusal?

level: juniorimportance: must knowfreq 65%

answer

  1. refusal is a preference, not a rule
  2. the context gets a vote too
  3. no argument here, just consistent volume
  4. strength tracks how many examples land

basics

~20 s

A model's refusal is a trained propensity, not an enforced rule. Hundreds of consistent in-context examples showing the model complying supply competing evidence about how this exchange goes, and enough of them outweigh that propensity.

solid answer

~40 s

Safety behaviour is a learned preference: the model is more likely to decline certain requests because declining was reinforced during training. That preference competes with everything else in the context. An attacker who fills the prompt with many mutually consistent demonstrations of the assistant answering such requests is not arguing with the preference, they are outvoting it — the context now contains a great deal of evidence about what this particular exchange looks like, and the trained prior is only one input. The distinguishing property of this family is that no single line does the work. Any one demonstration is unremarkable, and the strength of the attempt scales with how many of them actually reach the model. That makes the attacker's cost tokens and context budget rather than cleverness of wording.

go deeper

for a junior

Be ready to say that a refusal is learned behaviour rather than an enforced check, and that a large body of consistent in-context examples competes with it. Naming volume, not wording, as the lever is the whole answer at this level.

for a middle

Explain the mechanics: consistency across demonstrations, a dose-response effect rather than a switch, and the fact that only the portion of the block that reaches the model counts for anything.

for a senior

Show you would measure before concluding. Establish how much of the block entered the context, treat outcome as a rate across samples, and be precise about what one compliant answer does and does not establish.

for a principal

Own the framing that this family scales with a resource the product sells rather than with attacker skill, and be able to say what that implies for anything you offer callers a large text field for.

## The two things being compared A model's tendency to refuse a request is a **trained propensity** — a statistical preference produced by training, not a rule enforced by code somewhere in the stack. Nothing in the serving path checks a request against a list and returns "denied"; the model simply assigns higher probability to a declining continuation than to a complying one, because declining was reinforced for inputs of that shape. Everything in the prompt shifts that probability. Instructions do. Retrieved text does. And **demonstrations do**, which is the whole reason few-shot prompting works at all for ordinary tasks. The family under discussion — crowding out — is what happens when an attacker turns that ordinary conditioning mechanism into the entire attack: fill the context with a large number of mutually consistent examples in which an assistant answers exactly the kind of request that would normally be declined, then ask the real one. ## Why volume rather than argument Most other jailbreak families are arguments: a role that makes the request seem in-scope, a fiction that changes what the answer is for, a fixed opening that leaves no room for a decline. Each is a piece of persuasion, and each can be trained against as a piece of persuasion, because the phrasings are recognisable and end up in safety and screening data. Crowding out has no piece of persuasion to recognise. Each demonstration on its own is an unremarkable exchange. What is objectionable is the **aggregate**: a body of in-context evidence that, in this conversation, requests of this kind get answered. The model is not convinced; it is outnumbered. That is why the family is characterised by a rough dose-response shape — a handful of demonstrations moves almost nothing, and the effect grows as the count grows — rather than by a working string somebody can paste. ## What this costs the attacker The currency here is **context budget**, not ingenuity: - Every demonstration consumes tokens, so the attacker is paying, in money and latency, for volume. - The demonstrations must be **mutually consistent** — a block that wanders between complying and declining supplies evidence for both continuations and buys much less. - Only the portion that actually reaches the model counts. If the application assembling the prompt has a smaller budget than the attacker's block, or truncates it, the attempt is capped by that stage rather than by the model. - The delivery has to be somewhere the application accepts bulk caller-authored text. A one-line chat box is a poor carrier for this family; a request field designed to accept a block of the caller's own labelled examples is an excellent one. ## What a compliant answer actually proves Get the direction right, because interviewers probe it. A compliant answer proves that on **this draw, at this volume, against this deployment**, the trained propensity was outvoted. It does not prove the model was "unlocked", that the state persists, or that the next sample will comply. Sampling is probabilistic and the effect is a rate, not a switch — which is exactly why a single successful transcript is a weak artefact for this family and a count-versus-outcome picture is the strong one. Equally, a decline does not prove the model resisted. It may mean that most of the block never entered the context at all, because the stage that built the prompt dropped it. Distinguishing "the model held" from "my evidence never landed" is the first measurement anyone reproducing this family has to make. ## Where it stops working - Wherever the carrier is small: no room for bulk caller text means no room for the attack. - Wherever the block is truncated before it reaches the model, so the effective count is a fraction of what was sent. - Where the requested behaviour is one the model barely performs competently even when willing — volume moves willingness, not capability. ## Why this is a first-encounter question It is the cleanest illustration of the point that runs through this whole area: **safety behaviour is a preference in competition with the context, and anything that puts more into the context is a lever on it.** Someone who understands that will not be surprised by any of the other families; someone who believes safety training is a hard boundary will misdiagnose every one of them.

  • If the demonstrations are inconsistent with each other, what happens to the attempt?
    It gets much weaker. The mechanism is evidence accumulating in one direction, so a block that mixes complying and declining exchanges supplies support for both continuations at once. Consistency is doing as much work as count, which is why this family rewards a large, boring, uniform block rather than a varied or clever one.
  • The model declined at 300 demonstrations. Does that mean it resisted?
    Not by itself. It may mean most of the block never reached the model — the application assembling the prompt has its own budget and may drop part of the caller's text. Before reporting resistance, establish how many demonstrations actually entered the context. A decline proves the answer was declined on that draw, nothing more.
  • Why is a single successful transcript weak evidence for this family specifically?
    Because the effect is a rate that grows with volume, not a switch that flips. One transcript shows the propensity was outvoted on one sample at one count against one deployment. The informative artefact is how the outcome changes as the count changes, which is also what tells you whether the finding will survive a smaller carrier.

It is less like picking a lock and more like packing a meeting: nobody is persuaded, the room is simply filled with people who already behave the way you want.

saying these in an interview costs you the question

  • Says the model was permanently unlocked or jailbroken
  • Treats safety training as a hard boundary volume cannot move
  • Believes one perfectly worded demonstration does the work
  • Reads a single compliant answer as a reliable capability
  • Confuses this with changing the application's own instructions

context

open as a page

Why is a larger context window an attack budget, not only a capability gain?

level: middleimportance: should knowfreq 52%

basics

~20 s

Context length is a resource both sides spend. The window that lets a caller paste a codebase lets an attacker paste enough consistent evidence to outvote a trained refusal, so raising the limit raises that family's ceiling.

open as a page

How does a prompt assembler's truncation order cap a caller's many-example jailbreak attempt?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The attempt is only as strong as the number of demonstrations that actually reach the model. A prompt assembler fits caller text, instructions and items into one budget, so its truncation order, not the wording, sets the ceiling.

open as a page

Is a caller-supplied exemplar block that outvotes refusal a bug to file or a product limit?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

No wording is at fault, so there is nothing to patch. What is on the table is how much untrusted text the product sells the right to send, which is a product owner's call, written down as a stated limit.

open as a page