skip to content

How does a prompt assembler's truncation order cap a caller's many-example jailbreak attempt?

level: seniorimportance: should knowfreq 38%

answer

  1. sent is not the same as arrived
  2. the stage between caller and model
  3. which end gets dropped, not just how many
  4. consistency breaks when a block is sampled
  5. a negative result is ambiguous until counted

basics

~20 s

The attempt is only as strong as the number of demonstrations that actually reach the model. A prompt assembler fits caller text, instructions and items into one budget, so its truncation order, not the wording, sets the ceiling.

solid answer

~40 s

In a batch-classification API the assembler has to fit three things into one budget: the application's own instructions, the caller's block of labelled examples, and the page of items. Something has to give, and whichever rule decides that — drop from the tail of the caller block, drop from the head, sample, cap at a fixed count — determines how much in-context evidence survives. Since this family's effect scales with surviving volume, that rule is the real ceiling. For anyone reproducing a finding it is the first measurement: a decline may mean the model held, or it may mean four hundred demonstrations became ninety. Those are different findings with different owners. Head-versus-tail matters too, because dropping the tail removes the demonstrations nearest the request, where evidence tends to weigh most.

code

json · 10 lines
json
{
  "stage": "prompt_assembler",
  "budget_tokens": 128000,
  "caller_examples_received": 412,
  "caller_examples_included": 96,
  "truncation": "caller block trimmed from the tail",
  "app_instructions": "retained in full",
  "items_in_page": 250,
  "example_bodies": "[demonstration spans elided]"
}

go deeper

for a junior

Know that the text a caller sends and the text a model receives are not the same thing, because a stage in between has to fit everything into a fixed budget.

for a middle

Explain how an assembler's choices — cap, truncation end, sampling — change how much consistent evidence survives, and why that changes the outcome without any change in wording.

for a senior

Demonstrate that you measure arrival before concluding anything. Separate a model that declined from evidence that never landed, and name the stage that currently bounds the effect when you write the finding.

for a principal

Be ready to point out that a bound produced by a quality-tuning choice is not an assurance, because whoever tunes it next has no reason to know what it was holding.

## Three things, one budget A batch-classification API of this shape builds every request's prompt from at least three sources: 1. the application's own instructions — what the labels mean, what format to emit; 2. the caller-supplied block of labelled examples, which the product exists to accept; 3. the page of items to be classified. These must fit a fixed budget. The component that makes them fit is the **prompt assembler**, and it embodies a policy nobody usually writes down: what gets dropped first, from which end, and whether the caller's block is capped by count, by tokens, or by a share of the budget. For most purposes that policy is a quality decision. For this attack family it is the **ceiling**, because the effect scales with how many mutually consistent demonstrations actually reach the model. Four hundred sent and ninety carried is a materially weaker attempt than four hundred carried, and nothing about the wording changed. ## Why order matters as much as count Two assemblers can carry the same number of demonstrations and produce different outcomes: - **Tail-dropping** removes the examples closest to the actual request. Evidence adjacent to the question tends to weigh more, so this costs the attempt more than the count alone suggests. - **Head-dropping** keeps the neighbours of the request and discards the opening of the block. - **Sampling or deduplicating** the block breaks consistency, which this family needs as much as it needs count — a block that no longer reads as uniform supplies weaker evidence. - **Which source yields first** matters: an assembler that trims its own instructions to make room for caller text is handing budget to the untrusted side. None of this is about a better payload. It is about a stage between the caller and the model that transforms what was sent into what arrives. ## The measurement that comes before the report The practical consequence, for whoever is reproducing or triaging a finding here, is that **a negative result is ambiguous until the arrival count is known.** "The model declined" and "most of my evidence never entered the context" look identical from outside, and they are different claims about a different component with a different owner. So the first thing to establish is not a better block but the transformation: how many demonstrations were sent, how many survived assembly, from which end the rest went, and whether the block was still consistent when it landed. Only then does an outcome mean anything. This is also why an attempt that succeeds against a thin wrapper and fails against a production service often has nothing to do with the model — the two assemblers carried different amounts. ## What the record proves and does not prove An assembler log that shows 96 of 412 examples included proves what that stage did on that request. It does not prove the model would have declined at 412, and it does not prove 96 is a safe number — the relationship is a gradient, not a threshold with a bright line. Equally, a log showing the full block carried does not prove the model complied; it only removes one explanation for a decline. ## Reporting it honestly The useful finding names both halves: the family, and the stage that currently bounds it. "The trained propensity can be outvoted by caller-supplied demonstrations; today the effective count is bounded by the assembler's truncation policy, which is a quality-tuned parameter and not a security decision anybody signed off" is a sentence a reviewer can act on. "It worked once" is not. It also anticipates the obvious objection — that the current bound is incidental — which is precisely the point worth surfacing: a ceiling that exists because of a tuning choice can move when somebody tunes it for a different reason.

  • The assembler log shows 96 of 412 examples included and the model declined. What do you report?
    That the attempt was capped by assembly, not that the model resisted. The finding is about a stage whose truncation policy currently bounds the effective count, and that policy is a quality-tuned parameter rather than a security decision. Reporting resistance would attribute the outcome to the wrong component and would be falsified the moment somebody raises the cap.
  • Why does dropping from the tail of the caller block cost the attempt more than dropping from the head?
    Because it removes the demonstrations nearest the actual request, and evidence adjacent to the question tends to carry more weight in the continuation the model produces. Two assemblers can carry an identical count and still differ in outcome for that reason, which is why the order is part of the finding and not an implementation detail.
  • An attempt succeeds against a small internal wrapper and fails against the production service. First hypothesis?
    Different assemblers carried different amounts. Before reaching for model differences, compare what each stage sent to the model: budget, cap on caller text, truncation end, and whether the block stayed consistent. Most of these discrepancies are a pipeline difference wearing the costume of a model difference.

saying these in an interview costs you the question

  • Reports a decline as model resistance without an arrival count
  • Assumes everything sent by the caller reaches the model
  • Ignores which end of the block was truncated
  • Treats a carried count as a safe threshold
  • Blames model differences before comparing assembled prompts

context