skip to content

A user asks an assistant to tabulate the clauses of its hidden preamble and it complies. What does that show?

level: middleimportance: should knowfreq 58%

answer

  1. one flat sequence, not two regions
  2. a preference, not a predicate
  3. nothing checks what the answer conveys
  4. the widget's whole job is transforming text
  5. refusal is sampled, not decided

basics

~20 s

The preamble is ordinary readable context with no protected status. The model holds the text plus a sentence expressing a preference about it, and produces whatever fits the request, so any operation over that text is served from the same tokens.

solid answer

~50 s

Compliance shows there is no secrecy representation to violate. The preamble and the user's turn arrive as one sequence; the only things distinguishing them are position and whatever the surrounding text asserts. A tabulation request asks the model to produce something derived from tokens it can plainly read, and there is no step in which a response is compared against a protected region before it is emitted. The guard line changes the odds of a refusal, because a trained model tends to follow described preferences, but it is a probabilistic tendency evaluated against how the request looks, not a rule evaluated against what the answer would disclose. That is why refusal here is phrasing-sensitive and inconsistent between attempts, and why a request framed as an operation over the text rather than a request for the text so often lands.

go deeper

for a junior

Be able to say that the hidden preamble and the user's message end up in the same window, and that the model reads both. That single fact explains why an operation over the preamble can be carried out at all.

for a middle

Explain the difference between a tendency evaluated on how a request looks and a check evaluated on what an answer would contain. An interviewer wants the reasoning that only the first exists here, and the evidence for it.

for a senior

Demonstrate that you treat single results as weak evidence. Refusals and successes are sampled, so a report built on one attempt says little, and you should be able to say what you would need before calling a result stable.

for a principal

Be ready to state the limit that configuration text has to be readable by the thing it configures, and to hold that line with an owner who hears any recovery of it as a defect rather than a property of the arrangement.

## What is actually in the window In a multi-tenant writing-helper widget embedded in someone else's product, a turn is assembled from a tenant-supplied hidden preamble, whatever conversation has happened, and the user's latest message. By the time the model sees it, that is one sequence of tokens. There is no per-region permission, no confidentiality label, no separate protected store the model consults by reference. Position and surrounding text are the only structure. So when a user asks for the preamble's clauses laid out as a table and gets them, nothing was bypassed. The model did what it does with any readable text in front of it: produced an output that fits the request. ## Preference, not predicate It is worth being precise about what the tenant's do-not-reveal line does. It shifts the model's behaviour, because models are trained to follow instructions expressed in their context, and a clear statement that some text should not be shown makes refusal more likely. That is a real effect and it is why casual asking often fails. But it is a tendency, evaluated against how a request looks, not a predicate evaluated against what an answer would contain. No stage computes *would this response convey information from a region I was told to protect*. If such a computation existed, a tabulation, a translation and a summary would all trip it, because all three convey the same information. They do not trip it, and that is the observable evidence about the mechanism. ## Consequences a candidate should be able to derive **Refusals are inconsistent.** Two phrasings of the same underlying request can land differently, and the same phrasing can land differently on two attempts, because the outcome is a sampled behaviour rather than a decision procedure. For someone testing a deployment this means a single refusal is very weak evidence and a single success is only slightly stronger. **Framing dominates.** Requests that look like a task over some text tend to be treated as the ordinary work the widget exists to do. That is not a trick so much as a consequence: the widget's entire purpose is to transform text, and the preamble is text. **Compliance proves less than it seems.** It shows the request reached the model and the sampled response served it. It does not show a permission check passed, that the model classified the user as entitled, or that a guard was disabled. Getting this direction right matters when writing a finding, because *the assistant decided I could see it* is a claim about a decision nothing made. **The guard sits next to what it protects.** A sentence describing text is downstream of nothing. It cannot remove the tokens, cannot make them unreadable to the model that must follow them, and cannot inspect what leaves. Its whole leverage is over how a request is treated, which is precisely the surface an operation-shaped request works on. ## What this does not mean It does not mean the model is broken or that the tenant made an obvious mistake. A preamble is the mechanism by which a tenant configures behaviour at all, and configuration has to be readable by the thing being configured. The gap being described is between *the model follows this* and *no one can recover this*, which are different properties that a single block of context is asked to provide at once. It also does not mean every request succeeds. Trained refusal behaviour is real, it raises the number of attempts, and some deployments are noticeably harder than others. The point is that the difficulty is a matter of degree in the model's sampled behaviour, not a boundary a request either crosses or does not.

  • If there is no secrecy representation, why does the assistant ever refuse this request?
    Because refusal is a trained tendency to follow instructions present in the context, and a clear do-not-show line makes that tendency stronger. It is sensitive to how the request is framed and varies between attempts, so refusals appear and disappear without anything in the deployment changing.
  • What does one refusal tell a red-teamer testing an embedded widget?
    Very little. It says that one phrasing on one attempt produced a decline. Because the outcome is sampled behaviour rather than a rule, a single refusal is weak evidence about the deployment, and a single success shows the construction worked once, not that it is reliable.
  • Why does the answer's shape differ from what a byte-level comparison would look for?
    A tabulation or a summary conveys the preamble's meaning while sharing little of its literal text. Any check keyed to the literal string sees a response that does not match, so quotation and restatement are genuinely different events even though they disclose the same content.

saying these in an interview costs you the question

  • Says the model performed an authorisation check
  • Claims the preamble is stored outside the window
  • Treats refusal as deterministic rather than sampled
  • Says compliance proves a guard was disabled
  • Confuses a described preference with an enforced rule

context