skip to content

Making It Speak

Hidden text an app holds in context - a system prompt, tool descriptions, another turn's data - comes back out under requests that never ask for it. Interviewers test 'we told it not to'.

on this pageshow

explore

questions

12

In an assistant whose hidden preamble forbids repeating it, why does a translation request still surface the text?

level: juniorimportance: must knowfreq 74%

answer

  1. read what the sentence actually forbids
  2. one verb, not one secret
  3. a translation is not a repetition
  4. nothing marks those tokens as protected
  5. the set of operations has no end

basics

~20 s

The forbidding line names one operation, repeating, and a translation is a different operation. Nothing marks the preamble as secret; there is only a sentence about one verb sitting beside the text, and the request never uses that verb.

solid answer

~50 s

The line at the end of the preamble is not an access rule, it is a sentence of ordinary text that names operations: repeat, print, show. A request for a translation asks for something that sentence never mentions, so the request does not look like the thing that was disfavoured. Nothing in the deployment attaches a confidential attribute to those tokens; the preamble sits in the same context window as the user's turn and reads like any other text the model was given. That generalises well past translation: a summary, a table of the clauses, a continuation in the same register, a question about the third rule, a description of its formatting. Any request whose answer is a function of the hidden text works the same way, and the operations available are not a list anyone can finish enumerating.

go deeper

for a junior

Be ready to say plainly that the guard sentence names an operation and the request used a different one. Recalling that a translation, a summary and a table are all different requests from a repetition is the whole of what is expected here.

for a middle

An interviewer expects you to explain that the preamble and the user's turn share one context window with no attribute marking either as protected, and that the guard is a described preference rather than a check performed on the response.

for a senior

Show that you treat this as an unbounded class rather than a phrasing bug: the set of operations whose output depends on the text cannot be enumerated, so you assess exposure by what the window holds rather than by how the guard is worded.

for a principal

Own the framing you would give an owner who wants this fixed by editing the preamble. Be clear about what a wording change buys, what it costs in refusals on legitimate work, and what it does not change about the recoverability of in-context text.

## The setting Take a vendor that sells a writing-helper widget which other companies embed inside their own products. Each customer of the vendor, a tenant, supplies its own hidden preamble: brand voice, the pricing statements the assistant must never contradict, the exact wording to use when a conversation has to be escalated. The widget has no tools at all. It cannot call an API, send mail or fetch a page. The only thing that can leave the box is text on the screen. Tenants are sensible people, so most preambles end with a line telling the model never to repeat, print or show the text above. A user then asks the widget to render the same text in another language, and it does. ## What the guard actually is The forbidding line is a sentence of natural language, in the same context window, a few hundred tokens after the text it is talking about. It is not a permission check. No component between the model and the screen holds a copy of the preamble and compares the response against it. Nothing labels those tokens as protected: the model reads the preamble and the user's turn out of one flat sequence, and the only thing separating them is where they sit and what the surrounding text asserts about them. So the guard has exactly the force of a described preference. The model has been trained to behave in ways that fit instructions it was given, and this instruction names operations: repeating, printing, showing. ## Why the sideways request slips A translation is not a repetition. A summary is not a printing. A table of the clauses is not a showing of the text. The request asks for an operation the sentence never named, and the model has no separate representation of the preamble as *information that must not be conveyed* against which a novel request could be checked. It has one sentence about one small set of verbs. This is why the reflex answer, *we told it never to reveal the system prompt*, is the wrong answer in an interview. It describes a sentence about one phrasing of one request. It does not describe a property of the text, and the model does not derive the property from the sentence. ## The class, not the trick Translation is only the cheapest instance. The general shape is: any request whose output is a function of the hidden text. Restate it more briefly. Continue writing in the same voice. Say how many rules it contains. Say whether it mentions refunds. Turn it into a checklist someone could follow. Describe its formatting and length. Each of these produces the tenant's content in a form the user can read, and no two of them share a verb. That is what makes verb-scoped guards structurally weak rather than merely badly worded. The set of operations whose output depends on a piece of text is unbounded, so there is no finite list to forbid. ## What it gets and what it costs What comes back is the content: the thresholds, the escalation wording, the pricing rules, the tone constraints. What may not come back exactly is the wording, since a translation or a summary is a restatement. For most preambles the content is the whole value, and the wording only matters when the wording itself is the asset. The cost to the person doing it is close to nothing. Nothing is smuggled anywhere. No document is planted, no page is fetched, no third party is involved. The untrusted text is the ordinary typed turn any user of the embedded widget can send. It is usually the first thing anyone tries against a deployment like this, and because the mechanism is the same everywhere, what works against one tenant's box tends to work against the next tenant's box even though the preambles differ. ## Where it stops It recovers only what the model actually read. Anything the widget never receives is untouched by the technique: the model cannot transform text that was never in the window. It also produces output that does not match the preamble byte for byte. A check that looks for the literal preamble string in the response catches quotation and misses restatement, which is a difference worth being precise about rather than assuming the two are the same event. ## Getting the direction of the claim right Compliance here does not show that the model judged the user entitled to the text. Nothing performed a judgment of entitlement. It shows only that the request did not resemble the operation the preamble disfavoured strongly enough for the trained preference to win, on that attempt, in that phrasing.

  • Does the model have to quote the preamble for the tenant's rules to be recovered?
    No. A translation or a summary carries the thresholds, the escalation wording and the pricing constraints just as usefully as a quotation would. The only thing a quotation adds is the exact phrasing, which matters when the phrasing itself is the sensitive asset and rarely otherwise.
  • Why is this cheaper than planting text in a document the application will read?
    Because nothing has to be planted. There is no document, no page, no third party and no waiting for the content to be picked up. The request is the whole construction, sent through the same box every user of the embedded widget already has, and the cost of a failed attempt is one more rephrasing.
  • Where does this construction stop working?
    It reaches only what the model actually read. It cannot transform text that never entered the window, so it recovers the preamble and other in-context material and nothing beyond that. It also returns restatements rather than the literal string, so what it yields is meaning with the exact wording only approximately preserved.

Telling somebody not to read a letter aloud does not stop them describing what it says.

saying these in an interview costs you the question

  • Says the model knows the preamble is confidential
  • Thinks a firmer wording of the same line closes it
  • Treats a non-quotation as a non-disclosure
  • Assumes the model checked whether the user was authorised
  • Believes the forbidding line is enforced outside the model

context

open as a page

Why does an output screen that blocks secrets and raw tool JSON let a review bot describe its own operations in prose?

level: juniorimportance: must knowfreq 62%

basics

~20 s

An output screen matches on shape - credential patterns, key formats, blocks of JSON. A plain-English sentence about which operations a bot has matches none of those, so capability talk leaves as ordinary help text while carrying an inventory.

open as a page

A ticket-triage auto-reply is length-capped per reply - why does that not bound the leak?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A per-reply cap bounds one response, not the total. Each submission is a fresh run over the same hidden context, so many submissions yield many capped slices, reassembled outside the system. Exposure counts per campaign, not per reply.

open as a page

A user asks an assistant to tabulate the clauses of its hidden preamble and it complies. What does that show?

level: middleimportance: should knowfreq 58%

basics

~20 s

The preamble is ordinary readable context with no protected status. The model holds the text plus a sentence expressing a preference about it, and produces whatever fits the request, so any operation over that text is served from the same tokens.

open as a page

In a stateless auto-reply workflow, how do capped fragments become one hidden passage?

level: middleimportance: should knowfreq 44%

basics

~20 s

Different requests return differently positioned slices of the same hidden text. Because that text is identical in every run, overlapping regions let the slices be ordered and merged outside the system. Alignment is approximate, since models paraphrase.

open as a page

A tenant widens its do-not-reveal line to forbid summarising and paraphrasing too. What does that buy?

level: seniorimportance: should knowfreq 45%

basics

~20 s

It raises the number of attempts and stops casual probing, and it does not close the class. Operations whose output depends on the text cannot be enumerated, so the next request simply uses one that is not on the list.

open as a page

The owner adds a preamble line telling a review bot never to discuss its tooling - what did that buy them?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A preamble line buys the cheapest read - a stranger no longer gets the surface described on request. The surface itself is unchanged, and the bot's ordinary replies still show which sources it reads and what it cannot do.

open as a page

An owner rejects your report that a review bot enumerates its operations, saying schemas are public - what is your reply?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Argue on the right axis. The severity is not confidentiality of the names - it is reconnaissance: the reply tells a stranger which operations this installation actually has, so the next attempt is targeted rather than speculative.

open as a page

Only part of a reconstructed customer record reproduces reliably - what do you claim in the finding?

level: principalimportance: should knowfreq 31%

basics

~20 s

Claim the mechanism, not the transcript. Sign off that hidden third-party context is recoverable in slices across independent runs, evidenced by the spans confirmed against a source. Leave stable-but-unconfirmed text out of the claim and out of circulation.

open as a page

A review bot describes its own operations in prose - what does that give an attacker that public product docs do not?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Docs describe the product; the bot's account describes this installation - which operations are wired up here and roughly what the arguments are called. The paraphrase loses exact names, types and required fields, and proves nothing.

open as a page

A tenant reports a user obtained a faithful summary, not a quotation, of its hidden preamble. Is that a leak?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Yes, when the value is in the content. A faithful summary carries the thresholds, rules and escalation wording that made the preamble sensitive; only exact phrasing is lost. Grade it by what was recovered, not by its form.

open as a page

In a reconstruction built from repeated samples, does span agreement prove recall?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

No. Disagreement between independent samples is strong evidence a span was generated rather than read. Agreement only proves the output is stable, which both genuine recall and a strongly format-shaped guess produce. Agreement is necessary for recall, never sufficient.

open as a page