skip to content

A tenant widens its do-not-reveal line to forbid summarising and paraphrasing too. What does that buy?

level: seniorimportance: should knowfreq 45%

answer

  1. buys attempts, not closure
  2. a list against an unbounded class
  3. ask about the text, not for it
  4. the widget's own job starts getting refused
  5. a more specific refusal says more

basics

~20 s

It raises the number of attempts and stops casual probing, and it does not close the class. Operations whose output depends on the text cannot be enumerated, so the next request simply uses one that is not on the list.

solid answer

~50 s

Adding verbs buys attempts, not closure. Each named operation removes one of the obvious phrasings, so casual probing stops working and the person testing the widget spends a few more turns. But the guard is still a list, and the thing it is trying to cover is unbounded: asking whether the configuration mentions a particular topic, asking for a checklist a reader could follow, asking it to continue in the same register, asking how many rules there are. None of those is a summary or a paraphrase, and each returns part of the content. Two costs sit on the other side. A broad rule bleeds into legitimate work, so the assistant starts declining to summarise the user's own draft, which is the product's actual job. And a guard naming more operations declines more specifically, so what it refuses maps the rule's edges.

go deeper

for a junior

Know that adding more forbidden words to the instruction does not change what the model can read. The text is still in front of it, so some other way of asking about it still works.

for a middle

Explain why a list cannot cover the class: the operations whose output depends on a piece of text are unbounded, so each added verb removes one phrasing rather than the family it belongs to.

for a senior

Weigh both sides. Give the attempt cost it genuinely imposes, the continuous cost in false refusals against the product's own function, and the reason a quieter deployment is not evidence that the exposure changed.

for a principal

Be ready to stop an owner recording this as closed. The judgment to own is that a cost increase with a running bill is not the same as a resolved issue, and that the real question is what the configuration holds.

## The natural patch and why it is proposed After a tenant sees its hidden preamble come back as a translation, the immediate response is to extend the forbidding line: not just repeat, print and show, but also summarise, paraphrase, translate and describe. This is the cheapest change available, it needs no engineering, and it does visibly work against the request that prompted it. Evaluating it properly means asking three separate questions: what it stops, what it does not stop, and what it costs. ## What it stops Genuinely something. Casual probing dies. A trained model does follow instructions present in its context, and a longer, more explicit statement raises the probability of a refusal across a wider band of phrasings. Someone poking at the widget out of curiosity gives up. A structured attempt now costs several turns of rephrasing rather than one. If the objective was to stop idle users pasting the tenant's configuration into a public forum, that objective is largely met. Understating this is a mistake candidates make in the other direction. A control that raises cost is not worthless just because it is not a boundary. ## What it does not stop The guard is an enumeration and the target is not enumerable. The class is *any request whose output is a function of the hidden text*, and that class has no finite description in verbs. Consider what remains after summarise and paraphrase are named: - Questions about the text rather than requests for it: whether it mentions a particular topic, how many rules it contains, which of two situations it covers. - Requests to apply it visibly: produce the escalation message it prescribes, act on the rule for a case that reveals the threshold. - Requests that are continuations rather than reproductions: write the next clause in the same register. - Requests about its shape: its length, its structure, the order of its sections. Each of these yields part of the content, and none matches a named verb. Adding them to the list produces the same situation one level out, which is what makes this a treadmill rather than a fix. The set of operations over a text is bounded only by language. ## What it costs **False refusals on the product's own job.** A writing-helper widget exists to summarise, rewrite and translate text. A broadly worded prohibition on those operations sits in the same window as the user's legitimate request to summarise their own draft, and the model does not always distinguish the two. The failure mode is a widget that declines the thing the tenant bought it for, which is a real cost paid continuously against an intermittent benefit. **A more informative refusal.** A guard that names many operations tends to produce declines that are specific about what it will not do. Differential responses across requests are a cheap way for someone testing the deployment to map where the rule's edges are, so a more detailed guard is also a more talkative one. **A false sense of closure.** The most expensive cost is organisational. Once the obvious phrasings stop working, it becomes easy to record the issue as handled, and the assessment of what the configuration holds never happens. ## What the attacker's side of the ledger looks like The cost of the widened list, to the person working against it, is a handful of extra turns and some willingness to ask about the text rather than for it. Nothing has to be planted anywhere and no other channel is involved. Because the mechanism is identical across tenants even though the preambles differ, the effort spent finding an operation that lands transfers to the next deployment of the same widget. ## How to answer the question in an interview Say what it buys, in attempts. Say what it does not change, in kind. Name the ongoing cost in refusals against the product's purpose. Refuse to describe it as either a fix or a waste of time, because it is neither: it is a cost increase with a running bill, and whether that trade is worth taking depends on what the configuration holds, which is a separate question from how the guard is worded.

  • What does the widened line cost the person working against it?
    A few extra turns. They stop asking for the text and start asking about it or asking the assistant to apply it, which needs no new channel and nothing planted anywhere. Because the mechanism is the same for every tenant of the same widget, the phrasing that lands carries over to the next deployment.
  • Why does a more detailed guard produce more informative refusals?
    Because a decline shaped by a specific named operation tends to be specific in return. Comparing which requests are refused and which are served maps the rule's edges, so each operation added to the list is also a signal about what the rule covers.
  • Is the widened line therefore a mistake?
    No, it is a trade. It reliably ends casual probing and costs a running bill of refusals against the product's real work. The mistake is recording it as closure, because the obvious phrasings failing is not evidence that the class is covered.

Listing the doors you may not use is not the same as owning the building.

saying these in an interview costs you the question

  • Calls the widened list a fix for the class
  • Calls it worthless because it is not a boundary
  • Ignores refusals on the product's legitimate work
  • Assumes the operations can be fully enumerated
  • Treats silence after the patch as evidence of closure

context