skip to content

A tenant reports a user obtained a faithful summary, not a quotation, of its hidden preamble. Is that a leak?

level: seniorimportance: nice to knowfreq 30%

answer

  1. grade the content, not the form
  2. quotation adds wording, not meaning
  3. the vendor holds the ground truth
  4. split the report into two questions
  5. say how often it reproduced

basics

~20 s

Yes, when the value is in the content. A faithful summary carries the thresholds, rules and escalation wording that made the preamble sensitive; only exact phrasing is lost. Grade it by what was recovered, not by its form.

solid answer

~50 s

Triage on the content recovered, not on the form it arrived in. If the preamble held pricing rules, escalation thresholds or competitive positioning, an accurate summary hands over everything actionable and the missing quotation marks change nothing. The exception is a preamble whose exact wording is the asset, such as contract or regulated language, where a paraphrase is materially weaker evidence and a materially weaker outcome. As the vendor you also hold the ground truth, since the tenant supplied the preamble, so you can check fidelity directly instead of guessing. Then separate the two halves of the tenant's report: that a request of this shape was served is a property of putting readable configuration in front of a model, while whether that particular content belongs in a preamble is the tenant's own data-handling decision, which no rewording of the guard line settles.

go deeper

for a junior

Know that a summary of hidden configuration still tells the reader what it said. The absence of quotation marks is not what decides whether information left.

for a middle

Be able to name the case where wording genuinely matters, such as contractual or licensed language, and explain why elsewhere the meaning carries all the sensitivity that made the text worth protecting.

for a senior

Show a triage method: compare against the original you hold, grade on matched content, record how often it reproduced, and separate what the widget did from what the tenant chose to put in its configuration.

for a principal

Own the message to the tenant. Be able to say what configuration text can and cannot be relied on to hold without either overstating a defect or implying that a rewrite of the guard line resolves the underlying arrangement.

## The report as it arrives A tenant embedding a writing-helper widget in its product files a report: a user asked the assistant for a short summary of its own configuration and got one. The summary is accurate. It never quotes. The tenant wants to know whether this is a security bug in the widget. This is a triage call, and the two ways of getting it wrong are symmetrical. Grading it as *not a disclosure because no protected string appeared* is the more common error. Grading it as a critical break of the widget when nothing about the deployment behaved unexpectedly is the other. ## Why form is the wrong axis The reason a preamble is sensitive is almost never its prose. It is sensitive because of what it says: the discount ceiling the assistant must not exceed, the conditions under which a conversation escalates, the claims about a competitor the brand will not make, the internal names of processes it references. Every one of those survives a summary intact. A reader who wanted to act on the content is in the same position with a paraphrase as with a transcript. So the question to ask is not *did the text come out*, it is *what did the reader end up knowing*. Sensitivity travels with the semantics. ## When wording genuinely matters There are real cases where a quotation is the stronger outcome and a paraphrase is weaker: - The wording itself is the asset, for example licensed or contractual language, or a form of words a legal team signed off. - The value lies in reproducing the configuration exactly, for instance to build something that behaves identically. - The finding needs to survive a dispute about fidelity, where an approximate restatement invites the argument that the model invented it. Outside those, treating the paraphrase as a lesser event is a form-over-substance mistake. ## Fidelity is checkable here, which is the vendor's advantage A triager on the vendor side is not guessing about accuracy. The tenant supplied the preamble, so the ground truth is on file. Compare the summary against it and grade on how much of the sensitive content matched. Someone with no access to the original would have a much harder problem, which is precisely why external reports of this kind arrive with weaker evidence than internal triage of the same event. The severity ought to move with what matched: a summary that reproduces the escalation threshold accurately is a different finding from one that gestures at the preamble's general shape. ## Bug or design limit, which is the actual question Split the tenant's report in two. The first half is *the widget served a request for an operation over its own configuration*. Nothing malfunctioned. Configuration has to be readable by the model it configures, and a widget whose purpose is transforming text will transform text that is in front of it. A guard sentence changes the odds and not the arrangement. Reporting that back honestly is the job: this is what a preamble can be relied on to do, which is shape behaviour, and it is not a container. The second half is *this particular content was in the preamble*. What a tenant chooses to place there, and how it treats content whose recovery would matter, is the tenant's own decision about its data. A triager should not answer that question inside a report about the widget, and should be careful not to imply that a wording change to the guard line settles it. ## Reproduction and honesty about it Before assigning severity, establish how the result behaves on repetition. A construction of this kind can succeed on one attempt and fail on the next several, because the outcome is sampled model behaviour. A finding that reproduced once is genuinely a finding, but the report should say how many attempts it took, since *a user can obtain this* and *a user obtained this once* support very different conclusions about exposure. Writing the rate down is what separates a usable report from an anecdote. ## What the tenant is entitled to hear A clear statement of what the preamble holds and what it does not: it holds behaviour reliably enough to be worth writing, it does not hold content away from anyone who can address the model, and the difference between quotation and restatement is a difference in form rather than in what the reader now knows.

  • How do you check fidelity when the output never quotes the original?
    On the vendor side you hold the tenant's preamble, so you compare the summary against it clause by clause and grade on how much of the sensitive content matched. That is a direct comparison, which is why internal triage of this event is far better evidenced than an external report of the same thing.
  • What do you tell the tenant a preamble can be relied on to hold?
    Behaviour. It shapes tone, refusals and the rules the assistant follows, and it does that well enough to be the normal way to configure the widget. It does not hold content away from a user who can address the model, because the text has to be readable by the model to have any effect at all.
  • Does the finding change if it only reproduced once in five attempts?
    It changes the severity, not the classification. One success shows the construction worked against this deployment, and the rate belongs in the report because exposure at one in five and exposure at five in five justify different responses. Reporting a single success as if it were reliable is the error to avoid.

saying these in an interview costs you the question

  • Says no disclosure occurred because nothing was quoted
  • Grades severity by output form rather than recovered content
  • Guesses at fidelity while holding the original preamble
  • Reports one lucky success as a reliable result
  • Answers the tenant with a rewording of the guard line

context