skip to content

Reconstruction from Fragments

Hidden text longer than one answer comes back in pieces, and an attacker who samples repeatedly can tell the stable recalled parts from confabulation. Interviewers use it to price a 'small' leak.

on this pageshow

explore

questions

4

A ticket-triage auto-reply is length-capped per reply - why does that not bound the leak?

level: juniorimportance: must knowfreq 58%

answer

  1. one reply is not one session
  2. the cap bounds a fragment
  3. nothing in the pipeline counts across runs
  4. the reset clears the server, not the notes
  5. cost is request count, not cleverness

basics

~20 s

A per-reply cap bounds one response, not the total. Each submission is a fresh run over the same hidden context, so many submissions yield many capped slices, reassembled outside the system. Exposure counts per campaign, not per reply.

solid answer

~40 s

The cap is a property of a single response, and the screening layer that reads each reply is also a property of a single response. Neither counts anything across runs. In an unattended triage workflow every ticket is its own stateless run, and the same worked examples get pasted into the prompt each time, so the hidden text stays put while the requests keep coming. Somebody submitting four hundred tickets gets four hundred short, individually unremarkable replies, each carrying a different slice of that text, and stitches them together outside the system. The interesting number is not the length of one reply but the amount of hidden context and the number of requests needed to walk it. Cost here is request volume, not cleverness.

go deeper

for a junior

Be ready to say plainly that a per-reply cap limits one reply and nothing else, and that many cheap submissions against the same hidden context add up. Recall that each run is fresh for the application but not for the person collecting the replies.

for a middle

Explain why both the cap and the per-response screen are stateless, and what that means for a component that only ever sees one candidate reply. An interviewer expects you to name the request count as the real cost.

for a senior

Demonstrate that you grade severity on the campaign rather than the response. Be ready to say what the collected set proves and what it does not, and to name the deployment property - the same examples pasted every run - that the whole construction rests on.

for a principal

Own the framing that per-response measurements are the wrong unit for this class of exposure, and be able to argue that position to an owner whose evidence is a log full of individually unremarkable replies.

## The setting Picture an unattended ticket-triage and auto-reply workflow: a public intake form, a model that classifies each submission and writes a short answer back to the sender, and no human anywhere in the loop. Two details matter. First, the reply is template-capped - a short paragraph, no more. Second, the prompt built for each run pastes in two or three previously resolved tickets as worked examples, and those examples are real: another customer's name, their order number, the internal resolution note somebody typed. So hidden text that belongs to a third party is sitting in the context of every single run, and each run answers whatever arrived in the form body. ## What the cap actually bounds A per-reply length cap bounds one reply. That is the whole of its scope. It says nothing about how many replies exist, whether they differ, or whether their contents can be laid side by side afterwards. The same is true of a per-response screening layer. A screen reads one candidate reply and emits a pass or a block. It holds no memory of the previous four hundred replies and performs no comparison between them. A slice of somebody's resolution note, arriving alone, in a plausible support-reply register, is not obviously a disclosure to a component that only ever sees it alone. Both obstacles are real, and both are stateless. That is the whole shape of this construction: it is aimed at exactly the gap between a per-response control and a multi-request campaign. ## The session boundary cuts the other way from how it reads A fresh session per ticket sounds protective. It does clear the server's context. It does not clear the attacker's notes, and the reconstruction is not happening inside a conversation - it is happening in a text file on somebody else's machine, where nothing resets. What the reset genuinely costs is continuity. There is no way to say keep going from where you stopped, because nothing remembers where anything stopped. Each run must be nudged toward an overlapping region of the same hidden text, so the collected slices are redundant and out of order. That makes the campaign longer and noisier. It does not make it smaller. ## What is being spent Request count. The economics are linear and dull: more hidden text means more requests, more redundancy means more requests, and requests are cheap, automatable and parallel. There is no escalation in sophistication once the first slice comes back - the same undistinguished request shape, repeated. An interviewer asking about this is usually checking whether a candidate prices an attack in ingenuity when it should be priced in volume. ## Getting the direction of each claim right Four statements that look equivalent and are not: | observation | what it proves | what it does not prove | | --- | --- | --- | | the reply was short | the response was short | the exposure was small | | the screen never fired | each reply scored below a threshold alone | the set of replies was harmless | | the session reset | the server kept nothing | the requester kept nothing | | one ticket looked benign | that ticket looked benign | four hundred of them did | The unit of exposure is the campaign. Grading the leak by the size of one response is the single most common wrong answer here, and it is wrong in a way that changes severity by orders of magnitude. ## Where the construction runs out on its own terms It is not universal, and saying where it fails is part of a good answer. It depends on the hidden text being the same in every run. If the pasted examples are drawn fresh from a large pool each time, the slices belong to different underlying texts and will not align into one passage - the collector ends up with confetti rather than a document. It also depends on there being enough hidden text to be worth walking; a two-line preamble is not a campaign target. And it is bounded above by confabulation: past the point where the model is actually reading something, it starts producing plausible text that no amount of collection will turn into evidence. ## What an interviewer is listening for That the candidate reframes the question from how much leaked in that reply to how much is reachable across runs, names the per-response cap and per-response screen as the specific things the construction is aimed at, and prices the attack in requests.

  • An output screen never fired across four hundred captured replies. What does that tell you?
    That every reply, judged on its own, scored below the screen's threshold. A per-response screen holds no state across runs and never compares one reply with another, so it cannot see a set. Passing means scored low in isolation, not harmless - and a slice of somebody's resolution note in ordinary support-reply register is exactly the shape that scores low.
  • What does the per-ticket session reset actually cost the person collecting fragments?
    Continuity. Nothing remembers where the last slice ended, so each run has to be steered back onto an overlapping region of the same text, and the collected material arrives redundant and unordered. That inflates the request count and the reassembly effort. It does not reduce the total recoverable, because the reassembly happens outside the system where nothing resets.
  • Why is another customer's data sitting in this prompt at all, from the attacker's point of view?
    From the attacker's side it does not matter why - what matters is that it is present in every run and stable across runs. Passively pasted context is a better target than the application's own instructions here: it is third-party data, it is bulkier, and it is often unnoticed because nobody thinks of a worked example as something the model can be walked through.

A photocopier that will only ever copy a single page is not a document control. Somebody feeding it three hundred pages leaves with the book.

saying these in an interview costs you the question

  • Grades the leak by the length of one reply
  • Assumes a session reset erases what the requester kept
  • Treats a screen that never fired as proof nothing leaked
  • Counts exposure per response instead of per campaign
  • Thinks a length cap also limits request volume

context

open as a page

In a stateless auto-reply workflow, how do capped fragments become one hidden passage?

level: middleimportance: should knowfreq 44%

basics

~20 s

Different requests return differently positioned slices of the same hidden text. Because that text is identical in every run, overlapping regions let the slices be ordered and merged outside the system. Alignment is approximate, since models paraphrase.

open as a page

Only part of a reconstructed customer record reproduces reliably - what do you claim in the finding?

level: principalimportance: should knowfreq 31%

basics

~20 s

Claim the mechanism, not the transcript. Sign off that hidden third-party context is recoverable in slices across independent runs, evidenced by the spans confirmed against a source. Leave stable-but-unconfirmed text out of the claim and out of circulation.

open as a page

In a reconstruction built from repeated samples, does span agreement prove recall?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

No. Disagreement between independent samples is strong evidence a span was generated rather than read. Agreement only proves the output is stable, which both genuine recall and a strongly format-shaped guess produce. Agreement is necessary for recall, never sufficient.

open as a page