Only part of a reconstructed customer record reproduces reliably - what do you claim in the finding?
answer
- two claims, only one supported
- mechanism stands on the confirmed subset
- over-claiming has asymmetric cost
- verification is an access decision
- the evidence is itself third-party data
basics
~20 sClaim the mechanism, not the transcript. Sign off that hidden third-party context is recoverable in slices across independent runs, evidenced by the spans confirmed against a source. Leave stable-but-unconfirmed text out of the claim and out of circulation.
solid answer
~50 sThere are two separable claims and only one of them is supported. The mechanism - that a per-response cap and a per-response screen do not bound what many independent runs return, and that pasted third-party context is reachable - is demonstrated by the confirmed subset and stands on its own. The reconstructed transcript is an inference with paraphrased joins and unverifiable spans, and asserting it as another customer's record risks triggering notification duties over text a model invented, which is worse than under-claiming. So I would file on the mechanism, quantify the evidence honestly - so many spans stable, so many confirmed - and treat the unconfirmed remainder as excluded rather than pending. I would also argue against attaching the full reconstruction to a widely readable ticket, since circulating a possibly-real customer record inside the tracker widens the exposure the finding is about.
code
text · 14 linesfiled as: LLM07 System Prompt Leakage
(reporter's wording: the assistant disclosed hidden context)
refiled as: LLM02 Sensitive Information Disclosure
(recovered text is another customer's ticket data, pasted into
the prompt as a worked example - not the app's own instructions)
evidence: 31 spans stable across 5 independent runs
12 of 31 confirmed against the source record -> claimed
19 of 31 stable but unconfirmed -> not claimed
... remaining spans varied between runs -> struck
claim: third-party context is recoverable in slices across runs;
the reconstructed transcript itself is NOT asserted as the recordgo deeper
Recall that a recovered passage is not automatically true, and that a finding should say which parts were actually checked. You are not expected to weigh notification duties at this level.
Be able to separate the demonstrated mechanism from the assembled transcript and to explain why only the confirmed spans support a claim about a real person's data.
Show that you quantify evidence in the report, exclude rather than defer the unconfirmed remainder, and think about where a possibly-real record is allowed to be stored and circulated.
Own the sign-off: which claim the organisation makes, the asymmetric cost of over-claiming, whether verification against a live record is authorised and by whom, and how severity is argued when every individual response looks unremarkable.
## The position you are in Somebody filed a finding against an unattended ticket-triage workflow. The attachment is four hundred captured auto-replies and a reconstructed passage said to be another customer's ticket: a name, an order number, two dates, an internal resolution note. Some spans reproduced identically across independent runs; some varied; twelve were checked against a real record and matched. You have to say what the organisation is asserting. This is a judgment call, not a technical one, and it has four parts that people tend to collapse into one. ## Part one: separate the mechanism from the artefact The mechanism claim is that hidden third-party context pasted into every run is recoverable in slices across many independent, individually unremarkable requests, and that a per-reply length cap and a per-response screening layer bound the fragment rather than the total because neither is stateful across runs. Twelve confirmed spans establish that. The claim does not get stronger with a nicer transcript. The artefact claim is that the reconstructed passage is that customer's record. It is an inference: joins were made on approximate overlaps between paraphrased fragments, coverage gaps are invisible, and stable-but-unconfirmed spans are exactly the format-shaped ones a model converges on unaided. Signing that claim means asserting specific personal data about a named individual on the strength of a merge. Under-claiming here costs you almost nothing, because the mechanism is what any remediation decision turns on. Over-claiming costs a great deal. ## Part two: the cost of being wrong points one way If you assert that a named customer's record leaked and the unconfirmed half was confabulated, you may have started a notification process about invented data, alarmed a real person, and spent the credibility the security function needs for the next finding. If you assert only the confirmed subset and the full record did in fact leak, the mechanism claim already forces the same response. The asymmetry is stark and it should drive the wording. ## Part three: the verification you would need has its own price Confirming the remaining spans means pulling the real ticket and reading another customer's record - an access decision with its own authorisation question, made by someone in security to prove that someone else could read it. Some organisations will approve that under a defined process; some will not; either way it is not a step a triager takes unilaterally because it would tidy up a finding. Deciding whether to seek it, and being able to justify stopping without it, is part of the call. ## Part four: handling of the artefact itself A reconstruction that may contain real third-party data is itself third-party data. Attaching it in full to a ticket readable by an engineering org copies the exposure into a system with wider access than the workflow ever had. Referencing it, restricting it, or reducing it to the confirmed spans are all defensible; pasting the whole transcript into the tracker because it is the evidence is not, and it is a common reflex worth naming. ## Filing it under the right item This often lands under the system-prompt leakage item because the phrase hidden context is in the report. That is the wrong item. The recovered text is another party's data that entered the prompt passively as a worked example - it is not the application's own instructions - so it belongs under sensitive information disclosure. The distinction matters beyond bookkeeping: it changes whose data is at issue, which obligations attach, and which team recognises the finding as theirs. ## What reproducibility means for severity One successful campaign proves the construction worked against one deployment on the days it ran. If the pasted examples rotate, later campaigns may recover confetti instead of a passage, so the demonstrated result may not generalise across time - and a stable result may equally mean nobody has touched that prompt in a year. Say which of those you established and which you assumed. Severity should be argued on the reachability of third-party context across runs, not on the length of any individual reply, because that per-response framing is precisely the reasoning the finding refutes. ## What an interviewer is listening for A candidate who signs the mechanism, quantifies the evidence, refuses to promote stable into confirmed, treats verification as an access decision with a cost, and thinks about where the reconstructed text is allowed to live. A candidate who says file the whole transcript, it is what we recovered has not noticed that the deliverable is itself a copy of the thing that leaked.
- The reporter wants the full reconstructed transcript attached to the ticket. What is your answer?No, or restricted. A reconstruction that may contain real third-party data is third-party data, and a tracker readable across engineering has wider access than the workflow ever had - attaching it copies the exposure the finding is about. The confirmed spans, or a redacted reference, carry the argument; the full transcript adds nothing to the mechanism claim.
- An owner argues the finding is weak because no single reply exposed anything meaningful. How do you answer?That is the finding, not a rebuttal. The per-response framing is exactly the reasoning being refuted: the cap and the screen each judge one reply, and the exposure is a property of the set. Severity should be argued on how much third-party context is reachable across runs and how cheaply, which is request volume, not on the size of any one reply.
- Would you seek approval to read the real customer record to confirm the remaining spans?It is a decision with a cost, not a free step. Someone in security would be reading a customer's record to prove somebody else could read it, which needs authorisation and a defined process. I would ask, be prepared for a refusal, and make sure the finding stands without it - because a mechanism demonstrated by twelve confirmed spans does not need the other nineteen.
- What do you tell an owner whose proposed answer is to tighten the reply cap?That it changes the slice size and the request count, not the reachability - the cap was never the thing being defeated, it is the thing the construction is built around. That is an observation about what the fix addresses, and where the discussion moves is the owner's call, but the finding should record explicitly that a tighter cap does not close what was demonstrated.
saying these in an interview costs you the question
- Asserts the whole reconstruction as a customer's record
- Promotes stable-but-unconfirmed spans into confirmed
- Files third-party data exposure as system-prompt leakage
- Pastes the full reconstruction into a widely readable ticket
- Argues severity from the size of one reply
- Reads a real customer record unilaterally to confirm spans