skip to content

A leaked index stores 800-token chunks, another 32-token — how do you grade the two disclosures?

level: seniorimportance: should knowfreq 33%

answer

  1. fixed width, variable amount of text
  2. two effects pull opposite ways
  3. a score is not a disclosure
  4. count what one person's chunks reveal
  5. ask what the table's columns omit

basics

~20 s

Per-vector fidelity falls as chunks lengthen, so short chunks come back near-verbatim and long ones as gist. But long chunks carry more content each, so grade by what a reader learns about a person, not by a reconstruction score.

solid answer

~50 s

Two effects pull opposite ways. A fixed-width vector holding 800 tokens has to compress far harder than one holding 32, so per-vector reconstruction fidelity is much worse for the long chunks — topic and a few entities rather than wording. But each long chunk is 25 times as much source text, so even a gist-level reconstruction can reveal who the matter concerns and what it is about, and the short-chunk index needs many recovered vectors to say the same thing. Grade by the disclosure, not the score: for the short-chunk index, near-verbatim lines of intake fields; for the long-chunk index, an accurate summary of each case note. Then ask the questions the fidelity number does not answer — which encoder, still callable, what query budget was assumed, and how many chunks concern any one person.

code

text · 10 lines
text
chunk tokens | vectors | recon. score | reviewer note
-------------|---------|--------------|------------------------------------
          16 |  50,000 |         0.93 | field values recovered, names intact
          32 |  50,000 |         0.84 | close paraphrase
         128 |  50,000 |         0.57 | gist plus most named entities
         512 |  50,000 |         0.34 | topic only
         800 |  50,000 |         0.29 | topic only
...
(assessment does not state: which encoder, whether it is still callable,
 queries spent per vector, or how many chunks belong to one client)

go deeper

for a junior

Know that a fixed-size vector holding more text keeps less of the wording, so short chunks reconstruct more faithfully than long ones. Know that low fidelity still is not the same as safe.

for a middle

Explain both effects: falling per-vector fidelity against rising content per vector. Be able to say what a reported reconstruction score omits — the encoder, its reachability, and the query budget behind the number.

for a senior

Show incident judgment: grade by what a reader learns about a named individual, insist the number is a lower bound, and know that re-embedding under a new encoder does nothing for vectors already copied.

for a principal

Own the framing for leadership: severity is per person and per corpus, not per vector, and encoder reachability is a decision made at design time that you cannot revoke during an incident.

## Why chunk length is the dial An encoder writes every chunk into the same fixed number of coordinates. A 32-token chunk and an 800-token chunk both come out as, say, one 1,024-dimensional vector. The long chunk therefore had to be compressed roughly twenty-five times harder, and the information that survives is the information the encoder considers most distinguishing — subject matter, a handful of salient entities — rather than the ordering and wording of the sentences. So per-vector fidelity really does fall as chunks lengthen, and a table of reconstruction scores by chunk length will show it clearly. That is the true half of the intuition. The mistake is going from there to "our chunks are long, so the leaked index is low impact." ## The second effect, which points the other way Each long chunk is *more content*. A gist-level reconstruction of an 800-token case note can still produce: the client's name, the counterparty, dates, the type of matter, and the substance of the advice. That is a disclosure by any regulator's definition. Meanwhile the 32-token index needs the attacker to recover and reassemble many neighbouring vectors to reach the same picture — which they can do, because the vectors are adjacent in the store and often adjacent in the space, but it is more work per unit of harm. Grading the two leaks therefore looks like this: | | 32-token chunks | 800-token chunks | |---|---|---| | Per-vector fidelity | high, near-verbatim | low, gist | | Content per vector | one field or line | a whole note | | Attacker work per person | recover and stitch many | recover a few | | Typical realised harm | exact field values | an accurate summary of the matter | Neither column is the safe one. They are different disclosures of the same corpus. ## What a fidelity table does not tell you When an exposure assessment arrives as a table of reconstruction scores against chunk length, the columns that are missing usually matter more than the ones present: 1. **Which encoder produced these vectors, and can an outsider still call it?** If the encoder was released openly or is hosted for anyone to query, the attacker's precondition is satisfied for free and every row of the table is optimistic. If it was a private fine-tune that no longer runs anywhere, the whole attack has a missing ingredient — which is a genuine mitigating fact and one of the few available after the vectors are gone. 2. **What query budget was assumed per vector?** A score produced by a cheap attempt is a lower bound on recoverability, not an upper one. A funded adversary buys more proposals and more refinement rounds and gets more back from the identical stored data. 3. **How many chunks concern one individual?** Severity is per person, not per vector. Fifty gist-level chunks about one client compose into something far more damaging than the score of any one of them suggests. 4. **Which content was embedded at all?** An index built over public policy documents and one built over intake notes score identically and are not remotely the same incident. ## The direction of every claim here A low reconstruction score bounds the attack that was run, not the vectors. This is the same discipline you apply to a robust-accuracy figure or a clean scan result: the measurement tells you what your attempt achieved, and someone with better proposals, more queries, or knowledge of the domain will do better against the same artefact. Write the finding as "our attempt at this budget recovered X", never as "only X is recoverable". ## What you actually do with the grading The practical output is not a severity number, it is a set of decisions: whether the encoder can be taken out of public reach (rarely, if it is a hosted or released one — so usually not), whether the leaked snapshot must be treated as a document-corpus breach for notification purposes (usually yes), whether affected individuals can be enumerated from the index's own metadata, and whether re-embedding the corpus under a private encoder changes anything for the *copies already taken* (it does not — the stolen vectors were made by the old encoder, and it is the old encoder's availability that governs them). That last point is the one people get wrong in an incident: rotating the encoder protects future vectors and does nothing for the ones already out of the building.

  • Does re-embedding the corpus under a new private encoder fix a leak that already happened?
    No. The copied vectors were produced by the old encoder, so their recoverability depends on that encoder's continued availability, not on what you use going forward. Rotation protects new writes. If the old encoder is openly released or hosted, you cannot withdraw it, which is why encoder choice is a durable exposure decision made long before the incident.
  • An assessment reports a 0.29 reconstruction score on the long-chunk index. What do you ask next?
    Which encoder, whether it is still callable by an outsider, how many queries per vector the attempt spent, and what the recovered text actually said about a named person. A score answers none of those. Ask for a handful of reconstructed chunks and read them — severity is decided by what a reader learns, not by a similarity metric against the source wording.
  • Would storing shorter chunks or longer chunks be the safer choice going in?
    Neither, and the trade should not be argued on that axis — chunk sizing is a retrieval-quality decision. Short chunks leak exact field values, long chunks leak accurate summaries, and the corpus is equally exposed either way. The control that changes the outcome is who can read the index and who can call the encoder, not how the text was sliced.

saying these in an interview costs you the question

  • Long chunks are safe because the reconstruction score is low
  • Quoting a fidelity score with no encoder or query budget stated
  • Grading severity per vector instead of per person
  • Treating a weak in-house attempt as an upper bound
  • Believing rotating the encoder protects vectors already copied
  • Arguing chunk size as a privacy control rather than a retrieval choice

context