You inherit a vector index from an acquired product with no source documents — what can you honestly tell counsel is recoverable?
answer
- refuse both of the easy answers
- one fact carries the whole assessment
- metadata may answer it without inverting
- a negative sample clears nothing
- present a funded choice, not a reassurance
basics
~20 sNot "nothing". The index is a recoverable copy of whatever was embedded, and you cannot bound its content without inverting a sample. Offer a choice: fund that assessment, or classify at corpus sensitivity and destroy the backups.
solid answer
~50 sRefuse the two easy answers. "It is only vectors, nothing is recoverable" is false, and "assume everything is recoverable verbatim" is unfundable. What you can establish cheaply: which encoder produced the vectors and whether an outsider can still call it, how many vectors and how long the chunks were, what metadata the rows carry, and where the backups live. What you cannot establish without work: what the text said — and the only honest way to find out is to invert a sample under controlled conditions, which is itself processing of the very data you are unsure about. So present a decision, not a reassurance: fund a scoped sample inversion and classify from the result, or skip it and treat the store at the acquired corpus's sensitivity with matching retention and deletion. Say plainly that a reassuring low score would be a lower bound on recoverability, never a clearance.
go deeper
Know the one-line position: a vector index is derived data, not anonymised data, so it keeps the classification of whatever was embedded until somebody proves otherwise.
Be ready to list what you would check first — which encoder, is it still callable, chunk length, row metadata, where the backups are — and to explain why the encoder question dominates the rest.
Show you can scope the assessment: small sample, named approver, short retention on the reconstructions, and the wording of the finding agreed before the result exists.
Own the call. Put two funded options to counsel rather than a reassurance, defend a claim that stays true if the store leaks later, and refuse to let a negative sample become a clearance.
## The chair you are sitting in An acquisition brought over a running product. The vector index came with it; the source documents did not, there is no chunk manifest, nobody who built it is still around, and there are nightly backups of unknown age. Counsel asks a simple question: **is there personal data in there, and can it be recovered?** They need an answer they can stand behind, not a research result. This is a principal-level question because the technical part is short and the judgment part is not. It is an assurance claim someone has to sign, a cost someone has to fund, and a disclosure position someone has to defend later. ## The two answers you must refuse **"It's just floats — nothing is recoverable."** This is the field's standard wrong answer and it will not survive contact with anyone who reads the literature. An embedding is a meaning-preserving function of its input; with access to the same encoder, much of the text can be reconstructed, and the tokens that survive best are names, dates and identifiers. Putting this in writing is the outcome you most want to avoid. **"Assume verbatim recovery of everything."** Equally unhelpful. It overstates the attack — reconstruction is approximate and degrades with chunk length — and it converts every inherited artefact into a maximum-severity finding, which is how assurance programmes lose credibility and budget. ## What you can establish cheaply, and what it buys Four facts are obtainable in a day and each moves the decision: 1. **Which encoder produced these vectors, and is it still callable?** The output dimension and any surviving configuration narrow it. This is the single most load-bearing fact, because recoverability depends on an attacker being able to embed candidate text with the *same* encoder. An openly released or commercially hosted encoder means the precondition is met for anyone, forever, and you cannot withdraw it. A private fine-tune whose weights are lost with the acquisition is a genuine and defensible mitigation — the one piece of real good news available. 2. **Chunk length and row count.** Governs whether reconstructions would be near-verbatim lines or gists, and how much text is in scope. 3. **The row metadata.** Source identifiers, tenant keys, timestamps and titles are often stored beside the vector in plain text. They frequently answer the personal-data question outright, without inverting anything, and they tell you whether affected individuals are even enumerable. 4. **Where the copies are.** Backups, replicas, analytics exports and a laptop somewhere. The population of copies decides what deletion can even mean. ## What you cannot establish without paying for it What the text actually said. The only way to know is to invert a sample and read it — and note the awkward property: that assessment is itself processing of data whose nature is unknown, run by people who then know what was in it. It needs a scope, an approver, a small sample, a short retention window for the reconstructions, and a written destruction step. Do not let it be run informally by a curious engineer. Design it to answer one question — *does recovered text from this index contain personal data* — not to characterise fidelity. A yes on a small sample settles the classification. A no on a small sample settles nothing, and you must say so before you run it, or the negative result will be quoted back at you as a clearance. ## The decision you actually put to counsel Frame it as two funded options and a claim you can defend: - **Option A — assess.** Scope a sample inversion, classify from what comes back, and accept that a weak result still leaves you at the acquired corpus's classification by default. Costs engineering time and creates a document that will be read in any future dispute. - **Option B — do not assess.** Classify the index and every backup at the sensitivity of whatever the acquired product handled, apply that retention and access control, and if the product is dead, destroy the store and the backups and record the destruction. Cheaper, and in many acquisitions the right call, because an index nobody can attribute to source documents is an asset with no owner and no upside. And the claim itself, worded so it stays true later: *"The index is derived data, not de-identified data. With access to the encoder that produced it, an attacker recovers approximate text including names and dates. We have not bounded what is recoverable; we have bounded what our own attempt at a stated budget recovered."* ## The line you never cross Do not tell anyone the vectors are anonymised. Do not let a low reconstruction score be recorded as a clearance. Do not claim deletion of the source documents removed the exposure — the index is a separate copy under a weak transform, and deleting the originals while retaining the index leaves the disclosure fully intact while creating a record that says the data was deleted. That last combination is the worst of the available outcomes, and it is the one that happens by default when nobody asks this question at all.
- The acquired product's source documents were already deleted. Does that help?No, and it can make things worse. The index is a separate copy under a weak transform, so the disclosure survives the deletion intact — while the deletion record now asserts the data is gone. If you inherit that combination, treat the index as the surviving copy of the corpus and resolve it explicitly rather than letting the paperwork stand unchallenged.
- Your sample inversion returns unreadable text. What may you write down?That your attempt, at the stated budget and against the encoder you had, recovered little. Not that the index is safe. Agree that wording before the assessment runs, because a negative result is the one most likely to be quoted out of context later, and a funded adversary with better proposals gets more from the identical stored vectors.
- Engineering wants to keep the index to warm-start the replacement product. What do you say?Then it is retained regulated content, not a technical artefact: it takes the acquired corpus's classification, access control, retention clock and deletion obligations, and it must be enumerable to a data subject. If nobody will fund that, the warm start is not free and the honest answer is to re-embed from documents you can actually account for.
saying these in an interview costs you the question
- Telling counsel the vectors are anonymised or de-identified
- Recording a weak reconstruction result as a clearance
- Claiming deleting source documents removed the exposure
- Running an unscoped inversion assessment informally
- Refusing to give any answer until a full study is funded
- Ignoring plaintext row metadata that already answers the question