skip to content

Text From an Embedding

An embedding is a lossy but largely invertible function of its input, so an index is a copy of the corpus under a weak transform. Interviewers probe whether you treat it as derived non-data.

on this pageshow

explore

questions

4

An attacker copies a vector index of embedded case notes, no source text, and can query the same encoder — why is that a disclosure?

level: juniorimportance: must knowfreq 55%

answer

  1. lossy is not the same as one-way
  2. the transform exists to preserve meaning
  3. who else can call that encoder?
  4. guesses can be scored by re-embedding
  5. approximate text is still a disclosure

basics

~20 s

An embedding is a lossy but largely invertible encoding of its input. With the same encoder — usually public or purchasable — much of the original wording can be reconstructed, so the index carries the documents' sensitivity, not less.

solid answer

~50 s

Treat a vector index as a copy of the corpus under a weak transform, not as a pile of anonymous numbers. The encoder's entire job is to preserve meaning, so each vector retains most of what distinguishes its chunk from every other chunk. An attacker who holds the vectors and can call the same encoder can build a stand-in that maps a vector back toward text, and can check a candidate by re-embedding it and seeing how close it lands — that verification loop is what makes the search converge. What comes back is approximate: names, dates, the gist of a clause at paraphrase fidelity rather than a verbatim span. That is usually more than enough for the disclosure to matter. The dependency is the encoder, and a hosted or openly released encoder makes that a purchase rather than a breach. So the index inherits the classification of whatever was embedded.

go deeper

for a junior

Be ready to say plainly that an embedding is not anonymisation: it is a meaning-preserving function, so with the same encoder much of the text comes back. Name the one thing the attacker needs — access to that encoder.

for a middle

Explain why discriminative retrieval and recoverability are the same property, and why an attacker who can re-embed a guess can score it. Be precise that the output is approximate text, not the source bytes and not the encoder's weights.

for a senior

An interviewer expects you to convert this into storage policy: the index and its backups take the corpus's classification, tenancy isolation, retention and deletion obligations, and its breach analysis. Say what you would ask first when a snapshot goes missing.

for a principal

Own the framing that derived data is only protected when the derivation destroys information. Be ready to defend classifying a vector store at corpus sensitivity against cost and convenience arguments, and to say what evidence would change your mind.

## The wrong mental model The common position is: "the vector store holds embeddings, not documents, so it does not need the document store's controls." It sounds reasonable — a stored row is a list of a few hundred or a few thousand floats with no readable text in it — and it is the reason vector indexes and their nightly backups routinely end up in looser storage, wider replication, and lower breach-severity tiers than the corpus they were built from. It is wrong for a simple reason: an embedding is a **function of the text, chosen specifically to keep the text's meaning**. Nothing about it was designed to destroy information. It is lossy — you do not get the bytes back — but it is nothing like a hash or a redaction, and it is not anonymisation. ## Why a vector is recoverable at all A text encoder maps a chunk to a point such that chunks meaning similar things land near each other and chunks meaning different things land apart. For that to be useful for retrieval, the map must be highly discriminative: two case notes that differ only in a name, a date or a clause must land in *different* places, or retrieval would not work. That discriminativeness is exactly the property an attacker exploits. If distinct texts reliably produce distinct vectors, then a vector narrows the set of texts that could have produced it — often very sharply. An adversary holding the vectors and able to call the same encoder is in a strong position because they can **verify**. They do not have to invert anything analytically. They can propose a candidate text, embed it, and measure how far the result sits from the stolen vector. A search that can score its own guesses converges. Published work on this class of attack trains a model to propose good candidates and then refines them against that closeness signal; the point for an interview is the property, not the procedure — an invertible-enough function plus an oracle that scores guesses is a recoverable function. ## What the attacker actually gets Be precise about the payoff, because overclaiming is as bad as underclaiming: - **Approximate text, not the source bytes.** Short chunks come back close to the original; longer ones come back as gist — the topic, some entities, the shape of the claim. - **Entities survive well.** Names, dates, amounts and rare terms are precisely the tokens that move a vector the most, so they are the tokens most likely to be recovered. That is the worst possible failure mode for case notes, patient records or intake forms. - **Not the encoder's parameters.** This is a privacy attack on the *data*, not model extraction. The attacker ends with text, not weights. - **Not a claim about the model's training set.** Recovering the text behind a stored vector says nothing about what the encoder was trained on; that is a different attack with a different answer. ## The limit, stated honestly The attack has a real precondition: **the same encoder, or query access to it.** The mapping from vectors back to text is specific to the encoder that produced them; a stand-in built against one encoder does not transfer to a different one. So the questions that decide severity are: 1. Which encoder produced these vectors, and can an outsider call it? An openly released or commercially hosted encoder means yes, cheaply. A privately fine-tuned one means the attacker needs an endpoint that will embed text for them — which many products expose, deliberately, as part of ingest or search. 2. How long is each chunk? Recovery degrades as chunks lengthen, because a fixed-width vector has to carry more text. 3. How many encoder queries will they pay for? The verification loop is metered if the encoder is a paid endpoint. That is a cost control, not a boundary. Note what is *not* on that list: the similarity metric, the index structure, whether vectors were normalised, or how many vectors there are. Those are retrieval engineering, and none of them remove information about the source text. ## The consequence for how you treat the store If the index is a weakly transformed copy of the corpus, then the corpus's rules follow it: the same access control, the same tenancy isolation, the same retention and deletion obligations, the same encryption at rest, the same classification label, and the same breach-notification analysis when a backup goes missing. A neighbouring tenant who can read vectors they should not see has read documents they should not see. An operator with backup access has document access. An inherited snapshot with no source manifest is still regulated content. The honest summary for a review: *"the index is derived data, but it is not de-identified data."* Derivation is not protection unless the derivation destroys information, and this one is built not to.

  • Does keeping the encoder private stop this?
    It raises the bill; it is not a boundary. The attacker needs query access, not the weights, and most products expose an embed path during ingest or search. A privately fine-tuned encoder does narrow the field — a stand-in built for a different encoder will not transfer — so treat encoder secrecy as a cost control layered under access control on the index, never as the reason the index is safe.
  • Is this the same as extracting the encoder itself?
    No. The payoff here is the text behind stored vectors; the encoder's parameters stay where they are. Model extraction is a different attack with a different limit — a query budget spent to obtain a functional stand-in — and it does not require possession of anyone's index. Conflating them leads teams to defend the endpoint when the exposure is the storage.
  • Does normalising or quantising the stored vectors help?
    Barely. Normalising discards only the vector's length, which carried little of the content, and quantisation trades precision for space — it coarsens the reconstruction rather than preventing it. Both are storage and retrieval decisions; neither converts the index into de-identified data, and neither should change how you classify a backup of it.

A photocopy through frosted glass is still the document. You cannot read the individual letters, but with the same glass you can work out what was underneath — and the names come through first.

saying these in an interview costs you the question

  • Embeddings are anonymised, so the index needs weaker controls
  • You cannot get text back from a list of floats
  • The encoder is private, therefore nobody can invert it
  • An index leak is automatically lower severity than a document leak
  • Confusing recovering the text with stealing the encoder's weights
  • Assuming only exact, verbatim recovery would count as disclosure

context

open as a page

What does an attacker need besides a stolen embedding to recover its text, and how faithful is the result?

level: middleimportance: should knowfreq 42%

basics

~20 s

They need query access to the same encoder plus a budget to spend on it. What comes back is a paraphrase: close to the original for short chunks, gist for long ones. The corpus and the encoder's weights are not required.

open as a page

A leaked index stores 800-token chunks, another 32-token — how do you grade the two disclosures?

level: seniorimportance: should knowfreq 33%

basics

~20 s

Per-vector fidelity falls as chunks lengthen, so short chunks come back near-verbatim and long ones as gist. But long chunks carry more content each, so grade by what a reader learns about a person, not by a reconstruction score.

open as a page

You inherit a vector index from an acquired product with no source documents — what can you honestly tell counsel is recoverable?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Not "nothing". The index is a recoverable copy of whatever was embedded, and you cannot bound its content without inverting a sample. Offer a choice: fund that assessment, or classify at corpus sensitivity and destroy the backups.

open as a page