skip to content

What does an attacker need besides a stolen embedding to recover its text, and how faithful is the result?

level: middleimportance: should knowfreq 42%

answer

  1. one precondition does all the work
  2. you must be able to score a guess
  3. same encoder, or nothing
  4. queries are metered, that is the dial
  5. paraphrase, not the source bytes

basics

~20 s

They need query access to the same encoder plus a budget to spend on it. What comes back is a paraphrase: close to the original for short chunks, gist for long ones. The corpus and the encoder's weights are not required.

solid answer

~50 s

Three things: the vectors, an oracle that will embed text with the *same* encoder, and enough queries to pay for the search. The oracle is the load-bearing item, because it lets the attacker score a guess — embed a candidate, measure how close it lands to the stolen vector, keep what is closer. A stand-in model trained against that same encoder proposes better starting candidates and cuts the query bill sharply, but it is built per encoder and does not transfer to a different one. What they do not need is the target's corpus, its documents, or the encoder's parameters. The payoff is approximate text at paraphrase fidelity: short chunks of a few dozen tokens come back close to the original in reported results, while passage-length chunks come back as topic, some entities, and the shape of the content rather than the wording.

go deeper

for a junior

Know the shortlist: the vectors, access to the same encoder, and queries to spend. Know that the result is approximate text rather than the original file, and that the encoder's weights are not part of the prize.

for a middle

Explain why an oracle that scores guesses is what makes the search work, and why a stand-in is a per-encoder one-time cost that then amortises across the whole index. Be exact about fidelity and about what is not required.

for a senior

Show you can triage with this: which encoder produced these vectors, is it still callable, how long were the chunks, and how would you bound what a funded attacker recovers rather than what your own quick attempt recovered.

for a principal

Own the economics. Argue about where the attacker's cost actually sits, what encoder choice and endpoint exposure commit you to for the life of the index, and what evidence you would accept before calling a leaked index low impact.

## The three inputs, and which one matters An adversary who has copied a vector index — a neighbouring tenant, an operator, someone holding a nightly backup — needs three things to turn a stored vector back into readable text. **1. The vectors.** They have these; that is the premise. **2. An oracle for the same encoder.** This is the real precondition and the whole threat model. The attacker must be able to submit text and get back a vector from *the encoder that produced the stolen ones*. That can be a released weight file they run themselves, a commercially hosted encoder anyone can call, or the victim product's own ingest or search path if it will embed attacker-supplied text. It cannot be a different encoder: vector spaces are not interchangeable, and a candidate scored against the wrong encoder tells you nothing about the right one. **3. A query budget.** If the oracle is metered, every candidate scored costs money. This is the quantity that decides whether the attack is a curiosity or a Tuesday, and it is the only dial a defender controls once the vectors are gone. ## Why the oracle is what makes it work The attacker is not solving an equation. They are running a **search that can grade itself**. Propose text, embed it, measure the distance to the stolen vector, prefer the candidate that lands closer. Any function you can query is a function whose inputs you can search for, and the more discriminative the encoder — the property retrieval demands — the sharper the grading signal. A trained stand-in changes the economics rather than the possibility. Instead of starting from nothing, the attacker uses a model that has learned, from many (text, vector) pairs bought from the oracle, to propose a plausible text for a given vector. Good proposals mean far fewer scored candidates per stolen vector. Two consequences follow, and both are worth saying out loud in an interview: - The stand-in is **per encoder**. Building it is a one-time cost against a given encoder; once built it amortises across every vector in the index, so a large stolen index is not proportionally more expensive to mine. - A widely released or commercially hosted encoder means somebody may already have paid that cost. Encoder popularity is an attacker convenience. ## What they do not need Stating the non-requirements is how you show you have the threat model straight: - **Not the corpus.** No sample of the victim's documents is required. This is not membership inference and there is no reference set to assemble. - **Not the encoder's weights.** Query access is sufficient; the parameters stay where they are. If the goal were the parameters, that would be extraction — a different attack with a different payoff. - **Not the retrieval stack.** Index structure, filtering and ranking configuration are irrelevant once the raw vectors are in hand. - **Not write access anywhere.** Nothing is poisoned and nothing is retrained. This is a read-only privacy attack on stored data. ## Fidelity: say what actually comes back The honest description is **approximate text**, and the fidelity is governed by how much text each vector had to carry: - Very short chunks — a query, a title, a line of a form, a few dozen tokens — come back close to the original in reported results, sometimes near-verbatim. - Passage-length chunks come back as gist: the subject matter, several entities, the general claim, but not the wording. - **Rare, distinctive tokens survive best.** Names, dates, amounts, identifiers and unusual terms are the tokens that move a vector the furthest from its neighbours, so they are the most recoverable. For case notes or intake records, that is precisely the content that makes the disclosure serious. Do not claim verbatim reconstruction, and do not dismiss the result because it is not verbatim. A reconstruction that gets a client's name, a date and the nature of a matter is a disclosure whether or not a single sentence matches. ## The direction of the claim One subtlety separates a careful answer from a sloppy one. If someone runs an inversion attempt against their own index and gets poor reconstructions, that bounds **the attempt that was made** — that stand-in, that query budget, that number of refinement rounds — not what is recoverable from those vectors. Better proposals and a larger budget recover more from the same stored data. A weak result is evidence about the attacker's effort, never a guarantee about the artefact.

  • Why can't the attacker use a different encoder they already have a stand-in for?
    Because vector spaces are not shared. Coordinates from one encoder carry no fixed meaning in another, so a candidate scored against the wrong encoder gives a distance that says nothing about the stolen vector. This is the one genuine specificity in the attack, and it is why the first triage question after an index leak is which encoder produced these vectors and who can still call it.
  • How does the cost scale if they stole a million vectors rather than a thousand?
    Sub-linearly in the part that matters. The expensive one-time item is building a stand-in against that encoder; after that, each additional vector costs only its own refinement queries, and the attacker can cherry-pick. So volume does not protect you, and "they would never mine all of it" is not a defence when they only need the rows about one person.
  • Does the reconstruction being only a paraphrase reduce the severity?
    Rarely, because the surviving tokens are the identifying ones. Names, dates, amounts and rare terms move a vector the most and come back first, so paraphrase fidelity still resolves who a record is about and roughly what it says. Severity should be argued from what was recovered about a person, not from a similarity score against the original wording.

saying these in an interview costs you the question

  • Claiming the attacker needs the encoder's weights
  • Claiming they need samples of the victim's corpus
  • Saying any embedding model will do as the oracle
  • Describing the output as byte-exact recovery
  • Treating a weak in-house reconstruction as proof the vectors are safe
  • Ignoring the query budget as the attack's real cost

context