What happens to an encoder-tuned planted passage when a corpus is re-embedded under a new model?
answer
- the fit was to a function, not a topic
- the document stays, the coordinates move
- two models share no coordinate meaning
- what survives is what it plainly says
- expiry set by someone else's release cadence
basics
~20 sThe fit dies while the document survives. A re-embed recomputes every vector with a different function, so a passage shaped against the old encoder's geometry lands somewhere unremarkable. What still ranks is whatever the passage plainly says.
solid answer
~50 sTuning a passage to an encoder means fitting it to the geometry of one specific function. When the pipeline's pinned model version is upgraded, every chunk is passed through the new encoder and the index is rebuilt, because vectors produced by two different models are not comparable coordinates. The planted document is not removed by any of this — it is still in the corpus, still eligible, simply no longer near the questions it was fitted to. Whatever coarse subject-matter content it has survives, because both encoders put text about the same topic in broadly similar company; the fitted part does not, because it was solved against a function that is gone. This is why the family is the fragile end of corpus poisoning rather than the sophisticated end: its lifetime is set by an upgrade schedule the attacker neither controls nor observes, and nothing announces the expiry.
code
json · 17 lines{
"encoder_pin": "[email protected]",
"top_k": 5,
"results": [
{"rank": 1, "chunk_id": "sub-4471#0", "score": 0.91, "source": "open-submission", "body": "[tuned span elided]"},
{"rank": 2, "chunk_id": "rev-0912#3", "score": 0.74, "source": "curated-review"}
]
}
...
{
"encoder_pin": "[email protected]",
"top_k": 5,
"results": [
{"rank": 1, "chunk_id": "rev-0912#3", "score": 0.68, "source": "curated-review"}
],
"probe": {"chunk_id": "sub-4471#0", "rank": 214, "score": 0.29, "indexed": true}
}go deeper
Remember the one-line fact: embeddings from different models are not comparable, so upgrading the model means recomputing every vector in the index.
Be ready to separate three things a re-embed does — keeps the document, changes its coordinates, destroys a fit — and to say why coarse topical meaning transfers while a fitted margin does not.
Show you can reason about the attacker's information position: retrieval stopping is an unlabelled event, and several very different causes produce the identical observation of absence.
Be able to tell an owner plainly what an upgrade did and did not establish, so a non-reproduction after a version bump is not filed as a resolved root cause.
### What tuning to an encoder actually commits you to A passage shaped against a particular embedding model is not shaped against a subject, a question, or a reader. It is fitted to the geometry of one function: this model, this version, this pooling and normalisation behaviour. The attacker found the text by evaluating that function repeatedly and keeping what moved the vector toward the region where anticipated questions land. Everything the construction is worth is a property of that function. ### What a re-embed is, and when it happens A pipeline pins an embedding model version, because vectors from two different models share no coordinate meaning at all — different training, often different dimensionality, and no reason for the same axis to mean the same thing. So the index cannot hold a mixture. When the pipeline upgrades that pin, every chunk in the corpus is passed through the new encoder and the index is rebuilt from scratch. That is a routine operational event driven by the pipeline's own release cadence, not a response to anything the attacker did — which is precisely what makes it hard to plan around. ### What survives it Three things happen at once, and separating them is the question: - **The document survives.** Nothing about a re-embed removes a submission, its metadata, or its eligibility. The plant is still in the corpus. - **The coordinates change.** Every chunk gets a new vector, including the plant's. - **The fit is lost.** Whatever distortion pulled the old vector toward the query region has no reason to do the same thing in a differently trained space. Transfer, when it happens, is accidental rather than designed. What does carry over is ordinary semantics. Two competent encoders both place text about the same subject in broadly similar company, so a passage that genuinely reads like it is about the topic keeps some standing after a re-embed. A passage whose only claim to rank was geometry becomes inert — present, indexed, and never returned. ### Why this makes the family the fragile one The common wrong answer is that this is the sophisticated version of corpus poisoning, because it engages with the machinery rather than merely writing convincing prose. The mechanics point the other way. The artefact is fitted to one encoder version, it is ended by an upgrade the attacker does not control, and it is ended independently by any later stage that scores the question and the passage together, because such a stage reads text and this text reads badly. It is the construction with the shortest and least predictable life in the family. The attacker's information position makes it worse. There is no failure signal. Retrieval simply stops, and stopping looks identical whether the corpus was re-embedded, the candidate budget tightened, a new competitor entered the region, or the submission was removed. Re-verification is not free either: it means learning the new pin, re-scoring offline against the new function, and submitting again. ### Direction of the claims A re-embed ending a plant's retrieval does not establish that the plant was found, examined or removed, and it does not establish that unreviewed text is no longer eligible for the candidate set. It establishes one thing: the coordinates the construction depended on no longer exist. Reading a post-upgrade non-reproduction as a resolved finding is the mistake this question exists to prevent.
- Why can the fitted passage keep no advantage in the new space when both encoders were trained on similar text?Because the advantage was never semantic. It came from where one particular function happened to place an unusual string relative to a query region — a quirk of that model's learned geometry. A differently trained model has no reason to reproduce the quirk, even if it agrees broadly about what the text is about. Coarse topical placement transfers; the fitted margin does not.
- What does the attacker actually see when a re-embed ends their plant?Absence. The passage stops being returned, and that looks the same as a tightened candidate budget, a new competitor in the region, a changed query mix, or a removed submission. There is no error and no notification, so the only way to distinguish them is to probe again — which costs another round of offline scoring against a pin they must first discover.
- Does the re-embed remove the document from the corpus?No. It recomputes vectors for everything already indexed. The submission, its text and its metadata are untouched and still eligible for retrieval; it simply is not near anything anymore. Confusing a lost fit with a removed document is how people mistake an upgrade for a remediation.
saying these in an interview costs you the question
- Calls encoder-tuned text the robust or sophisticated form of poisoning
- Assumes vectors from two models are comparable
- Thinks a re-embed deletes the planted document
- Expects a fitted margin to transfer to a new encoder
- Reads a post-upgrade non-reproduction as a fixed root cause