skip to content

Why does mass-duplicating one planted passage buy no reach when a pipeline collapses near-copies?

level: middleimportance: should knowfreq 44%

answer

  1. position in the space, not the count
  2. copies land on the same coordinates
  3. one representative survives ingest
  4. differing enough to survive means aimed elsewhere
  5. each neighbourhood costs a distinct document

basics

~10 s

Near-identical text embeds to nearly the same vector, so copies chase the same questions rather than new ones, and ingest that keeps one representative removes most of them. Reach follows position, not count.

solid answer

~50 s

Two effects stack. First, even with no deduplication anywhere, fifty copies of one passage sit in the same place in the embedding space: they match the same phrasings and compete with each other for slots on those queries. Volume there buys depth on one question — several of the retrieved slots — never breadth across other questions. Second, ingest pipelines commonly keep a single representative of near-duplicate text, so most of the copies never become separate candidates at all. The consequence for whoever is doing the writing is the real point: covering a wider family of buyer questions costs *genuinely differing* documents. And "differing" has to mean differing in the ways the encoder registers — which is the same thing as landing in different query neighbourhoods. The cost of surviving collapse and the cost of gaining breadth are one cost, paid per distinct, attributable post.

go deeper

for a junior

Remember that identical text ends up in the same place in the vector space, so extra copies chase the same questions. Volume is not the lever people expect it to be.

for a middle

Explain both effects and keep them separate: copies share a position regardless of deduplication, and near-copy collapse removes most of them at ingest. Then draw the consequence for cost.

for a senior

Be ready to re-read someone else's result with this in mind: ask how many distinct passages a large write count actually represented, and how many question neighbourhoods moved as a result.

for a principal

The framing to own is that breadth is priced per distinct document, so a corpus-poisoning risk statement should be expressed in documents-per-question-family, not in raw write counts.

## The move this question kills Someone who has just learned that a planted passage owns a narrow query neighbourhood reaches for the obvious scaling lever: post it again, under more listings, more accounts, more product pages. On a keyword-matching system that instinct is partly sound. On a retrieval pipeline it is close to worthless, for two independent reasons. ## Reason one: copies occupy the same position First-stage retrieval embeds the query and each stored passage separately and ranks by vector distance. Text that is identical, or paraphrased so lightly that the encoder cannot tell, produces vectors that are essentially the same point. Fifty copies are therefore fifty candidates at the same coordinates: they are near the same phrasings, far from the same phrasings, and the questions they can reach are exactly the questions one copy could reach. What they *do* buy is depth on that one neighbourhood — with no collapse anywhere, several of the retrieved slots for that question could be copies of the same claim, pushing genuine passages out of the context the generator sees. That is a real effect on one question. It is still not breadth, and it is the most conspicuous shape a corpus can take. ## Reason two: near-copy collapse Many ingestion paths keep one representative of duplicated or near-duplicated text. Where that exists, the fifty copies do not even arrive as fifty candidates: one survives, the rest are dropped or merged before anything is retrievable. Note what this means for the writer's feedback loop — from outside the system, a collapsed copy and a never-ingested copy look identical, because there is no way to query the index and count what is in it. Collapse can sit at more than one stage; the attacker-relevant fact is not where it runs but that repetition of the same text is not a reliable multiplier. ## Why the two costs turn out to be the same cost The repair looks obvious: make the copies different. But how different? - Cosmetic variation — reordered sentences, swapped synonyms, a different sign-off — may or may not clear a near-copy check, and it does not move the passage to a new place in the embedding space. It still answers the same question. - Variation that *does* clear collapse reliably is variation the encoder registers: different claims, different product attributes, different vocabulary buyers actually use. But a passage that lands somewhere else in the embedding space is, by definition, a passage aimed at a different set of questions. So the effort to survive deduplication and the effort to widen coverage are not two bills. They are one bill: **each new query neighbourhood costs one genuinely distinct document.** | What you post | What it can reach | What it costs | | --- | --- | --- | | One passage | The phrasings near it, where it outranks genuine passages | One post | | Fifty copies of it | The same phrasings; possibly more slots on them | Fifty attributable posts, most collapsed | | Fifty distinct passages | Up to fifty neighbourhoods, minus overlaps | Fifty posts and fifty pieces of real writing | ## The costs that are not about retrieval at all On a marketplace, each post is attributable to a listing and an account. Volume of near-identical text under listings one seller controls is the pattern a marketplace's own content review is most likely to notice — the writer is paying attributable posts for candidates that a collapse may discard anyway. And unlike the retrieval effect, that cost is paid immediately and cannot be undone by editing later. ## What this does to a finding Someone triaging a report that says "we planted 200 documents and the assistant repeats our claim" should ask how many *distinct* passages those 200 were, and how many distinct questions moved. If the 200 were near-copies aimed at one question, the honest description is one moved neighbourhood at a cost of 200 attributable posts — an expensive result, not a broad one. The count of writes is not the size of the effect; the count of moved question neighbourhoods is. ## Directions to keep straight - A collapsed copy proves the pipeline kept one representative, not that the text was judged harmful. - A passage appearing in a retrieved set proves it ranked, not that it displaced anything — you have to look at what else was there. - Failing to see an effect after posting copies proves nothing about ingestion timing; from outside, a dropped copy, a collapsed copy and a copy awaiting the next build are indistinguishable.

  • Do duplicates buy anything at all, if nothing collapses them?
    Depth on a single question. Several of the retrieved slots for that one phrasing can be copies of the same claim, squeezing genuine passages out of the context the generator sees. It never reaches a new question, and it is the most visible shape a corpus can take, so the marginal value falls fast.
  • Would light paraphrasing be enough to clear a near-copy check?
    Unreliably, and that is the trap. Cosmetic edits may or may not clear the check, and even when they do they leave the passage in the same region of the embedding space, so it still answers the same question. Variation the encoder actually registers is variation that aims the passage elsewhere — which is the breadth you were trying to buy anyway.

saying these in an interview costs you the question

  • Treats posting volume as a direct multiplier on reach
  • Assumes paraphrasing reliably defeats near-copy collapse
  • Confuses number of writes with number of questions moved
  • Thinks duplicate passages reach queries a single copy could not
  • Reads a missing effect as proof the copies were rejected

context