Why does a nightly Haystack indexing pipeline keep adding duplicate chunks?
answer
- Ids are fingerprints of pipeline output
- Change the text, change the identity
- OVERWRITE needs a row to find
- A shrinking file leaves orphans
- Delete by source metadata first
basics
~20 sBecause a Haystack Document's auto-generated id is a hash of its content and meta. Any change to the extracted text, splitter settings or metadata produces new ids, so OVERWRITE finds nothing to replace and writes the chunks as new rows.
solid answer
~50 sUnless you set `Document(id=...)` yourself, Haystack derives the id from the document's content and its `meta`. That makes `DuplicatePolicy.OVERWRITE` idempotent only while the pipeline emits byte-identical output. Upgrade a PDF converter, flip a `DocumentCleaner` flag, change `split_length`, or add one metadata key at conversion time, and every chunk hashes differently — the writer sees no collision and appends. The store grows by a full corpus per run, and retrieval starts returning several near-identical chunks. The fix is not a policy: it is an explicit delete-before-write keyed on stable metadata. Call `filter_documents` with a filter on `meta.file_path` (or your own source key), pass those ids to `delete_documents`, then write the fresh chunks. The alternative is deterministic ids you assign from a source key plus chunk index — but that still needs a delete step when a file shrinks.
code
python · 18 linesfrom haystack import Document
from haystack.document_stores.in_memory import InMemoryDocumentStore
store = InMemoryDocumentStore()
# Same text, one extra meta key -> different id -> not a duplicate
a = Document(content="Refunds are processed in 5 days.", meta={"file_path": "faq.md"})
b = Document(content="Refunds are processed in 5 days.", meta={"file_path": "faq.md", "run": 2})
print(a.id == b.id) # False
# The reliable re-index: delete by source, then write
store.write_documents([a])
stale = store.filter_documents(
filters={"field": "meta.file_path", "operator": "==", "value": "faq.md"}
)
store.delete_documents([d.id for d in stale])
store.write_documents([b])
print(store.count_documents()) # 1go deeper
Know that a Haystack Document's id is generated from its content and meta rather than being random, so identical input yields the same id and changed input yields a new one.
Explain why OVERWRITE stops working when the pipeline's output text or metadata changes, and name the concrete triggers: converter upgrades, cleaner flags, splitter settings, added meta keys.
Show the operational fix: a stable source key in every chunk's meta, filter-then-delete before re-writing, and an assertion on document count in CI. Be ready to explain why a shrinking document is the case policies cannot cover.
Own the index lifecycle. Decide whether the index is a rebuildable cache or durable state, whether re-indexing is incremental or blue/green with a pointer swap, and how deletions in the source system propagate to retrieval.
## The mechanism When you build a `Document` without an `id`, Haystack computes one — a SHA-256 hash derived from the document's content, any binary blob, and its `meta` dict. This is deliberate: it makes the same input produce the same id everywhere, with no coordination and no database round-trip. It also means **the id is a fingerprint of the pipeline's output, not of the source file**. Every stage of an indexing pipeline touches that fingerprint: - The converter sets `file_path` in meta and produces text. A library upgrade that extracts one extra space changes the hash of every chunk downstream. - `DocumentCleaner` rewrites content, and its `keep_id` flag defaults to `False`, so cleaned documents get new ids by design. - `DocumentSplitter` produces entirely new Documents and stamps `source_id`, `split_id`, `split_idx_start` and `page_number` into their meta. Change `split_length` or `split_overlap` and both the content and the split metadata shift. - Any per-file meta you attach at conversion time is inside the hash too, so adding an `ingested_at` timestamp guarantees a fresh id on every single run. ## Why OVERWRITE does not save you `DuplicatePolicy.OVERWRITE` replaces a stored document *with the same id*. If the ids moved, there is no collision to resolve, so the writer simply inserts. The store doubles. Nothing errors, nothing warns, and `documents_written` reports the expected number — which is why this is usually discovered by a retrieval regression ("the top 5 results are five copies of the same paragraph") or by a storage-cost alert, not by the pipeline. The reverse mistake exists too: someone stamps a run timestamp into meta precisely so that they can tell runs apart, and thereby guarantees the corpus grows forever. ## The correct pattern: delete by source, then write The unit of change is the *source document*, not the chunk. So make the source identifiable in every chunk's metadata — converters already give you `meta["file_path"]`, and you can attach your own stable key (a URI, a CMS id) via the converter's `meta` argument. Re-indexing one file then becomes: 1. `store.filter_documents(filters={"field": "meta.file_path", "operator": "==", "value": path})` 2. `store.delete_documents([d.id for d in stale])` 3. run the indexing pipeline for that file with `DuplicatePolicy.OVERWRITE` This is correct regardless of how the chunk text changed, and it handles the case a policy can never handle: a file that now yields *fewer* chunks than before. No write policy can remove rows the current run does not produce. For a full-corpus rebuild the same logic scales up: index into a new store or a new index name and swap the pointer once it is populated, rather than mutating the live index in place. That also gives you an instant rollback. ## The deterministic-id alternative Instead of accepting content hashes, assign ids yourself after splitting — for example `f"{source_uri}::{split_id}"`. Then a re-run of a changed file overwrites chunks 0..n in place regardless of text drift. It is cheaper than delete-then-write (one pass, no read) and it survives converter upgrades. Its weakness is the shrinking file: if yesterday's version produced 40 chunks and today's produces 30, chunks 30..39 keep their ids, are not touched by the write, and remain in the store as stale content that retrieval will happily return. So you still need either a delete pass or a tombstone convention. Deterministic ids reduce churn; they do not remove the lifecycle problem. ## Detecting it early - Assert on `count_documents()` after a re-run in an integration test: a stable corpus should produce a stable count. - Compare `documents_written` against the expected chunk count, and alert when the store's total grows by roughly a full corpus. - Keep a stable source key in meta from day one. Retrofitting one is painful precisely because you cannot identify which existing rows came from which file. ## What to say in the interview The answer that lands is: "ids are content hashes, so idempotence is a property of the pipeline being deterministic, not of the write policy — and since a shrinking document can never be handled by a write policy at all, deletion keyed on stable source metadata has to be part of the design." That reframes it from a Haystack quirk into a lifecycle decision, which is what the question is really probing.
- Which pipeline changes silently change every chunk id?A converter upgrade that extracts text differently, any `DocumentCleaner` flag change (its `keep_id` defaults to False, so cleaning re-ids by default), new `split_by`/`split_length`/`split_overlap` settings — which alter both content and the splitter's own meta keys — and any metadata you add at conversion time, since meta is inside the hash. A run timestamp in meta guarantees it every night.
- Why can no DuplicatePolicy handle a source document that now yields fewer chunks?A write policy only decides what happens to documents the current run actually produces. Chunks from the previous run that the new run no longer emits are never presented to the writer, so nothing considers them. They stay in the store, keep their embeddings, and keep being retrieved. Only an explicit `delete_documents` call removes them.
- What do you gain and lose by assigning deterministic ids yourself?You gain stability across converter and splitter changes: chunk N of a source always overwrites chunk N. You lose the safety of a content hash — a bug that renumbers chunks silently overwrites the wrong content — and you still need a delete pass for the tail when a document shrinks. It is a churn optimisation, not a lifecycle solution.
- How would you rebuild a whole corpus safely?Index into a fresh store or index name rather than mutating the live one, verify counts and run a retrieval smoke test against it, then swap the pointer your query pipeline uses. The old index stays intact as an immediate rollback and can be dropped once the new one has served traffic.
saying these in an interview costs you the question
- Assuming DuplicatePolicy.OVERWRITE makes any re-index idempotent
- Adding an ingestion timestamp to meta without realising it re-ids everything
- Thinking the store dedupes by content similarity behind the scenes
- Believing deleted source files disappear from the index automatically
- Treating a growing document count as normal for a stable corpus