In Haystack, what does DocumentWriter's DuplicatePolicy control?
answer
- Four members, one enum
- Matching is by document id
- The default is not the safe one
- InMemory resolves the vague one to a raise
- Replace means replace, not merge
basics
~20 sDuplicatePolicy tells a Haystack DocumentWriter what to do when an incoming document's id already exists in the store: OVERWRITE replaces the stored copy, SKIP keeps it, FAIL raises DuplicateDocumentError, and NONE defers to the store's own default.
solid answer
~40 s`DocumentWriter(document_store=..., policy=...)` takes a `DuplicatePolicy` from `haystack.document_stores.types`, and it is matched on **document id only** — never on content similarity. `OVERWRITE` replaces the stored document, `SKIP` leaves the stored copy and drops the incoming one, `FAIL` raises `DuplicateDocumentError`, and `NONE` — the default — means "whatever this store considers its default", which `InMemoryDocumentStore` resolves to `FAIL`. That default is the one that bites: a second run of an indexing pipeline over unchanged files blows up rather than being a no-op, so most production indexing pipelines set `OVERWRITE` explicitly. Because ids are content hashes unless you set them yourself, `OVERWRITE` gives you idempotent re-indexing only while the pipeline produces byte-identical text and metadata.
code
python · 19 linesfrom haystack import Document
from haystack.components.writers import DocumentWriter
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.document_stores.types import DuplicatePolicy
store = InMemoryDocumentStore()
writer = DocumentWriter(document_store=store, policy=DuplicatePolicy.OVERWRITE)
docs = [Document(content="Haystack pipelines are explicitly wired.")]
print(writer.run(documents=docs)) # {'documents_written': 1}
print(writer.run(documents=docs)) # {'documents_written': 1} - same id, replaced
print(store.count_documents()) # 1
# Same text, different meta -> different id -> not a duplicate at all
store.write_documents(
[Document(content="Haystack pipelines are explicitly wired.", meta={"v": 2})],
policy=DuplicatePolicy.OVERWRITE,
)
print(store.count_documents()) # 2go deeper
Know the four names and that they apply when a document with the same id is already stored. Being able to say "the default raises on InMemoryDocumentStore, so set OVERWRITE for a re-runnable pipeline" is enough here.
Explain that matching is by id only, that ids are content-derived by default, and that OVERWRITE replaces rather than merges. Interviewers expect you to reason about what a second pipeline run does.
Demonstrate that you have designed for idempotent re-indexing: deterministic ids or stable pipeline output, an explicit delete path for chunks that vanish, and checking documents_written to catch silent SKIPs.
Frame it as an identity-and-lifecycle question, not a flag. Who owns document identity, how deletions propagate from the source system, and whether the store is treated as a rebuildable cache or as durable state are the calls you should be making.
## What the policy is `DuplicatePolicy` is an enum in `haystack.document_stores.types` with four members: `NONE`, `SKIP`, `OVERWRITE`, `FAIL`. It is passed either directly to a store's `write_documents(documents, policy=...)` or, in a pipeline, to the `DocumentWriter` component at construction time: `DocumentWriter(document_store=store, policy=DuplicatePolicy.OVERWRITE)` The writer holds the policy for its lifetime; it is not a per-run input. ## Matching is by id, and only by id A "duplicate" in Haystack means a document whose `id` already exists in the store. It has nothing to do with the text being similar, or even identical-but-differently-tagged. If you construct `Document(content="x")` twice you get the same id both times, because the id defaults to a hash derived from the content and the meta. If you construct `Document(content="x", meta={"v": 1})` and `Document(content="x", meta={"v": 2})`, you get two different ids and the store cheerfully holds both copies of the same sentence — no policy will stop that. This is the single most important thing to say in an interview: DuplicatePolicy is an id-collision policy, not a deduplication feature. ## The four members **OVERWRITE** — the incoming document replaces the stored one wholesale. Meta, embedding and content are all replaced, not merged; a field present on the stored copy and absent on the incoming one disappears. This is what you want for an indexing pipeline that is re-run on a schedule, because it makes the write idempotent for unchanged inputs. **SKIP** — the stored document wins and the incoming one is discarded, typically with a warning logged. Useful when the store is authoritative and you are backfilling from a lower-quality source, and cheap when writes are expensive. The trap is that a genuinely updated document with an id you assigned yourself will be silently ignored. **FAIL** — raises `DuplicateDocumentError` (from `haystack.document_stores.errors`) on the first collision. Loud and safe for a one-shot bulk load where a collision means your input set is wrong, but hostile to any pipeline that runs more than once. **NONE** — the literal meaning is "I have no opinion; use the store's default". It is the default value both on `DocumentWriter` and on the protocol's `write_documents` signature. Each store resolves it: `InMemoryDocumentStore` resolves `NONE` to `FAIL`, and other integrations pick their own resolution. Because it is indirection, `NONE` is the wrong thing to leave in production code — say what you mean. ## The failure people actually hit A developer builds an indexing pipeline, runs it, gets documents in the store, runs it again over the same folder, and gets `DuplicateDocumentError`. The cause is the `NONE`-resolves-to-`FAIL` chain, and the fix is one argument. The mirror-image failure is subtler: they set `OVERWRITE`, re-run, and the store *doubles* in size. That is not a policy bug — it means the ids changed between runs, because the content or the meta changed (a new metadata key, a different splitter setting, a converter upgrade that extracts whitespace differently). `OVERWRITE` can only overwrite a row it can find by id. ## Taking control of ids If you set `Document(id=...)` explicitly, the auto-hash is bypassed and you own the identity. A common production pattern is a deterministic id built from a stable business key plus the chunk index — `f"{document_uri}::{chunk_index}"` — so that re-indexing a changed file overwrites exactly the chunks it should. The cost is that if a file gains or loses chunks, the surplus old chunks keep their ids and stay in the store until you delete them; a hash-based id at least never *wrongly* overwrites. Neither scheme removes the need for an explicit deletion step when a source document shrinks or disappears. ## What the writer reports `DocumentWriter.run(documents=...)` returns `{"documents_written": n}`. Under `SKIP`, `n` counts what actually landed, so comparing it against the input length is a cheap way to detect that a nightly run is being silently skipped rather than applied. Under `FAIL` you get an exception instead of a count, which in a pipeline aborts the whole run — there is no partial-success mode where the good documents land and the collisions are reported. ## Choosing one Re-runnable indexing pipeline: `OVERWRITE`. One-shot bulk load where a collision indicates a bug in your input assembly: `FAIL`. Backfill against an authoritative store, or an append-only ingest where re-processing is expensive: `SKIP`. Almost never `NONE` in code you intend to keep.
- What does DuplicatePolicy.NONE actually do?It means "use this store's default". It is the default value on both `DocumentWriter` and `write_documents`, and each store resolves it independently — `InMemoryDocumentStore` resolves it to `FAIL`, so a second run over unchanged input raises `DuplicateDocumentError`. Because the behaviour is indirection through the store, production code should set an explicit policy rather than rely on it.
- Does OVERWRITE merge the incoming metadata with what is already stored?No. The stored document is replaced wholesale — content, meta and embedding. A meta key that exists on the stored copy but not on the incoming one is gone after the write. If you need a merge you have to read the document back with `filter_documents`, combine the meta yourself, and write the combined document.
- Two Documents hold the same sentence but different meta. Will SKIP deduplicate them?No. The auto-generated id is derived from content *and* meta, so different meta means different ids and no collision at all — both are written. DuplicatePolicy is an id-collision policy, not a near-duplicate detector. Real deduplication needs either ids you assign from a business key, or a similarity pass before the writer.
- How do you make re-indexing genuinely idempotent?Either keep the pipeline byte-stable so the content hashes repeat and `OVERWRITE` lands on the same ids, or assign deterministic ids yourself from a source key plus chunk index. Both still need an explicit delete step for chunks that no longer exist, since neither policy can remove rows the current run does not produce.
saying these in an interview costs you the question
- Thinking DuplicatePolicy detects similar or near-duplicate content
- Assuming the default policy makes re-indexing a safe no-op
- Believing OVERWRITE merges metadata into the stored document
- Expecting SKIP to apply updates to an existing document
- Setting OVERWRITE and assuming the store can no longer grow on re-runs