In Haystack, what changes when you swap InMemoryDocumentStore for a database store?
answer
- The protocol is four methods plus serialization
- Writer stays, retriever moves
- Vector width is decided at creation
- Keyword scoring is the backend's, not Haystack's
- Nothing could fail before; now it can
basics
~20 sThe DocumentStore protocol keeps write_documents, filter_documents, count_documents and delete_documents identical, so the indexing pipeline is unchanged. What changes is the store class and its connection config, the fixed vector dimension, the store-specific retriever, and who owns durability.
solid answer
~50 sHaystack's `DocumentStore` is a protocol, not a base class: any implementation offers `write_documents`, `filter_documents`, `count_documents`, `delete_documents` plus `to_dict`/`from_dict`. That is why `DocumentWriter` and the whole indexing chain survive the swap untouched — you change one constructor. What does not survive: the store class now comes from a separate `haystack_integrations` package with its own dependency and connection settings; the vector dimension and distance function are fixed when the index or table is created, so the store and the embedder must agree; the embedding retriever is store-specific and must be swapped with it; keyword search semantics change, since a database store uses its backend's text search rather than the in-memory BM25 implementation; and warm-up, connection failures and index lifecycle become real operational concerns. Durability moves too — the in-memory store dies with the process, so tests that relied on a fresh empty store now need explicit cleanup.
go deeper
Know that InMemoryDocumentStore holds documents in the process and is meant for development and tests, and that production stores are separate integration packages you install alongside Haystack.
Name the protocol methods and explain why the indexing pipeline is unchanged by a swap while the retriever is not. Know that vector dimension is fixed when the index is created.
Show migration experience: config threading, an index-creation deployment step, a delete-then-write re-index path, and a retrieval smoke test against the real store because keyword scoring and filter edge cases do not port.
Frame the one-way doors. Embedding dimension and distance function are baked into the index, so model choice, re-index strategy and cutover plan are a single coupled decision — and the abstraction's value is that it postpones, not removes, that commitment.
## The protocol is the portable part `DocumentStore` in Haystack is a Python protocol. An implementation must provide: - `write_documents(documents, policy=DuplicatePolicy.NONE) -> int` - `filter_documents(filters=None) -> List[Document]` - `count_documents() -> int` - `delete_documents(document_ids) -> None` - `to_dict()` / `from_dict()` for pipeline serialization Everything an indexing pipeline needs is in that list. `DocumentWriter` calls `write_documents`; your re-index logic calls `filter_documents` and `delete_documents`; your tests call `count_documents`. Swap the object handed to `DocumentWriter(document_store=...)` and the converter, cleaner, splitter, embedder and writer are all unchanged. That is the real payoff of Haystack's explicit wiring: the storage decision is one constructor call, not a rewrite. ## What the protocol does not cover **Construction and connection.** Integration stores live in separate distributions under the `haystack_integrations` namespace, each with its own install and its own constructor arguments — connection strings, index or table names, an option to recreate the index. Those arguments are where all the environment-specific configuration now lives, and they are the part that must be threaded through your deployment config rather than hard-coded. **Vector dimension and distance function.** An in-memory store is happy to hold vectors of any size and picks a similarity function at construction. A database-backed store creates an index or a table with a fixed vector dimension and a fixed distance metric. If the store was created for 768 dimensions and your embedder produces 1536, writes are rejected — and there is no way to "fix" it other than recreating the index and re-embedding. Changing embedding model is therefore an index-rebuild operation, not a config change. **The retriever.** Each store integration ships its own retriever components, because the query has to be expressed in the backend's language. Moving stores means moving the retriever in the query pipeline as well. The rest of the query pipeline — prompt builder, generator — is untouched. **Keyword search.** `InMemoryDocumentStore` implements BM25 in Python, configurable through its own constructor. A search-engine-backed store uses that engine's analysis and scoring instead. Scores are not comparable across the two, and neither are the tokenization rules, so a keyword result set verified against the in-memory store is not evidence about production. This bites hardest in hybrid setups where score fusion is involved. **Filter edge cases.** Every store translates Haystack's filter dicts into its own query language. The common cases behave identically; the edges — missing keys, nulls, numeric-versus-string coercion, nested meta paths — can differ. Validate real filters against the real store. **Lifecycle and failure.** The in-memory store cannot fail to connect, cannot be unavailable, and cannot be shared between processes. A database store can do all three. Errors surface where there previously were none, index creation becomes a deployment step, and two processes writing concurrently is now a real scenario rather than an impossibility. ## Durability, and the testing habit it creates `InMemoryDocumentStore` lives in the process and dies with it, which makes it perfect for tests and for demos; it can also serialize itself to disk and be reloaded, which covers small static corpora. A team that develops entirely against it acquires an unstated assumption — "each run starts empty" — that is false everywhere else. Migrating usually surfaces this as tests that pass alone and fail in sequence, and as re-indexing logic that was never written because it was never needed. ## How to make the swap cheap - Construct the store in exactly one place and inject it into `DocumentWriter` and the retriever. Never construct one inline in a component. - Keep the embedding model and the store's vector dimension declared together, so a mismatch is a config error at startup rather than a write failure in the middle of a batch. - Write the delete-then-write re-index path from day one, against the protocol methods, so it does not have to be invented under pressure. - Run at least a smoke test of retrieval against the real store in CI. Protocol compatibility guarantees the calls work; it guarantees nothing about the ranking you get back. ## The interview answer Say that the protocol makes the *plumbing* portable and the *behaviour* only mostly portable: the writer and indexing chain are untouched, but the retriever, the vector dimension, the keyword scoring and the operational failure modes are all store-specific. Then name the one-way door — changing embedding dimension means rebuilding the index — because that is the decision an interviewer actually wants to hear you anticipate.
- Which methods does the DocumentStore protocol require?`write_documents(documents, policy)`, `filter_documents(filters)`, `count_documents()`, `delete_documents(document_ids)`, plus `to_dict`/`from_dict` so a store embedded in a pipeline can be serialized. Notice what is absent: nothing about vector search. Similarity search is expressed through store-specific retriever components, which is exactly why retrievers do not port across stores.
- Why is changing embedding model a bigger deal on a database-backed store?Because the index or table is created with a fixed vector dimension and distance function. A new model with a different dimension cannot be written into the existing index at all, so you must create a new index, re-embed the whole corpus, and cut over. Even a same-dimension model change requires a full re-embed, since old and new vectors are not comparable.
- Your retrieval tests pass against InMemoryDocumentStore but rank differently in production. Why?Keyword scoring is not portable. The in-memory store implements BM25 in Python with its own tokenization; a search-engine-backed store applies its own analyzers and scoring. Absolute scores and often the ordering differ, and in hybrid pipelines that difference propagates into score fusion. Ranking assertions belong in a test that runs against the real store.
- What does the in-memory store let you get away with that a real one does not?Assuming a fresh empty store on every run, assuming writes never fail, assuming a single process owns the data, and never writing a deletion path. Those assumptions are invisible until they break, which is why teams hit re-indexing and concurrency problems on the day they migrate rather than during development.
saying these in an interview costs you the question
- Assuming a store swap is a drop-in with zero pipeline changes
- Thinking the DocumentStore protocol includes vector search methods
- Expecting a store to adapt to whatever embedding dimension arrives
- Trusting BM25 rankings measured against the in-memory store
- Treating in-memory persistence to disk as equivalent to a database