How does LlamaIndex's IngestionPipeline avoid re-embedding unchanged documents?
answer
- two mechanisms, not one
- content hash versus document identity
- a strategy enum decides insert or replace
- stateless process means stateless dedup
- churning ids defeat both
basics
~20 sTwo separate mechanisms: an IngestionCache that skips a transformation when the same input hash and transformation were seen before, and an attached docstore that compares each document's id and hash to decide insert, skip or upsert. Both need stable document ids.
solid answer
~40 s`IngestionPipeline(transformations=[...])` runs its transformations in order and returns nodes. Re-run economy comes from two independent parts. First, the **cache**: each transformation's output is cached under a key derived from the input nodes plus the transformation, so identical input plus identical transformation returns the stored result rather than recomputing — that is what saves embedding calls. Second, the **docstore**: pass `docstore=SimpleDocumentStore()` and the pipeline records each document's id and content hash, then applies a `docstore_strategy` — `DocstoreStrategy.UPSERTS` re-processes only documents whose hash changed and skips the rest, `DUPLICATES_ONLY` just avoids duplicate inserts, and `UPSERTS_AND_DELETE` also removes documents absent from the new batch. Both depend on **stable document ids**; with generated per-run ids everything looks new and the pipeline re-embeds the corpus. `pipeline.persist(dir)` and `load(dir)` keep that state between processes.
code
python · 15 linesfrom llama_index.core import SimpleDirectoryReader
from llama_index.core.ingestion import IngestionPipeline, DocstoreStrategy
from llama_index.core.node_parser import SentenceSplitter
from llama_index.core.storage.docstore import SimpleDocumentStore
docs = SimpleDirectoryReader("./data", filename_as_id=True).load_data()
pipeline = IngestionPipeline(
transformations=[SentenceSplitter(chunk_size=512, chunk_overlap=32)],
docstore=SimpleDocumentStore(),
docstore_strategy=DocstoreStrategy.UPSERTS,
)
nodes = pipeline.run(documents=docs)
pipeline.persist("./pipeline_storage") # docstore + cache survive the processgo deeper
Know that an ingestion pipeline chains transformations over Documents and returns nodes, and that re-running it blindly can duplicate content unless something tracks what was already ingested.
Explain the two mechanisms separately — a content-hashed cache that skips recomputation, and a docstore that compares document ids and hashes — and name the docstore strategies that decide insert, skip or replace.
Diagnose the real failures: unstable ids, non-persisted docstore and cache, vector stores that cannot delete by ref doc, and unbounded cache growth. Be ready to describe how you made a nightly ingest genuinely incremental.
Own the refresh design end to end — full rescan versus source change feed, where dedup state lives so many workers share it, what a re-embedding event costs when the embedding model changes, and who is allowed to trigger one.
## The construct ``` from llama_index.core.ingestion import IngestionPipeline, DocstoreStrategy from llama_index.core.storage.docstore import SimpleDocumentStore pipeline = IngestionPipeline( transformations=[splitter, extractor, embed_model], docstore=SimpleDocumentStore(), vector_store=vector_store, docstore_strategy=DocstoreStrategy.UPSERTS, ) nodes = pipeline.run(documents=docs, num_workers=4) ``` `run()` applies each transformation in sequence — parsing Documents into nodes, running metadata extractors, embedding — and returns the resulting nodes. If a `vector_store` is attached, embedded nodes are written there as well. On a second run over a mostly unchanged corpus, two different things stop work from happening, and interviews reward being precise about which is which. ## Mechanism one: the cache The pipeline carries an `IngestionCache`. For each transformation it computes a key from the input nodes and the transformation itself; if that key is present, the stored output is returned instead of running the step. This is what avoids paying an embedding provider twice for text that has not changed, and it also skips LLM-backed metadata extractors, which are frequently the most expensive stage in the pipeline. Things to know about the cache in production: - It is keyed by **content**, not by document id. Two identical chunks from different files hit the same entry. - It grows. `pipeline.cache.clear()` exists precisely because an unbounded local cache on a long-lived ingestion box is a real operational problem. `IngestionCache` can be backed by a remote store (Redis, for example) so that several workers share it. - Changing a transformation's configuration changes the key — a new chunk size or a new embedding model correctly invalidates, rather than silently reusing stale vectors. ## Mechanism two: the docstore The cache saves recomputation; it does not stop the pipeline from inserting the *same document again* under a new node id. That is the docstore's job. When a docstore is attached, the pipeline looks up each incoming Document by `id_` and compares its stored hash: - **Unseen id** → process and insert. - **Known id, hash unchanged** → skip entirely. - **Known id, hash changed** → under `DocstoreStrategy.UPSERTS`, delete the old nodes for that document and re-process it. `DocstoreStrategy` options are `UPSERTS`, `DUPLICATES_ONLY` and `UPSERTS_AND_DELETE`. The last also removes documents that exist in the store but are absent from the batch — correct only when you pass the *complete* corpus each run; pass a delta batch with that strategy and you will delete everything you did not resend. Upserting into a vector store requires that store to support deletion by document reference. Not every backend does this equally well, and "my upserts leave orphan vectors behind" is usually a store-capability problem, not a pipeline bug. ## The failure that dominates Almost every "my pipeline re-embeds everything every night" report is the same root cause: **document ids are not stable**. `SimpleDirectoryReader` generates a fresh id per Document unless you pass `filename_as_id=True`; a custom reader does whatever you wrote. With churning ids, the docstore sees an entirely new corpus, `UPSERTS` degrades to insert-everything, and duplicates accumulate in the vector store. The content cache may still spare you the embedding calls — which is why the symptom is sometimes "duplicates but no extra cost", and sometimes "extra cost too" if the text also changed slightly. Second most common: state that does not survive the process. An in-memory `SimpleDocumentStore` and the default cache both vanish when the job exits. Use `pipeline.persist("./pipeline_storage")` and `pipeline.load("./pipeline_storage")`, or back the docstore and cache with a real database, if the pipeline runs as a scheduled job. ## Operational shape `run(num_workers=N)` parallelises across a process pool; there is an async `arun()` for I/O-bound transformations. Sizing here is about the embedding provider's rate limits more than about CPU. And a pipeline that both dedups and caches still cannot fix an upstream that hands it slightly-different text every run — normalising whitespace and rendering deterministically in the reader is what makes hashing work at all.
- What is the difference between what the cache saves and what the docstore saves?The cache is content-keyed and prevents recomputation — no second embedding call for text already transformed. The docstore is identity-keyed and prevents duplicate or stale records — it decides whether a document is inserted, skipped or replaced. You can have the cache hit while the docstore still inserts a duplicate, which is exactly what happens when document ids churn.
- When is DocstoreStrategy.UPSERTS_AND_DELETE the wrong choice?Whenever a run receives only a delta rather than the whole corpus. That strategy removes documents present in the docstore but missing from the batch, so an incremental run of ten changed files would delete everything else. Use it only for full-corpus refreshes, and prefer explicit deletion driven by a source change feed otherwise.
- Your pipeline runs as a nightly container and still re-processes everything. What do you check first?Document id stability — `filename_as_id=True` on the reader, or a source-derived id_ in a custom reader — then state durability. An in-memory docstore and cache die with the process, so persist them with `pipeline.persist()`/`load()` or back them with a shared database. Those two causes account for nearly all such reports.
saying these in an interview costs you the question
- Thinks the cache alone prevents duplicate documents
- Assumes an in-memory docstore survives between job runs
- Uses UPSERTS_AND_DELETE with incremental batches
- Believes changing chunk size still reuses cached vectors
- Ignores that generated document ids change every run