How do you keep a LlamaIndex index in sync when source documents change or are deleted?
answer
- derived data, silently stale
- the id must survive the run
- hashes decide what gets re-embedded
- removals are never inferred
- one default flag leaves a ghost behind
basics
~10 sGive every document a stable id, then call index.refresh_ref_docs(documents): LlamaIndex compares each document's hash against the docstore and re-inserts only what changed. Removals need an explicit index.delete_ref_doc(doc_id, delete_from_docstore=True).
solid answer
~50 sSync depends on two things: stable document ids and a docstore. Set `id_` on each `Document` to something derived from the source (a path, a primary key), not a random UUID, so the same source always maps to the same `ref_doc_id` on its nodes. Then `index.refresh_ref_docs(documents)` hashes each incoming document, compares it with the hash recorded in the docstore, and upserts only the changed or new ones — it returns a list of booleans telling you which were touched. For a single document, `index.update_ref_doc(document)` is delete-then-insert. Deletions are not inferred: `refresh_ref_docs` only sees the documents you hand it, so a source file that disappeared stays in the index until you call `index.delete_ref_doc(doc_id, delete_from_docstore=True)`. That flag defaults to `False`, which removes the nodes but leaves the document record behind — enough to make a later refresh think the old version is still current.
code
python · 9 linesfrom llama_index.core import Document, VectorStoreIndex
docs = [Document(text="policy v1", id_="policy.md")]
index = VectorStoreIndex.from_documents(docs)
docs = [Document(text="policy v2", id_="policy.md")]
changed = index.refresh_ref_docs(docs) # [True] -> re-embedded
index.delete_ref_doc("policy.md", delete_from_docstore=True)go deeper
Know that an index does not update itself and that refresh_ref_docs exists. Remember that documents need stable ids for any update to find the right nodes.
Explain the hash comparison in the docstore, what the returned booleans tell you, and why delete_ref_doc needs delete_from_docstore=True to fully remove a source.
Design the whole sync job: id derivation, the deleted-id diff against the previous run, metadata that must stay out of the hash, and the monitoring that catches a refresh doing nothing or doing everything.
Own freshness as a contract. Decide the acceptable staleness window, who runs the rebuild pipeline, how global invalidation from a model or chunking change is rolled out, and whether sync belongs in the framework at all versus upstream of it.
## Why sync is a design problem, not an API call A vector index is derived data. The moment the source changes, the index is wrong in a way nothing detects: queries still succeed and still look plausible, they just answer from stale text. LlamaIndex gives you the primitives, but the correctness depends on decisions you make at ingest time. ## Stable ids are the precondition Every `Document` carries an `id_`. If you do not set it, you get a fresh UUID on every run, so the same file ingested twice is two unrelated documents and every refresh is an insert. Derive the id from the source: file path, row key, URL. Nodes produced from a document record it as `ref_doc_id`, and that link is what lets LlamaIndex find every chunk belonging to a source and replace or remove it as a unit. ## refresh_ref_docs `index.refresh_ref_docs(documents)` walks the list you pass. For each one it computes the document hash and looks up the hash the docstore recorded for that id. Unchanged means skip — no embedding call, no write. Changed means delete the old nodes and insert the new ones. Absent means insert. The return value is a list of booleans in the same order, so you can log or assert how much actually moved; a refresh that reports all `True` on an unchanged corpus is a sign your ids or your text normalisation are unstable. Because the comparison is over the document's content hash, incidental churn matters. If your loader injects a timestamp into metadata that participates in the hash, every document looks changed on every run and you re-embed the entire corpus nightly while believing you built an incremental pipeline. ## Deletions are your responsibility Nothing in `refresh_ref_docs` can know that a file was removed — it only sees what you pass. The pattern is to keep the set of source ids from the previous run, diff it against the current set, and call `index.delete_ref_doc(doc_id, delete_from_docstore=True)` for each id that vanished. The `delete_from_docstore` argument defaults to `False`. With the default, the nodes go but the document record and its hash remain in the docstore. Re-adding that source later can then be skipped by a hash comparison against a record whose nodes no longer exist — an index that reports nothing to do while returning nothing for that document. Passing `True` is almost always what you want. ## The no-docstore trap All of this needs a docstore, because the hashes and the document-to-node mapping live there. An index created with `VectorStoreIndex.from_vector_store(vector_store)` has no docstore: the integration stores node text alongside the vectors (`stores_text=True`), which is enough to answer queries but carries no document hashes. Refresh has nothing to compare against. If your architecture is a stateless service attaching to a shared remote store, either persist and share the docstore alongside it, or move the sync problem out of LlamaIndex entirely and manage upserts and deletes against the vector store by your own stable ids. ## Operating shape In production, treat refresh as a job with observability: count how many documents reported changed, alert when that number is zero for a corpus you know is changing or equal to the corpus size for one you know is not. Run the delete diff in the same job so additions and removals cannot drift apart. And accept that some changes cannot be patched incrementally at all — a chunking-parameter change or an embedding-model change invalidates every vector regardless of whether the source text moved, and the honest response is a full rebuild rather than a refresh.
- Why does delete_ref_doc default to leaving the document in the docstore?The default `delete_from_docstore=False` removes the nodes but keeps the document record and its hash, which is only sensible if you intend to keep the record for bookkeeping. In practice it creates a ghost: a later `refresh_ref_docs` compares against a hash whose nodes no longer exist, decides nothing changed, and leaves that source unretrievable. Pass `True` unless you have a specific reason not to.
- Your nightly refresh re-embeds the entire corpus even though almost nothing changed. What do you check?First the document ids — if the loader assigns fresh UUIDs each run, every document is new. Then the hash inputs: metadata that changes every run, such as an ingestion timestamp or a mutable file attribute, makes identical text hash differently. Normalise metadata and derive ids from the source identity, then verify with the boolean list `refresh_ref_docs` returns.
- When is refresh the wrong tool entirely?When the invalidation is global rather than per document. Changing the embedding model, or changing chunk size so node boundaries move, invalidates every vector even though no source text changed — hashes still match, so refresh does nothing and the index quietly stays inconsistent. Those changes call for a full rebuild into a new index artifact, cut over once it is verified.
saying these in an interview costs you the question
- Lets the loader assign random document ids and expects refresh to work
- Assumes deleted source files disappear from the index automatically
- Calls delete_ref_doc without delete_from_docstore and considers the document gone
- Tries to refresh an index attached with from_vector_store and no docstore
- Thinks changing chunk size or embedding model can be handled incrementally