How do you persist a LlamaIndex index and reload it without re-embedding?
answer
- an index is more than its vectors
- two calls to get it back, not one
- name your index if there are several
- the remote store already has the data
- same embedding model on both sides
basics
~10 sCall index.storage_context.persist(persist_dir="./storage") to write the docstore, index store and default vector store to disk, then rebuild with StorageContext.from_defaults(persist_dir="./storage") and load_index_from_storage(storage_context). Re-calling from_documents() would re-embed everything.
solid answer
~40 sAn index is not just vectors — it is a `StorageContext` holding a docstore (the nodes), an index store (the index's own structure and id), and a vector store. `index.storage_context.persist(persist_dir="./storage")` writes those to JSON files in that directory. To reload, build a storage context pointing at the same directory with `StorageContext.from_defaults(persist_dir="./storage")` and pass it to `load_index_from_storage(storage_context)`; the vectors come back as stored, so no embedding calls are made. If the directory holds more than one index, give each one an id with `index.set_index_id("faq")` and pass `index_id="faq"` on load, otherwise the load is ambiguous. The important caveat: with an external vector store such as Chroma, Pinecone or Weaviate, the vectors were never in the local files — that store owns them, and you reattach with `VectorStoreIndex.from_vector_store(vector_store)`.
code
python · 13 linesfrom llama_index.core import (
Document,
StorageContext,
VectorStoreIndex,
load_index_from_storage,
)
index = VectorStoreIndex.from_documents([Document(text="hello")])
index.set_index_id("faq")
index.storage_context.persist(persist_dir="./storage")
storage_context = StorageContext.from_defaults(persist_dir="./storage")
reloaded = load_index_from_storage(storage_context, index_id="faq")go deeper
Remember the two calls that bring an index back: StorageContext.from_defaults(persist_dir=...) then load_index_from_storage(). Know that persist() is explicit and nothing is saved without it.
Explain what lives in the docstore, the index store and the vector store, and why reloading avoids the embedding bill. Be able to name set_index_id and the index_id argument for multi-index directories.
Handle the external-store case cleanly: what persist() actually wrote, why from_vector_store is the reattachment path, and what you lose when there is no docstore. Talk about pinning the embedding model to the artifact.
Treat the index as a versioned build artifact with an owner and a rebuild pipeline. Decide where it lives, how a model change forces regeneration, and how services get a consistent index without racing each other to rebuild it.
## The storage context is the real unit of persistence People think of a LlamaIndex index as "the vectors". It is actually an object over a `StorageContext` with several independent stores: - **docstore** — the `TextNode` objects themselves: text, metadata, relationships, and the hash of each source document. - **index store** — the index's own metadata: its type, its id, and for structured index types the structure (a tree's parent/child links, a keyword table's mapping). - **vector store** — the embeddings. - **graph store / property graph store** — used by graph index types. Persisting means serialising whichever of these are local. ## The default, all-local flow ``` index.storage_context.persist(persist_dir="./storage") ``` writes JSON files into that directory — `docstore.json`, `index_store.json`, and for the in-memory default store `default__vector_store.json`. Reloading is two steps, and the two-step shape is what candidates get wrong: ``` storage_context = StorageContext.from_defaults(persist_dir="./storage") index = load_index_from_storage(storage_context) ``` `from_defaults(persist_dir=...)` reconstitutes the stores from the files; `load_index_from_storage()` reads the index store to work out which index type to instantiate and hands you a live index object. No embedding model is called, which is the whole point: rebuilding with `from_documents()` would re-embed every node and re-bill you for data you already have. ## Multiple indexes in one directory A storage context can hold several indexes. Each index has an id, and if you never set one it is a random UUID. Call `index.set_index_id("faq")` before persisting, then load with `load_index_from_storage(storage_context, index_id="faq")`. With more than one index present and no id supplied, the load cannot decide and fails — an error people meet the first time they persist two indexes into the same folder. ## The external-vector-store case This is the part interviewers push on. When you built the index against a real vector store — wired in via `StorageContext.from_defaults(vector_store=...)` — the embeddings went to that system, not to your process. `persist()` then writes only the docstore and index store; there is nothing local to write for the vectors. Reloading with `load_index_from_storage()` against a bare `persist_dir` gives you an index whose vector store is the empty local default, and queries return nothing. The correct reattachment is: ``` index = VectorStoreIndex.from_vector_store(vector_store) ``` which constructs an index directly over the existing remote data with no local files at all. This works because most integrations set `stores_text=True`: they keep the node text alongside the vector, so retrieval can return whole nodes without consulting a docstore. That convenience has a consequence — with no docstore, the operations that need document hashes and node-to-document links, notably refresh and delete-by-document, have nothing to work from. If you want those, persist the docstore too and pass both the vector store and the persisted directory when constructing the storage context. ## Practical shape A typical service does: try to load; on failure, build and persist. Keep the embedding model identical across build and load — vectors are only comparable within the model that produced them, and reloading an index while a different `Settings.embed_model` is active silently degrades every query rather than raising. Treat the persist directory as a build artifact, versioned with the embedding model name, so swapping models forces a visible rebuild rather than a quiet mismatch.
- You persisted an index built on Pinecone, then load_index_from_storage() returns nothing. Why?Because the vectors were never in the persist directory — the remote store owns them, and `persist()` only wrote the docstore and index store. Loading from a bare directory gives you an index whose vector store is the empty local default. Reattach with `VectorStoreIndex.from_vector_store(vector_store)` after constructing the store client, optionally combining it with the persisted docstore if you also need refresh and delete-by-document to work.
- What breaks if you reload an index under a different embedding model?Nothing raises, and that is the danger. The stored vectors came from the old model; the query is embedded with the new one, so the two live in incompatible spaces and similarity scores become meaningless. Retrieval quietly returns poor results. Pin the embedding model alongside the index artifact and rebuild whenever it changes — treat the model identifier as part of the index's version.
- How do you keep two indexes in one persist directory?Call `index.set_index_id("name")` on each before persisting into the shared directory, then load with `load_index_from_storage(storage_context, index_id="name")`. Both indexes share the docstore, so nodes are stored once even though two index structures reference them. Without explicit ids each index gets a random UUID and an id-less load cannot choose between them.
saying these in an interview costs you the question
- Thinks re-running from_documents() is how you reload a saved index
- Assumes persist() also uploads or downloads a remote vector store's data
- Forgets StorageContext.from_defaults and calls load_index_from_storage on a path
- Believes swapping the embedding model after a load is harmless
- Says the index is only the vectors, ignoring the docstore and index store