When should a LlamaIndex app move off SimpleVectorStore to Chroma, Pinecone or Weaviate?
answer
- default trades scale for zero setup
- the trigger is structural, not aesthetic
- one process versus many replicas
- wiring is easy, operating is not
- exact scan versus approximate index
basics
~20 sMove when the index stops fitting one process: the corpus exceeds memory, several replicas must share it, writes must be incremental rather than a whole-file rewrite, or durability and concurrent access become requirements. Until then SimpleVectorStore is faster to iterate on.
solid answer
~50 s`SimpleVectorStore` is the default because it removes a dependency: vectors sit in a Python dict, search is a brute-force scan, and persistence is a JSON file inside the storage context. That is genuinely fine for prototypes and small, stable corpora. The signals to move are structural, not aesthetic — the corpus no longer fits in process memory, more than one replica needs the same index and you do not want each rebuilding it, updates must be incremental instead of rewriting a whole file, or you need durability and concurrent readers and writers. Moving is a wiring change on the LlamaIndex side: construct the store, pass it via `StorageContext.from_defaults(vector_store=...)`, and reattach later with `VectorStoreIndex.from_vector_store()`. The real cost is what comes after — you now own a datastore's lifecycle, collection naming, and the fact that filter and hybrid capabilities differ between integrations, so a switch is not free.
code
python · 11 linesimport chromadb
from llama_index.core import Document, StorageContext, VectorStoreIndex
from llama_index.vector_stores.chroma import ChromaVectorStore
client = chromadb.PersistentClient(path="./chroma")
store = ChromaVectorStore(chroma_collection=client.get_or_create_collection("docs"))
storage_context = StorageContext.from_defaults(vector_store=store)
index = VectorStoreIndex.from_documents(
[Document(text="hello")], storage_context=storage_context
)go deeper
Know that LlamaIndex defaults to an in-memory SimpleVectorStore and that a real deployment usually wires a vector database in through a StorageContext instead.
Name the concrete limits of the default — process memory, whole-file persistence, no sharing — and show the storage_context wiring plus the from_vector_store reattachment that replaces it.
Reason about what does not port: approximate versus exact search, filter capability differences, and the loss of refresh and delete-by-document when there is no docstore. Re-evaluate retrieval quality after the move.
Own the operating commitment. State the trigger you would act on, weigh managed versus self-hosted versus embedded against what your team can run, and keep the choice reversible by keeping backend-specific features out of the application.
## What the default actually is `SimpleVectorStore` holds embeddings in memory and compares the query vector against every stored vector. There is no index structure and no approximation, so results are exact — which is a genuine advantage while you are still tuning chunking and prompts, because retrieval quality is not confounded by an approximate search. It serialises to JSON with the rest of the storage context, which means a save rewrites the file rather than appending. ## The four signals to move **Memory.** The whole index lives in the process. Once vectors plus node text no longer fit comfortably alongside the application, you are choosing between a bigger instance forever and a store designed to hold this data. **Sharing.** With several replicas, each one either loads its own copy of the index at startup — slow boots, duplicated memory, and a window where replicas disagree after a rebuild — or they all point at one store. The second is the only shape that scales operationally. **Write pattern.** Rewriting a JSON file on every change is fine for a nightly rebuild and untenable for continuous ingestion. External stores accept upserts and deletes per record. **Durability and concurrency.** In-memory means a crash loses everything since the last persist, and nothing coordinates a writer against readers. Once freshness has an SLA, that is not acceptable. Search latency is often assumed to be the trigger, and it usually is not the first one — a brute-force scan over a modest corpus is quick. It becomes the trigger once the corpus is large enough that a linear scan per query dominates response time. ## What changes in the code Surprisingly little. You construct the integration's store object, wrap it in a storage context, and build as usual: ``` storage_context = StorageContext.from_defaults(vector_store=store) index = VectorStoreIndex.from_documents(documents, storage_context=storage_context) ``` Later processes attach with `VectorStoreIndex.from_vector_store(store)` and skip the build entirely. Retrieval code is unchanged, which is the point of the abstraction. ## What does not port cleanly The abstraction is not perfectly uniform, and pretending otherwise is what makes migrations painful: - **Metadata filtering.** All integrations accept filters, but the supported operators and how they perform differ by backend, so a filter-heavy retrieval design tuned on one store can behave differently on another. - **Hybrid and sparse search.** Some backends support it natively and expose it through their integration; others do not, so a design that depends on it constrains your choice of store. - **Docstore semantics.** Most integrations keep node text alongside the vector, so `from_vector_store()` alone answers queries. But without a docstore you lose document hashes, and with them refresh-by-hash and delete-by-document. Decide early whether the docstore is persisted and shared too. - **Approximate search.** External stores generally use an ANN index, so results are approximate and tunable. Quality differences you attribute to your chunking may in fact be recall differences from index parameters. ## Framing the decision at a lead level The question behind the question is what you are willing to operate. A managed service converts the problem into a bill and a vendor dependency; a self-hosted store converts it into capacity, backups and upgrades your team owns; an embedded store keeps the data local and single-writer, which is a real fit for a single-node service and a dead end for a fleet. The strongest answer names the trigger you would act on, states that the retrieval code is portable while the operational commitment is not, and keeps the choice reversible by not letting store-specific features leak into the application until they have earned it.
- What changes in retrieval quality when you move to an external store?Search usually stops being exact. `SimpleVectorStore` scans every vector, while most integrations use an approximate nearest-neighbour index with tunable parameters, so recall becomes a setting rather than a guarantee. Results can shift even with identical embeddings and chunking. Re-run your retrieval evaluation after the migration rather than assuming parity, or you will misattribute the difference to your pipeline.
- After moving to an external store, do you still need the docstore?Only if you want the operations that depend on document hashes — `refresh_ref_docs` and delete-by-document. Most integrations store node text with the vector, so queries work from `from_vector_store()` alone. If you skip the docstore, sync moves out of LlamaIndex: you own upserts and deletes against the store by your own stable ids. Decide this deliberately rather than discovering it when the first document changes.
- How do you keep the choice of store reversible?Keep store-specific behaviour out of the application. Build and reattach through `StorageContext` and `from_vector_store()`, express filters in LlamaIndex's own metadata filter objects rather than backend query syntax, and avoid depending on a native capability until it has earned the lock-in. Then also keep the index reproducible from source, so switching backends is a rebuild rather than a data migration.
saying these in an interview costs you the question
- Reaches for a hosted vector database on day one for a few hundred documents
- Assumes swapping stores is transparent because the retrieval code is unchanged
- Believes exact and approximate search return the same results
- Forgets that dropping the docstore removes refresh and delete-by-document
- Treats the decision as a benchmark exercise rather than an operating commitment