skip to content

Indexes & Vector Stores

You will learn the index types — VectorStoreIndex, SummaryIndex, KnowledgeGraphIndex and TreeIndex — along with vector store integrations such as Pinecone, Chroma and Weaviate, index construction, persistence via the storage context, and refresh. Interviewers ask because picking an index type and keeping it in sync with changing source data are design decisions, not defaults.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

questions

6

In LlamaIndex, what does VectorStoreIndex.from_documents() actually do?

level: juniorimportance: must knowfreq 78%

answer

  1. one pass over the data at build time
  2. which model gets called, and which does not
  3. chunk first, then embed
  4. where the vectors sit by default
  5. nothing on disk until you say so

basics

~20 s

VectorStoreIndex.from_documents() splits each document into nodes, calls the embedding model once per node, and writes those vectors plus the node text into a vector store — by default an in-memory SimpleVectorStore that vanishes when the process exits.

solid answer

~40 s

`VectorStoreIndex.from_documents(documents)` runs the whole build pipeline in one call. Each `Document` goes through the configured node parser to produce `TextNode` chunks, every node is sent to the embedding model, and the resulting vector plus the node's text and metadata are written into a vector store. If you pass no `storage_context`, that store is an in-memory `SimpleVectorStore`, and the docstore and index store are in memory too — nothing touches disk until you call `persist()`. The build cost is one embedding call per node (batched), not an LLM call: unlike `TreeIndex` or `PropertyGraphIndex`, a `VectorStoreIndex` makes no LLM calls at construction. The arguments you actually reach for are `show_progress=True`, `embed_model=` to override the global `Settings.embed_model`, and `storage_context=` to point the build at a real store.

code

python · 4 lines
python
from llama_index.core import Document, VectorStoreIndex

documents = [Document(text="LlamaIndex embeds each node at build time.")]
index = VectorStoreIndex.from_documents(documents, show_progress=True)

go deeper

for a junior

Be able to say the three steps out loud: split into nodes, embed each node, store the vectors. Know that the default store is in memory and that nothing is saved unless you persist it.

for a middle

Explain why the embedding model — not the LLM — is what runs at build time, and quantify the cost as roughly one embedding per node. Mention that you can skip parsing by passing nodes straight to the VectorStoreIndex constructor.

for a senior

Talk about build economics on a real corpus: batching, re-embedding on every rebuild, and why teams persist early. Be ready to say when you would build from nodes produced by your own pipeline instead of from_documents().

for a principal

Own the decision of what gets embedded at all. Argue about corpus scope, embedding-model choice and its migration cost, and whether rebuilds should be a batch job with its own budget rather than something a service does at startup.

## What the call is `VectorStoreIndex` is LlamaIndex's default index type: a flat collection of embedded chunks you later search by vector similarity. `from_documents()` is the convenience constructor that turns raw `Document` objects into a queryable index in one line. ## The three stages **1. Documents become nodes.** A `Document` holds a blob of text plus metadata. Retrieval does not operate on whole documents — it operates on chunks, which LlamaIndex calls `TextNode` objects. `from_documents()` runs the configured node parser (taken from `Settings` unless you override it) to split each document, and every node keeps a `ref_doc_id` pointing back at the document it came from. That back-pointer is what later makes updates and deletes possible. **2. Nodes become vectors.** Each node's text is sent to the embedding model — `Settings.embed_model` unless you pass `embed_model=` explicitly. Embedding calls are batched, so a 10,000-node corpus is not 10,000 round trips, but it is still 10,000 pieces of text billed and processed. This is the dominant cost and the dominant latency of the build. `show_progress=True` prints a progress bar so a long build does not look hung. **3. Vectors land in a store.** The `(id, vector, text, metadata)` tuples go into the vector store held by the index's `StorageContext`. With no `storage_context` argument, LlamaIndex builds a default one whose vector store is `SimpleVectorStore` — a plain in-memory dict with brute-force similarity search — plus an in-memory docstore and index store. Everything lives in the Python process and is lost on exit unless you persist it. ## What it does NOT do It does not call the LLM. People assume indexing means "the model reads my documents"; for `VectorStoreIndex` no generative model is involved, only the embedding model. Index types that summarise at build time — `TreeIndex`, `DocumentSummaryIndex`, `PropertyGraphIndex` — are the ones that spend LLM tokens up front. It does not deduplicate. Calling `from_documents()` twice with the same input gives you two independent indexes; feeding the same document into one index twice gives duplicate nodes. Incremental updates are what `insert()`, `insert_nodes()`, `update_ref_doc()` and `refresh_ref_docs()` are for. It does not persist. `persist()` is a separate, explicit call. ## Constructing without the helper `from_documents()` is sugar. If you already have nodes — because you ran your own splitting or transformation step — construct the index directly as `VectorStoreIndex(nodes)`; the embedding and storage stages are identical, only the parsing stage is skipped. And if the vectors already exist in an external store from a previous run, `VectorStoreIndex.from_vector_store(vector_store)` attaches an index object to them without re-embedding anything. ## Cost control in practice Because every node is embedded, the two knobs that actually move the bill are how many nodes you produce and which embedding model you use. Re-running `from_documents()` on an unchanged corpus re-embeds all of it, which is why teams that iterate on prompts and retrieval settings persist the index early and reload it rather than rebuilding on every script run. Building against a remote store also means the write path is network-bound, so a large first build is often store-throughput-limited rather than embedding-limited.

  • If you never set embed_model, which model performs the embedding?
    The one on the global `Settings.embed_model`. LlamaIndex falls back to an OpenAI embedding model unless you assign a different one, which is why a build can fail on a missing API key even though you never wrote any model code. Setting `Settings.embed_model` once at startup, or passing `embed_model=` to the constructor, makes the choice explicit and keeps the same model in use when the index is later reloaded.
  • What happens if you call from_documents() twice with the same documents?
    You get two separate indexes, each with its own full set of embeddings — there is no shared cache and no deduplication. If both builds target the same external vector store, you end up with duplicate vectors for the same content, which silently degrades retrieval because near-identical chunks crowd out the top results. Use `insert()` or `refresh_ref_docs()` against an existing index instead of rebuilding.
  • Does building a VectorStoreIndex ever call the LLM?
    No. Construction only touches the embedding model. The LLM enters at query time, when a response synthesizer turns retrieved nodes into an answer. Other index types differ: `TreeIndex` and `DocumentSummaryIndex` spend LLM calls during construction to write summaries, and `PropertyGraphIndex` uses the LLM to extract triples, so their build cost is an LLM bill rather than an embedding bill.

saying these in an interview costs you the question

  • Thinks the LLM reads and summarises documents during VectorStoreIndex construction
  • Assumes the index is written to disk automatically without persist()
  • Believes embeddings are computed lazily at query time
  • Expects repeated from_documents() calls to deduplicate existing content
  • Thinks from_documents() stores whole documents rather than chunked nodes

context

open as a page

When would you choose SummaryIndex or TreeIndex over VectorStoreIndex in LlamaIndex?

level: middleimportance: must knowfreq 70%

basics

~20 s

Choose by the shape of the question. VectorStoreIndex suits targeted lookups over a large corpus. SummaryIndex keeps every node in order and is for whole-document synthesis over a small set. TreeIndex spends LLM calls at build to create a summary hierarchy for large-scope questions.

open as a page

How do you persist a LlamaIndex index and reload it without re-embedding?

level: middleimportance: must knowfreq 72%

basics

~10 s

Call index.storage_context.persist(persist_dir="./storage") to write the docstore, index store and default vector store to disk, then rebuild with StorageContext.from_defaults(persist_dir="./storage") and load_index_from_storage(storage_context). Re-calling from_documents() would re-embed everything.

open as a page

What does building a PropertyGraphIndex in LlamaIndex cost versus a VectorStoreIndex?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A VectorStoreIndex build spends one embedding call per node. A PropertyGraphIndex additionally runs LLM extractors over every chunk to pull out entities and relations, so the build is an LLM bill, is far slower, and produces non-deterministic structure that changes between runs.

open as a page

How do you keep a LlamaIndex index in sync when source documents change or are deleted?

level: seniorimportance: should knowfreq 52%

basics

~10 s

Give every document a stable id, then call index.refresh_ref_docs(documents): LlamaIndex compares each document's hash against the docstore and re-inserts only what changed. Removals need an explicit index.delete_ref_doc(doc_id, delete_from_docstore=True).

open as a page

When should a LlamaIndex app move off SimpleVectorStore to Chroma, Pinecone or Weaviate?

level: principalimportance: should knowfreq 38%

basics

~20 s

Move when the index stops fitting one process: the corpus exceeds memory, several replicas must share it, writes must be incremental rather than a whole-file rewrite, or durability and concurrent access become requirements. Until then SimpleVectorStore is faster to iterate on.

open as a page