skip to content

What does building a PropertyGraphIndex in LlamaIndex cost versus a VectorStoreIndex?

level: seniorimportance: should knowfreq 40%

answer

  1. the build is inference, not bookkeeping
  2. one model call per chunk, at least
  3. same input, different graph each run
  4. a schema is the tame-it lever
  5. in-memory graph store is a demo default

basics

~20 s

A VectorStoreIndex build spends one embedding call per node. A PropertyGraphIndex additionally runs LLM extractors over every chunk to pull out entities and relations, so the build is an LLM bill, is far slower, and produces non-deterministic structure that changes between runs.

solid answer

~50 s

`PropertyGraphIndex.from_documents(documents)` runs a list of `kg_extractors` over each chunk. The default set includes `SimpleLLMPathExtractor`, which prompts the LLM to emit subject-predicate-object triples, and `ImplicitPathExtractor`, which is free because it just reads the node relationships already present. So the build cost is roughly one LLM call per chunk on top of embeddings, which is one to two orders of magnitude more expensive and slower than a vector build. You also get variance: two runs over identical text yield different triples, so the graph is not reproducible. `SchemaLLMPathExtractor` constrains extraction to declared entity and relation types, which trades recall for consistency and is what makes graph indexes operable. Triples land in a property graph store — `SimplePropertyGraphStore` in memory by default, or a graph database integration for anything real. It is worth it when questions are multi-hop or entity-centric, and wasteful when similarity search already answers them.

code

python · 9 lines
python
from llama_index.core import Document, PropertyGraphIndex
from llama_index.core.indices.property_graph import SimpleLLMPathExtractor

documents = [Document(text="Acme acquired Initech in 2021.")]
index = PropertyGraphIndex.from_documents(
    documents,
    kg_extractors=[SimpleLLMPathExtractor(max_paths_per_chunk=10)],
    show_progress=True,
)

go deeper

for a junior

Know that LlamaIndex can build a graph index and that it uses the LLM to pull entities and relations out of text, unlike a vector index which only embeds.

for a middle

Explain the kg_extractors list and which entries cost model calls, and be able to contrast the build cost with one embedding per node. Mention that the default graph store is in memory.

for a senior

Talk about non-determinism and entity resolution, why SchemaLLMPathExtractor is the practical lever, and what a changed document does to graph consistency. Be clear that KnowledgeGraphIndex is the deprecated predecessor.

for a principal

Own the build-cost decision: what fraction of the corpus justifies extraction, how the recurring bill is budgeted, whether a graph database becomes a system you now operate, and what question types would otherwise go unanswered.

## What is actually being built A `VectorStoreIndex` build is mechanical: chunk, embed, store. A `PropertyGraphIndex` build is an inference task. Each chunk is fed to extractors that produce labelled nodes (entities) and relations between them, and the resulting graph is stored alongside optional embeddings of the graph nodes. The index therefore encodes claims the model made about your text, not just the text itself. ## The extractor list is the cost model `kg_extractors` is a list, and each entry has a very different price: - `ImplicitPathExtractor` reads relationships already recorded on nodes — previous/next, parent/child — and costs nothing. - `SimpleLLMPathExtractor` prompts the LLM per chunk for free-form triples, bounded by `max_paths_per_chunk`. This is the expensive default. - `SchemaLLMPathExtractor` does the same but constrains output to a declared set of possible entity and relation types, so the graph has a vocabulary you can query against. Cost scales with chunk count times extractor count. A corpus that embeds in a couple of minutes can take hours and a substantial model bill to extract, and every rebuild repeats it. ## Non-determinism is the operational headache Because extraction is generative, the same document processed twice gives different triples: different surface forms for the same entity, relations phrased differently, occasional hallucinated links. Two consequences follow. First, the graph does not converge — entity resolution is on you, or you accept duplicates like "Acme Corp" and "Acme Corporation" as distinct nodes. Second, incremental updates are messy: re-extracting a changed document produces triples that may not line up with the ones its neighbours contributed, so the graph drifts. `SchemaLLMPathExtractor` is the main lever against this, because a closed vocabulary of entity and relation types removes most of the surface-form variance at the cost of missing anything outside the schema. ## Storage The default `SimplePropertyGraphStore` keeps the graph in memory and serialises with the rest of the storage context. That is fine for demos and hopeless at scale, because traversal over a large graph in Python is slow and the whole graph must fit in the process. Real deployments point `property_graph_store` at a graph database integration, which moves traversal into a system designed for it and lets several services share one graph. `embed_kg_nodes` controls whether graph nodes also get embeddings, which is what allows finding an entry point into the graph by similarity before traversing structurally. ## KnowledgeGraphIndex is the predecessor Older material and much training data describes `KnowledgeGraphIndex` with `max_triplets_per_chunk` and `include_embeddings`. In current llama-index-core that class is the deprecated ancestor; `PropertyGraphIndex` supersedes it with typed nodes and relations, pluggable extractors, and real graph-store integrations. Naming it as the current API in an interview dates you, so say plainly that it is legacy. ## When the cost is justified Graph indexes earn their price on questions vector search structurally cannot answer: multi-hop connections ("which suppliers are two steps removed from this sanctioned entity"), aggregation over entities, and traversal where the relevant text chunks share no vocabulary with the question. They lose on ordinary lookup questions, where a vector index answers just as well for a fraction of the build cost. The honest engineering answer is usually to keep a vector index as the workhorse and add a graph index over the narrow, entity-dense subset where hops are the actual requirement — and to budget the extraction pass as a recurring cost, not a one-off.

  • How does SchemaLLMPathExtractor change the tradeoff?
    It constrains extraction to a declared set of entity and relation types instead of free-form triples, so the graph gains a stable vocabulary you can actually write queries against and duplicate surface forms drop sharply. The cost is recall: anything the schema does not anticipate is dropped rather than captured. That is usually the right trade for a production graph, because an unqueryable graph of inconsistent labels has no value however complete it is.
  • Why does the default SimplePropertyGraphStore stop working at scale?
    It holds the entire graph in process memory and traverses it in Python, so both footprint and traversal latency grow with the graph and nothing is shared between replicas. A graph database integration pushes traversal into a purpose-built engine, gives you persistence and concurrent access, and lets several services query one graph. Swapping it is a constructor argument, but the migration also means owning schema and operations for another datastore.
  • Should a changed document trigger a full graph rebuild?
    Ideally no, but partial re-extraction is genuinely awkward: the new triples for that document come from a fresh generation and may use different surface forms than the untouched neighbours contributed, so the graph drifts toward inconsistency. A declared schema mitigates most of this. Where consistency matters more than freshness, teams rebuild the graph on a schedule and treat it as a batch artifact rather than something updated per document.

saying these in an interview costs you the question

  • Assumes graph index construction costs about the same as embedding
  • Expects two builds over identical documents to produce identical graphs
  • Cites KnowledgeGraphIndex as the current API rather than the deprecated one
  • Ships the in-memory graph store to production
  • Reaches for a graph index for ordinary single-hop lookup questions

context