skip to content

What does LlamaIndex's HierarchicalNodeParser produce, and what must you store?

level: seniorimportance: should knowfreq 34%

answer

  1. one document, several chunk sizes
  2. precision to match, size to read
  3. leaves get embedded, all get stored
  4. parent links are ids, not nodes
  5. docstore must be persisted

basics

~20 s

It chunks each document several times at decreasing sizes and links the levels with PARENT and CHILD relationships. You embed only the leaf nodes returned by get_leaf_nodes, but every node of every level must go into the docstore or the parent links dangle.

solid answer

~50 s

`HierarchicalNodeParser.from_defaults(chunk_sizes=[2048, 512, 128])` splits each document at every size in the list and wires the resulting nodes together with `NodeRelationship.PARENT` and `CHILD`, so a 128-token leaf knows its 512-token parent and that its 2048-token grandparent. The intended usage is asymmetric: index only `get_leaf_nodes(nodes)` in the vector store so the embedded unit is small and precise, but call `storage_context.docstore.add_documents(nodes)` with the **full** list so the coarser levels can be resolved by id later. Retrieval then matches leaves and can trade them up for their parent when enough siblings hit, giving the LLM a coherent passage instead of a handful of fragments. The costs are a bigger docstore, an ingest that runs the splitter once per level, and node ids that all change on re-ingest — so the docstore has to be persisted and versioned alongside the vector index.

code

python · 10 lines
python
from llama_index.core import StorageContext, VectorStoreIndex
from llama_index.core.node_parser import HierarchicalNodeParser, get_leaf_nodes

parser = HierarchicalNodeParser.from_defaults(chunk_sizes=[2048, 512, 128])
nodes = parser.get_nodes_from_documents(documents)

storage_context = StorageContext.from_defaults()
storage_context.docstore.add_documents(nodes)          # every level
index = VectorStoreIndex(get_leaf_nodes(nodes), storage_context=storage_context)
storage_context.persist(persist_dir="./storage")

go deeper

for a junior

Know that this parser chunks the same document at several sizes and records which small chunk sits inside which larger one, so retrieval can later show more context than it matched on.

for a middle

Explain the asymmetry: get_leaf_nodes goes into the vector index, the complete node list goes into the docstore, and parent relationships are ids that need that docstore to resolve.

for a senior

Show the operational side — persisting and versioning the docstore with the vector store, handling id churn on re-ingest, and judging when the storage and ingest cost is not worth it for short documents.

for a principal

Frame it as a two-store consistency problem: the hierarchy couples a vector database and a document store that must be refreshed, backed up and restored as one unit, and that coupling is the real cost you are signing up for.

## The problem it solves Flat chunking forces one size to serve two different jobs. Small chunks embed precisely — the vector represents one idea, so it matches a specific query — but they are lousy context, because the answer usually needs more than the matched sentence. Large chunks are good context but muddy vectors: several topics in one embedding means nothing matches strongly. Overlap only blurs the boundary; it does not resolve the conflict. `HierarchicalNodeParser` resolves it by refusing to pick one size. ## What it produces `HierarchicalNodeParser.from_defaults(chunk_sizes=[2048, 512, 128])` splits each document three times, once per entry in the list, each level using sentence-based splitting. It returns one flat Python list containing nodes from **all** levels, wired together by relationships: - each 512-token node has `PARENT` pointing at the 2048-token node it fell inside, - each 2048-token node has `CHILD` entries listing its 512-token pieces, - the same holds between 512 and 128, - and all of them keep `SOURCE` back to the original document. Helpers exported from `llama_index.core.node_parser` slice that flat list: `get_leaf_nodes(nodes)` returns the finest level, `get_root_nodes(nodes)` the coarsest. ## The asymmetric storage rule This is the part candidates get wrong, and it is the whole question: - **Vector index: leaves only.** `VectorStoreIndex(get_leaf_nodes(nodes), storage_context=storage_context)`. Embedding every level would triple the embedding bill and pollute results, because a query would match a leaf and its parent and its grandparent — three hits carrying largely the same text. - **Docstore: everything.** `storage_context.docstore.add_documents(nodes)` with the complete list. Parent ids stored on a leaf are just ids; without the coarser nodes in a lookup store, nothing can trade a leaf up for its parent, and the hierarchy is dead weight. Forget the second line and the pipeline appears to work — retrieval returns leaves, answers are just a bit thin — which is exactly why it is a good interview probe. ## What consumes the hierarchy At query time the point is to retrieve fine and read coarse: match on leaves, then substitute the parent when a sufficient share of that parent's children were retrieved. LlamaIndex ships a retriever for exactly this pattern, and it needs the storage context precisely because it resolves parent ids through the docstore. The effect is that a query hitting three adjacent leaves yields one coherent 512-token passage rather than three fragments the LLM has to stitch together, and it de-duplicates redundant sibling hits for free. ## Costs and operational consequences - **Ingest time** scales with the number of levels: the splitter runs once per entry in `chunk_sizes`. - **Docstore size** is roughly the corpus text stored once per level — three levels means the text is held three times over, even though only one level is embedded. - **Persistence is mandatory.** The docstore must be persisted (`storage_context.persist(...)`) and loaded back, and it must stay in sync with the vector store. A vector index restored from a managed vector database while the docstore is rebuilt from scratch produces dangling parent ids. - **Re-ingestion churns ids.** Reparsing a document generates new node ids at every level, so a partial refresh has to remove the old family, not merely add the new one, or the docstore accumulates orphans. - **Level choice matters.** The default `[2048, 512, 128]` roughly triples each step. Levels that are too close together produce parents barely larger than their children, which buys nothing for the storage cost. ## When not to use it If your documents are short — support tickets, product descriptions, FAQ entries — a single level plus neighbour expansion through `PREVIOUS`/`NEXT` gives most of the benefit at a fraction of the storage and operational complexity. The hierarchy pays off on long, structured documents: manuals, contracts, standards, research papers, where a leaf genuinely needs a section-sized parent to be interpretable. ## The one-line summary to give Multi-level chunks linked by parent/child, leaves in the vector index, every level in a persisted docstore, so that retrieval can be precise while the context handed to the model stays coherent.

  • Why not just embed every level instead of only the leaves?
    Because the levels contain the same text at different granularities. Embedding all three multiplies the embedding bill and the index size, and a query then matches a leaf, its parent and its grandparent — three near-duplicate results consuming your top-k with one passage. Embedding leaves keeps vectors precise and distinct; the coarser levels are needed only for lookup by id, which the docstore provides far more cheaply.
  • Your parent lookups return nothing after a restart. What did you forget?
    Persisting and reloading the docstore. Parent relationships hold node ids, not node objects, so if the vector store survives the restart but the docstore was in memory, every parent id points at nothing. `storage_context.persist(persist_dir=...)` at ingest and loading the same storage context at query time is the fix, and the two stores must be kept in sync as a unit.
  • How do you choose the chunk_sizes list?
    Aim for a clear ratio between levels — the default 2048/512/128 is roughly 4x per step — so a parent is meaningfully more context than its children. Levels that are too close deliver almost no extra context for the storage they cost. Set the leaf size by what embeds precisely for your content and the root by what you are willing to put in a prompt, then fill the middle.
  • When is a flat parser plus neighbour expansion the better call?
    When documents are short or unstructured — tickets, FAQ entries, chat logs — a single chunk level with PREVIOUS/NEXT expansion recovers most of the context benefit without storing the corpus several times over or coupling two stores that must be refreshed together. The hierarchy earns its complexity on long structured documents such as manuals, contracts and papers.

saying these in an interview costs you the question

  • Embeds every level into the vector index
  • Adds only leaf nodes to the docstore
  • Thinks parent nodes are stored inside their children
  • Skips persisting the docstore alongside the vector store
  • Picks chunk sizes so close together that parents add no context

context