What are NodeRelationship.PREVIOUS and NEXT for on LlamaIndex nodes?
answer
- links, not copies of neighbours
- chunking loses reading order
- expand context at read time
- ids need a docstore to resolve
- chain stops at document edges
basics
~20 sThey are pointers stored in each node's relationships dict that link it to the chunks immediately before and after it in the original document. They let you expand a retrieved chunk back into its surrounding context instead of relying only on chunk_overlap.
solid answer
~50 sEvery node LlamaIndex produces carries a `relationships` dict keyed by `NodeRelationship`. Splitters set `SOURCE` (back to the originating document) always, and `PREVIOUS`/`NEXT` when `include_prev_next_rel` is on, which is the default for node parsers in llama-index-core 0.14.x. Each entry is a `RelatedNodeInfo` holding the neighbour's `node_id`, not the neighbour itself — resolving it requires a docstore, which is why the convenience accessors `node.prev_node` and `node.next_node` fail if the nodes were never persisted. The practical value is context expansion: retrieval returns the one chunk whose embedding matched, and you widen it to its neighbours so the answer is not cut off at a chunk boundary. That is a cheaper and more surgical alternative to raising `chunk_overlap`, because the duplication happens at read time for the few retrieved nodes rather than at write time for the whole corpus.
code
python · 9 linesfrom llama_index.core.schema import NodeRelationship
node = nodes[5]
print(node.relationships[NodeRelationship.SOURCE].node_id)
nxt = node.relationships.get(NodeRelationship.NEXT)
if nxt is not None:
neighbour = docstore.get_node(nxt.node_id)
context = node.text + "\n" + neighbour.textgo deeper
Know that each node records links to the chunk before it, the chunk after it, and the document it came from, and that these links are how you get back to surrounding context.
Explain that relationships store node ids in a RelatedNodeInfo rather than the nodes themselves, that include_prev_next_rel is on by default, and that resolving links needs a docstore.
Argue the read-time versus write-time tradeoff concretely: small distinct embedded units plus neighbour expansion usually beat large chunks with large overlap on both cost and retrieval precision.
Treat the node graph as persisted state with a lifecycle — ids change on re-ingest, the docstore must be backed up and versioned alongside the vector store, and any component that caches node ids needs an invalidation story.
## The relationships dict A `BaseNode` in LlamaIndex has a `relationships: Dict[NodeRelationship, RelatedNodeInfo]` field. `NodeRelationship` is an enum with `SOURCE`, `PREVIOUS`, `NEXT`, `PARENT` and `CHILD`. A `RelatedNodeInfo` records the related node's `node_id` plus a little bookkeeping — crucially it is a **reference**, not an embedded copy. Nodes form a linked structure, not a nested one. Splitters populate these during parsing: - `SOURCE` — set on every node, pointing at the `Document` it was cut from. This is what powers citation and what `ref_doc_id` reads. - `PREVIOUS` / `NEXT` — set between consecutive chunks of the same document when the parser's `include_prev_next_rel` flag is on, which is the default. - `PARENT` / `CHILD` — set by parsers that build multiple levels rather than a flat sequence. ## Why sequence links matter Chunking destroys reading order as far as the vector index is concerned: a retriever returns the k nodes whose embeddings are nearest the query, with no guarantee they are adjacent or in order. But answers frequently need adjacency. A question about "the exception to that rule" matches a chunk stating the rule, while the exception lives in the following chunk. Overlap helps only if the exception happens to fall within the overlap window. `PREVIOUS`/`NEXT` let you fix this at read time: take the retrieved node, follow its neighbour ids, fetch those nodes from the docstore and hand the LLM a contiguous window. The cost is paid once per retrieved node per query, instead of being baked into every stored vector. ## Read-time versus write-time context This is the tradeoff an interviewer is probing: - **Bigger `chunk_overlap`** duplicates text into the stored, embedded representation. It inflates the index, inflates embedding spend, and makes top-k results redundant because neighbouring vectors are near-identical. - **Neighbour expansion via relationships** keeps the embedded units small and distinct — which is generally better for retrieval precision — and reconstitutes context only for the handful of nodes actually returned. A common configuration is therefore small chunks with modest overlap plus neighbour expansion, rather than large chunks with large overlap. ## The docstore requirement Because relationships store ids, following them requires somewhere to look ids up. In a typical in-memory build the nodes live in the `StorageContext`'s docstore and everything just works; the accessors `node.prev_node`, `node.next_node`, `node.source_node`, `node.parent_node` and `node.child_nodes` return the `RelatedNodeInfo` for you to resolve. If you embed nodes into an external vector store and never persist the node set, the ids point at nothing, and neighbour expansion silently returns nothing useful. Persisting the docstore is part of shipping this pattern, not an optimisation. ## Relationships are per document, not global The first node of a document has no `PREVIOUS` and the last has no `NEXT`; the chain does not run across document boundaries. Code that walks neighbours must handle the missing key rather than assuming a link always exists. Similarly, re-ingesting a document produces new node ids, so stale ids held elsewhere will dangle — one more reason to treat the docstore as the source of truth for the node graph. ## A related but distinct mechanism LlamaIndex also offers `SentenceWindowNodeParser`, which takes a different route to the same goal: it makes each node a single sentence but stashes a window of surrounding sentences in the node's metadata under a window key, to be swapped in later. That trades index size for a much finer retrieval unit. The relationship-based approach keeps ordinary chunks and follows links; the window approach keeps tiny chunks and carries their context along. Both beat simply inflating `chunk_overlap`. ## What to say in an interview Name the enum, say the entries hold ids rather than nodes, note that `include_prev_next_rel` is on by default, and then make the real point: sequence relationships exist so that the unit you embed can stay small and precise while the unit you show the model stays large enough to answer with. Add that this only works if the nodes are in a docstore you actually persisted.
- You follow node.next_node and get nothing back. What are the likely causes?Either the node is the last chunk of its document, so no NEXT entry exists, or the parser ran with `include_prev_next_rel` disabled, or — most often in production — the nodes were embedded into an external vector store and the docstore was never persisted, so the stored id resolves to nothing. Relationships hold ids, and ids need a lookup store to become nodes.
- Why prefer neighbour expansion over simply doubling chunk_overlap?Overlap pays at write time for the entire corpus: every stored vector carries duplicated text, the index grows, embedding cost rises, and top-k results become near-duplicates of one another. Neighbour expansion pays at read time for only the handful of retrieved nodes, so the embedded units stay small and distinct — which improves retrieval precision — while the prompt still receives contiguous context.
- Does the PREVIOUS/NEXT chain span across different source documents?No. Splitters link consecutive chunks within a single document only, so the first node of each document has no PREVIOUS and the last has no NEXT. Code that walks the chain must handle the absent key. If you need cross-document ordering — pages of one manual split into separate files, for example — you have to encode that yourself in metadata.
saying these in an interview costs you the question
- Thinks relationships embed the neighbouring node itself
- Assumes neighbour lookup works without a persisted docstore
- Believes the prev/next chain runs across documents
- Says overlap and neighbour expansion solve different problems
- Forgets that re-ingestion changes node ids