In LlamaIndex, what is the difference between a Document and a Node?
answer
- one is the source, one is retrievable
- loaders return the first kind
- chunks keep a link home
- metadata copied downhill at parse time
- subclass relationship in the schema module
basics
~20 sA Document is one whole source item as loaded — a file, a page, an API record. Nodes are the pieces produced from it, each with its own id, metadata, embedding, and a stored reference back to the source Document.
solid answer
~40 sLoading produces `Document` objects: one unit of source material with `text`, a `metadata` dict, and an `id_`. Indexing does not embed Documents — a node parser turns each Document into `TextNode` objects, and those nodes are what get embedded, stored, retrieved and stuffed into the prompt. `Document` is actually a subclass of `TextNode`, so the two share the same fields; the distinction is role, not shape. Each node keeps a `NodeRelationship.SOURCE` entry pointing at its Document's id (readable as `node.ref_doc_id`), which is what makes citation back to the original file possible and what lets you delete or refresh everything derived from one source. Metadata you attach at load time is copied onto every node parsed from that Document, which is why ingestion-time metadata is the leverage point for filtered retrieval.
code
python · 14 linesfrom llama_index.core import Document, VectorStoreIndex
from llama_index.core.node_parser import SentenceSplitter
doc = Document(
text="Refunds are processed within 5 business days.",
metadata={"source": "policy.md", "team": "support"},
id_="policy.md",
)
nodes = SentenceSplitter(chunk_size=256).get_nodes_from_documents([doc])
print(nodes[0].metadata["team"]) # inherited from the Document
print(nodes[0].ref_doc_id) # 'policy.md'
index = VectorStoreIndex(nodes)go deeper
Be able to say that a loader gives you Documents, that Documents get split into nodes, and that nodes are the things stored and retrieved. Naming SimpleDirectoryReader as the thing producing Documents is enough at this level.
Explain the mechanics: nodes carry their own ids and embeddings, a SOURCE relationship points back to the Document, and Document metadata is copied onto every node at parse time. Knowing that Document subclasses TextNode is a good signal.
Show the operational consequence — stable document ids are what make deletion, refresh and citation work, and metadata missed at ingestion time means re-ingesting the corpus. Talk about how you verify provenance survives into retrieved results.
Own the contract: which system assigns document identity, what metadata schema every source must supply, and whether chunking belongs inside LlamaIndex at all when an upstream platform already produces retrievable units.
## The two objects LlamaIndex splits ingestion into two representations. A `Document` is what a reader returns: one unit of source material, exactly as the loader saw it. `SimpleDirectoryReader` returns one Document per text file; a PDF reader may return one Document per page; a database reader returns one per row. A Document holds `text`, a `metadata` dict, an `id_`, and the metadata-visibility fields described below. A `Node` — concretely `TextNode`, and its relatives `ImageNode` and `IndexNode` — is a retrievable unit. Nodes are produced from Documents by a node parser during indexing or an ingestion pipeline run. Each node has its own `id_`, its own copy of the metadata it inherited, optionally its own embedding, and a `relationships` map linking it to its source and to neighbouring nodes. ## They are the same class family In `llama_index.core.schema`, `Document` is a subclass of `TextNode`. That surprises people, and it is worth saying out loud in an interview: the difference is not a different data structure, it is a different job. Document means "raw source unit, not yet retrievable"; TextNode means "unit that will be embedded and returned by a retriever". Because they share a base, `get_content()`, `metadata`, `id_` and the metadata-exclusion fields behave identically on both. ## Why the split exists Embedding models and context windows are bounded, and retrieval quality depends on returning a focused passage rather than a whole file. So the pipeline must split. But once you split, you lose provenance — and provenance is what turns an answer into a citable answer. The Document/Node split keeps both: the Document is the durable identity of the source, and every node derived from it carries `relationships[NodeRelationship.SOURCE]`, a `RelatedNodeInfo` holding that Document's id. `node.ref_doc_id` reads it directly. That back-pointer does real operational work: - **Citation.** After retrieval you can report which file, page or record the passage came from, because the node still knows its origin — and the metadata copied onto it usually carries `file_name`/`file_path` too. - **Deletion.** Removing a source means removing every node it produced. `index.delete_ref_doc(doc_id, delete_from_docstore=True)` does exactly that, using the SOURCE relationship, instead of forcing you to track chunk ids yourself. - **Refresh.** `index.refresh_ref_docs(documents)` compares incoming Documents against stored ones by id and hash, and re-processes only what changed. All three depend on **stable Document ids**. A generated UUID per run makes yesterday's nodes orphans; that is why `SimpleDirectoryReader(filename_as_id=True)` exists, and why a custom reader should set `id_` from something durable in the source system. ## Metadata flows downhill Metadata set on a Document is copied onto every node parsed from it. This is the single most consequential fact about the split. Anything you want to filter on at query time — tenant, product version, department, publication date, source URL — has to be on the Document at ingestion time, or you will be re-ingesting the corpus to add it later. Nodes may also gain their own metadata afterwards from extractor transformations, but the Document-level fields are the ones that come free and stay consistent. Both Documents and Nodes carry `excluded_embed_metadata_keys` and `excluded_llm_metadata_keys`, so metadata can be present for filtering while being hidden from the embedding text or from the prompt. ## What this looks like in practice You rarely construct nodes by hand. The usual flow is: reader returns `List[Document]` → you attach or enrich metadata → an index constructor or an `IngestionPipeline` runs transformations that parse Documents into nodes and embed them → the vector store holds nodes. You *can* build a `TextNode` directly and index it, which is what you do when another system already owns chunking. ## Common failure modes - Treating the Document as the retrievable unit and wondering why whole 80-page PDFs come back as "a result". - Attaching metadata after chunking, per node, by hand — inconsistent and expensive, when Document-level metadata would have propagated. - Letting reader-generated ids change on every run, which breaks deletion, refresh and deduplication all at once. - Assuming embeddings live on Documents; they live on nodes.
- A source file is deleted from the corpus — how do you remove what it contributed to the index?Delete by source id rather than by chunk. `index.delete_ref_doc(doc_id, delete_from_docstore=True)` walks the SOURCE relationship and removes every node derived from that Document, plus the docstore entry. It only works if the Document's id was stable across runs — a random per-run id leaves the old nodes stranded with no way to name them.
- Does metadata added to a Document reach the nodes, or do you have to set it per node?It propagates. The node parser copies the Document's `metadata` dict onto each node it produces, so Document-level metadata is the cheap, consistent way to make filtered retrieval possible. Per-node metadata is added later by extractor transformations or by hand, and is the right place only for facts that genuinely differ between chunks, such as an extracted section title.
- Can you skip Documents entirely and index TextNode objects directly?Yes. `VectorStoreIndex(nodes)` accepts nodes you built yourself, which is normal when an upstream system already chunked the corpus. The cost is provenance: unless you populate `relationships[NodeRelationship.SOURCE]` and stable ids yourself, you lose citation back to a source and the ref-doc deletion and refresh paths stop working.
A Document is the book you put on the shelf; nodes are the individual paragraphs a librarian photocopies and files, each stamped with which book it came from.
saying these in an interview costs you the question
- Says Documents are what get embedded and retrieved
- Claims Document and Node are unrelated classes
- Thinks nodes lose all knowledge of their source file
- Believes metadata must be re-attached to every chunk manually
- Assumes a random per-run document id is harmless