skip to content

Node Parsing & Chunking

You will learn node parsing: SentenceSplitter, TokenTextSplitter and SemanticSplitter, the NodeParser API, tuning chunk size and overlap, metadata propagation, and parent/child and next/previous node relationships. Interviewers ask because chunking is the highest-leverage knob in a RAG pipeline, and they want reasoning rather than a default value you copied.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

questions

6

In LlamaIndex's SentenceSplitter, what do chunk_size and chunk_overlap measure?

level: juniorimportance: must knowfreq 72%

answer

  1. counted in tokens, not characters
  2. the ceiling, not the target
  3. boundary insurance costs duplication
  4. default 1024 with 200 overlap
  5. overlap larger than size is rejected

basics

~20 s

Both are token counts, not characters. chunk_size (default 1024) caps how many tokens a node may hold, measured with the tokenizer LlamaIndex is configured with; chunk_overlap (default 200) repeats that many trailing tokens at the start of the next node.

solid answer

~40 s

In llama-index-core 0.14.x, `SentenceSplitter(chunk_size=1024, chunk_overlap=200)` works in **tokens**, counted by the tokenizer on `Settings.tokenizer` (tiktoken by default), not characters or words. `chunk_size` is an upper bound: the splitter first breaks text on paragraph and sentence boundaries and then packs whole sentences together until adding one more would exceed the budget, so real nodes are often noticeably smaller than the number you set. `chunk_overlap` makes consecutive nodes share a tail/head region so a fact straddling a boundary still appears whole in at least one node — it costs storage and embedding calls proportional to the duplication. The overlap must be smaller than the chunk size; constructing the splitter the other way round raises a `ValueError`. Because the counting tokenizer is usually not your embedding model's tokenizer, treat the number as approximate.

code

python · 7 lines
python
from llama_index.core.node_parser import SentenceSplitter

splitter = SentenceSplitter(chunk_size=256, chunk_overlap=32)
nodes = splitter.get_nodes_from_documents(documents)

for node in nodes[:3]:
    print(len(splitter._tokenizer(node.text)), node.text[:60])

go deeper

for a junior

Be able to say the numbers are token counts, name the defaults of 1024 and 200, and explain in one sentence that overlap repeats text so a fact on a boundary is not cut in half.

for a middle

Explain that chunk_size is an upper bound applied after sentence and paragraph splitting, so nodes come out smaller, and that overlap should be scaled as a percentage of chunk size rather than left at the default.

for a senior

Show that you check the produced token distribution instead of trusting the parameter, and that you account for the mismatch between the counting tokenizer and the embedding model's tokenizer when sizing near a model limit.

for a principal

Own the cost model: overlap multiplies embedding spend, index footprint and re-ingest time across the whole corpus, so the chunking parameters are a budget decision that should be measured against retrieval quality, not a default copied from a tutorial.

## What the splitter is doing A `NodeParser` in LlamaIndex turns each input document into a list of `TextNode` objects — the units that get embedded, stored and retrieved. `SentenceSplitter` (from `llama_index.core.node_parser`) is the default choice for prose. Its two headline knobs, `chunk_size` and `chunk_overlap`, are the highest-leverage numbers in the whole pipeline, and the first thing an interviewer checks is whether you know what they are counted in. ## Tokens, not characters Both values are **token counts**. The splitter holds a `tokenizer` callable, which defaults to the global one on `Settings.tokenizer` — tiktoken's encoding for OpenAI models unless you replace it. So `chunk_size=512` means "at most about 512 tokens as that tokenizer counts them", which for English prose is roughly 350–400 words and roughly 2,000 characters, but for code, CJK text or heavy punctuation the ratio is very different. The practical consequence: the tokenizer doing the counting is usually **not** the tokenizer of the embedding model you will send the chunk to. If you set `chunk_size` right at your embedding model's input limit, a chunk that measures 512 tokens under tiktoken can measure more under the embedding model's own tokenizer and get silently truncated by the provider. Leave headroom, or install a tokenizer that matches the model. ## chunk_size is a ceiling, not a target `SentenceSplitter` is boundary-aware. It splits the text on the paragraph separator, then on sentence boundaries, then (if a single sentence is still too big) on a secondary regex, and finally on the plain separator. It then merges those pieces back together into chunks, stopping before the budget is exceeded. Because it refuses to cut mid-sentence when it can avoid it, a `chunk_size=1024` run routinely yields nodes of 300–900 tokens, and short documents produce a single small node. Seeing nodes smaller than `chunk_size` is normal and not a bug. Defaults quoted here are llama-index-core 0.14.x: `chunk_size=1024`, `chunk_overlap=200`, `separator=" "`, and a paragraph separator of three newlines. ## What overlap actually buys Overlap duplicates the last `chunk_overlap` tokens of one chunk at the head of the next. The purpose is boundary insurance: a definition whose subject is in one sentence and whose predicate is in the next survives intact in at least one node, and the retriever therefore has a node whose embedding represents the complete statement. Without overlap, a fact split across a boundary is represented by two half-facts, and neither embeds close to the query. The cost is linear duplication. At `chunk_size=512, chunk_overlap=200` you are re-embedding and re-storing roughly 40% of your corpus, and near-duplicate neighbours crowd the top-k results, so several returned nodes may carry the same sentence. The usual working range is 10–20% of `chunk_size`; the 200-token default is deliberately generous relative to a small chunk size and should be reduced when you reduce `chunk_size`. ## Guardrails and where the values live The constructor validates the relationship between the two: an overlap larger than the chunk size raises a `ValueError` rather than silently clamping. You can set the numbers in two places — globally via `Settings.chunk_size` / `Settings.chunk_overlap`, which reconfigures the default node parser, or explicitly by passing an instance, e.g. `VectorStoreIndex.from_documents(docs, transformations=[SentenceSplitter(chunk_size=256, chunk_overlap=32)])`. An explicitly passed splitter is what runs; the global setting only supplies the default when you pass nothing. Mixing both and forgetting the precedence is a very common source of "my chunk size had no effect". ## Reading the result Always inspect the output rather than trusting the number: `nodes = splitter.get_nodes_from_documents(docs)` then look at `len(nodes)` and the token length of a sample of `node.text`. A distribution with a long tail of 40-token nodes usually means the source had many short paragraphs; a distribution pinned exactly at the ceiling usually means you are chunking something without sentence structure, where `TokenTextSplitter` would be the more honest tool.

  • Which tokenizer does SentenceSplitter count with, and why does that matter?
    It uses the callable on `Settings.tokenizer`, tiktoken's OpenAI encoding by default, unless you pass your own `tokenizer` to the splitter. That is usually a different tokenizer from your embedding model's, so the count is an approximation. If you size chunks right at the embedding model's input limit, the provider may truncate text the splitter believed fit. Leave headroom or install a matching tokenizer.
  • You cut chunk_size from 1024 to 256 but left chunk_overlap at 200. What goes wrong?
    You are now duplicating roughly 78% of every chunk. Storage, embedding cost and index size balloon, and retrieval degrades because the top-k results are near-identical windows of the same passage instead of distinct evidence. Overlap should scale with chunk size — commonly 10–20% — so 256 tokens pairs with roughly 25–50 tokens of overlap.
  • Why are many of your nodes far below chunk_size even though the documents are long?
    `SentenceSplitter` respects structure first: it splits on paragraphs and sentences and only packs whole units up to the budget, so a chunk closes as soon as the next sentence would overflow. Documents made of short paragraphs, headings, list items or table rows therefore produce many small nodes. That is expected behaviour, not misconfiguration; if you need tightly packed fixed windows, `TokenTextSplitter` is the tool.

saying these in an interview costs you the question

  • Says chunk_size counts characters or words
  • Thinks every node comes out at exactly chunk_size
  • Keeps a 200-token overlap after shrinking chunks to 256
  • Assumes overlap larger than chunk_size is silently clamped
  • Believes the splitter's tokenizer matches the embedding model

context

open as a page

How does document metadata change what LlamaIndex's node parser emits?

level: middleimportance: must knowfreq 50%

basics

~20 s

Metadata is copied onto every node, prepended to the node's text for embedding and for the LLM, and counted against chunk_size — so long metadata shrinks the real text budget and can even make it non-positive. Use excluded_embed_metadata_keys and excluded_llm_metadata_keys to control what leaks where.

open as a page

What are NodeRelationship.PREVIOUS and NEXT for on LlamaIndex nodes?

level: middleimportance: should knowfreq 42%

basics

~20 s

They are pointers stored in each node's relationships dict that link it to the chunks immediately before and after it in the original document. They let you expand a retrieved chunk back into its surrounding context instead of relying only on chunk_overlap.

open as a page

When would you pick TokenTextSplitter over SentenceSplitter in LlamaIndex?

level: middleimportance: should knowfreq 52%

basics

~20 s

Pick TokenTextSplitter when the text has no reliable sentence structure — logs, code, transcripts, scraped markup — and you want tight, uniform token windows. SentenceSplitter is the better default for prose because it refuses to cut mid-sentence.

open as a page

What does LlamaIndex's HierarchicalNodeParser produce, and what must you store?

level: seniorimportance: should knowfreq 34%

basics

~20 s

It chunks each document several times at decreasing sizes and links the levels with PARENT and CHILD relationships. You embed only the leaf nodes returned by get_leaf_nodes, but every node of every level must go into the docstore or the parent links dangle.

open as a page

What does LlamaIndex's SemanticSplitterNodeParser cost you at ingest time?

level: seniorimportance: should knowfreq 36%

basics

~20 s

It requires an embed_model and embeds sentence groups across the whole corpus before any node exists, so ingestion becomes an extra full embedding pass with the latency, spend and rate limits that implies. It also gives no hard token ceiling on the chunks it emits.

open as a page