When would you pick TokenTextSplitter over SentenceSplitter in LlamaIndex?
answer
- semantic integrity versus uniform windows
- one respects sentences, one does not
- transcripts, logs and code break sentences
- the two defaults differ in overlap
- measure the node-length distribution
basics
~20 sPick TokenTextSplitter when the text has no reliable sentence structure — logs, code, transcripts, scraped markup — and you want tight, uniform token windows. SentenceSplitter is the better default for prose because it refuses to cut mid-sentence.
solid answer
~50 sBoth live in `llama_index.core.node_parser` and both budget in tokens, but they differ in what they protect. `SentenceSplitter` is hierarchy-aware: it splits on the paragraph separator, then sentences, then a secondary punctuation regex, and only then on plain whitespace, packing whole units up to `chunk_size`. That keeps semantic units intact but yields uneven node sizes and can under-fill chunks badly on text with no punctuation. `TokenTextSplitter` splits on a separator with backup separators and fills windows to the budget, so it will cut through a sentence to keep chunks uniform. Use it for log lines, source code, ASR transcripts without punctuation, or when you need predictable node sizes for a fixed context budget. In llama-index-core 0.14.x their defaults differ too — both default to `chunk_size=1024`, but `SentenceSplitter` defaults to 200 overlap versus 20 for `TokenTextSplitter`, so switching one for the other silently changes duplication by an order of magnitude.
code
python · 7 linesfrom llama_index.core.node_parser import SentenceSplitter, TokenTextSplitter
prose = SentenceSplitter(chunk_size=512, chunk_overlap=64)
logs = TokenTextSplitter(chunk_size=512, chunk_overlap=64, separator="\n")
prose_nodes = prose.get_nodes_from_documents(article_docs)
log_nodes = logs.get_nodes_from_documents(log_docs)go deeper
Know that both splitters exist, that SentenceSplitter tries not to cut sentences in half, and that TokenTextSplitter fills fixed token windows regardless of punctuation.
Explain the split cascade — paragraph, then sentence, then clause punctuation, then whitespace — and name concrete content types where sentence awareness buys nothing, such as logs, code and unpunctuated transcripts.
Show that you validate the choice on real data by inspecting node-length distributions and reading samples, and that you restate chunk_size and chunk_overlap explicitly instead of inheriting class defaults.
Frame it as a data-contract question: heterogeneous corpora usually need per-source parsers rather than one global splitter, and the ingestion layer should route documents by type so that a single default never silently degrades one whole content class.
## Two splitters, one budget, different priorities LlamaIndex ships several text splitters that all implement the same `NodeParser` surface — you call `get_nodes_from_documents(docs)` and get `TextNode`s back — but they disagree about what to sacrifice when text does not divide evenly. `SentenceSplitter` optimises for **semantic integrity**. `TokenTextSplitter` optimises for **uniformity**. Everything else follows from that. ## How SentenceSplitter decides It applies a cascade of splitting functions in order of decreasing preference: the paragraph separator first (three newlines by default), then a sentence tokenizer, then a secondary regex that breaks on clause punctuation, then the plain `separator` (a space). It splits only as deep as it must: if paragraph-level pieces already fit the budget, sentences are never touched. It then merges the resulting pieces back into chunks, closing a chunk as soon as adding the next piece would exceed `chunk_size`. The payoff is that a node almost never ends mid-sentence, so the embedding represents a complete thought and the text handed to the LLM reads naturally in the prompt. The price is variance: node sizes swing widely, and pathological input — a 5,000-token paragraph with no internal punctuation, minified JSON, a wall of code — forces the splitter all the way down to whitespace splitting, where its advantage evaporates. ## How TokenTextSplitter decides `TokenTextSplitter` splits the text on its `separator` with `backup_separators` as fallbacks, then packs pieces into windows of `chunk_size` tokens with `chunk_overlap` carried between them. It has no notion of sentences or paragraphs, so it will happily end a chunk halfway through a clause. What you get in return is nodes that are consistently near the budget, which matters when you are computing a context budget precisely — for example, "top-k of 6 at 400 tokens each must fit in 2,400 tokens of prompt" is a promise `TokenTextSplitter` keeps and `SentenceSplitter` does not. ## The content types that decide it for you - **Prose, documentation, articles, policies** — `SentenceSplitter`. Sentence boundaries are real and carry meaning. - **Source code** — neither is ideal; a language-aware splitter is better, but between these two, `TokenTextSplitter` at least does not pretend that periods in `obj.method()` are sentence ends. - **Raw ASR transcripts without punctuation** — `SentenceSplitter` degrades to whitespace splitting and gives you nothing; `TokenTextSplitter` gives you predictable windows. - **Log lines and CSV-like records** — `TokenTextSplitter` with a newline separator, or a purpose-built parser. - **Scraped HTML converted to text** — clean it first; both splitters will otherwise chunk boilerplate. ## The default trap In llama-index-core 0.14.x, `SentenceSplitter` defaults to `chunk_size=1024, chunk_overlap=200` while `TokenTextSplitter` defaults to `chunk_size=1024, chunk_overlap=20`. Swapping the class without restating both numbers changes your duplication factor roughly tenfold, which shows up as a surprising change in index size, embedding bill and top-k redundancy. Always pass both explicitly rather than relying on class defaults; it also documents intent for the next reader. ## Wiring the choice in Either splitter can be installed globally as `Settings.node_parser`, or passed per index build through `transformations=[...]`. Explicit beats global: when both are present the explicitly passed splitter runs, and mixing the two is a common cause of "my splitter change did nothing". Because both are `TransformComponent`s, the same instance can be reused wherever a transformation list is accepted. ## How to justify the choice in an interview The strong answer is not a rule but a measurement: chunk a representative sample both ways, print the token-length distribution, and read a handful of nodes. If the `SentenceSplitter` output contains complete, self-contained passages with a reasonable size spread, keep it. If it produced a bimodal mess of 30-token fragments and 1,000-token blobs, the input lacks the structure that splitter assumes, and the uniform windows of `TokenTextSplitter` — or a domain-specific parser — will retrieve better.
- What is the risk of switching splitters without restating chunk_overlap?The class defaults differ. In llama-index-core 0.14.x `SentenceSplitter` defaults to an overlap of 200 tokens while `TokenTextSplitter` defaults to 20, so moving from one to the other silently changes how much text is duplicated across nodes — index size, embedding cost and near-duplicate top-k results all shift. Pass `chunk_size` and `chunk_overlap` explicitly so the behaviour is stated rather than inherited.
- Your SentenceSplitter output is full of 20-token nodes. What does that tell you?The source is not prose in the shape the splitter assumes — typically headings, bullet lists, table rows or navigation boilerplate converted to text, each separated by paragraph breaks. Fix it upstream by cleaning or merging short blocks, or switch to a splitter that packs uniform windows. Tiny nodes embed poorly because they carry too little context for a query vector to match.
- How do you make a splitter choice reviewable rather than a matter of taste?Chunk a representative sample with each candidate, print the token-length distribution and read a random sample of nodes for readability and self-containment, then run the same evaluation query set through both indexes. The comparison is cheap, it is reproducible, and it turns an opinion into two numbers plus a handful of examples anyone can check.
saying these in an interview costs you the question
- Claims SentenceSplitter never produces oversized nodes
- Uses sentence splitting on punctuation-free transcripts
- Swaps splitter classes and inherits a different default overlap
- Thinks TokenTextSplitter understands sentence boundaries
- Picks a splitter without ever looking at the produced nodes