skip to content

When would you pick TokenTextSplitter over SentenceSplitter in LlamaIndex?

level: middleimportance: should knowfreq 52%

answer

  1. semantic integrity versus uniform windows
  2. one respects sentences, one does not
  3. transcripts, logs and code break sentences
  4. the two defaults differ in overlap
  5. measure the node-length distribution

basics

~20 s

Pick TokenTextSplitter when the text has no reliable sentence structure — logs, code, transcripts, scraped markup — and you want tight, uniform token windows. SentenceSplitter is the better default for prose because it refuses to cut mid-sentence.

solid answer

~50 s

Both live in `llama_index.core.node_parser` and both budget in tokens, but they differ in what they protect. `SentenceSplitter` is hierarchy-aware: it splits on the paragraph separator, then sentences, then a secondary punctuation regex, and only then on plain whitespace, packing whole units up to `chunk_size`. That keeps semantic units intact but yields uneven node sizes and can under-fill chunks badly on text with no punctuation. `TokenTextSplitter` splits on a separator with backup separators and fills windows to the budget, so it will cut through a sentence to keep chunks uniform. Use it for log lines, source code, ASR transcripts without punctuation, or when you need predictable node sizes for a fixed context budget. In llama-index-core 0.14.x their defaults differ too — both default to `chunk_size=1024`, but `SentenceSplitter` defaults to 200 overlap versus 20 for `TokenTextSplitter`, so switching one for the other silently changes duplication by an order of magnitude.

code

python · 7 lines
python
from llama_index.core.node_parser import SentenceSplitter, TokenTextSplitter

prose = SentenceSplitter(chunk_size=512, chunk_overlap=64)
logs = TokenTextSplitter(chunk_size=512, chunk_overlap=64, separator="\n")

prose_nodes = prose.get_nodes_from_documents(article_docs)
log_nodes = logs.get_nodes_from_documents(log_docs)

go deeper

for a junior

Know that both splitters exist, that SentenceSplitter tries not to cut sentences in half, and that TokenTextSplitter fills fixed token windows regardless of punctuation.

for a middle

Explain the split cascade — paragraph, then sentence, then clause punctuation, then whitespace — and name concrete content types where sentence awareness buys nothing, such as logs, code and unpunctuated transcripts.

for a senior

Show that you validate the choice on real data by inspecting node-length distributions and reading samples, and that you restate chunk_size and chunk_overlap explicitly instead of inheriting class defaults.

for a principal

Frame it as a data-contract question: heterogeneous corpora usually need per-source parsers rather than one global splitter, and the ingestion layer should route documents by type so that a single default never silently degrades one whole content class.

## Two splitters, one budget, different priorities LlamaIndex ships several text splitters that all implement the same `NodeParser` surface — you call `get_nodes_from_documents(docs)` and get `TextNode`s back — but they disagree about what to sacrifice when text does not divide evenly. `SentenceSplitter` optimises for **semantic integrity**. `TokenTextSplitter` optimises for **uniformity**. Everything else follows from that. ## How SentenceSplitter decides It applies a cascade of splitting functions in order of decreasing preference: the paragraph separator first (three newlines by default), then a sentence tokenizer, then a secondary regex that breaks on clause punctuation, then the plain `separator` (a space). It splits only as deep as it must: if paragraph-level pieces already fit the budget, sentences are never touched. It then merges the resulting pieces back into chunks, closing a chunk as soon as adding the next piece would exceed `chunk_size`. The payoff is that a node almost never ends mid-sentence, so the embedding represents a complete thought and the text handed to the LLM reads naturally in the prompt. The price is variance: node sizes swing widely, and pathological input — a 5,000-token paragraph with no internal punctuation, minified JSON, a wall of code — forces the splitter all the way down to whitespace splitting, where its advantage evaporates. ## How TokenTextSplitter decides `TokenTextSplitter` splits the text on its `separator` with `backup_separators` as fallbacks, then packs pieces into windows of `chunk_size` tokens with `chunk_overlap` carried between them. It has no notion of sentences or paragraphs, so it will happily end a chunk halfway through a clause. What you get in return is nodes that are consistently near the budget, which matters when you are computing a context budget precisely — for example, "top-k of 6 at 400 tokens each must fit in 2,400 tokens of prompt" is a promise `TokenTextSplitter` keeps and `SentenceSplitter` does not. ## The content types that decide it for you - **Prose, documentation, articles, policies** — `SentenceSplitter`. Sentence boundaries are real and carry meaning. - **Source code** — neither is ideal; a language-aware splitter is better, but between these two, `TokenTextSplitter` at least does not pretend that periods in `obj.method()` are sentence ends. - **Raw ASR transcripts without punctuation** — `SentenceSplitter` degrades to whitespace splitting and gives you nothing; `TokenTextSplitter` gives you predictable windows. - **Log lines and CSV-like records** — `TokenTextSplitter` with a newline separator, or a purpose-built parser. - **Scraped HTML converted to text** — clean it first; both splitters will otherwise chunk boilerplate. ## The default trap In llama-index-core 0.14.x, `SentenceSplitter` defaults to `chunk_size=1024, chunk_overlap=200` while `TokenTextSplitter` defaults to `chunk_size=1024, chunk_overlap=20`. Swapping the class without restating both numbers changes your duplication factor roughly tenfold, which shows up as a surprising change in index size, embedding bill and top-k redundancy. Always pass both explicitly rather than relying on class defaults; it also documents intent for the next reader. ## Wiring the choice in Either splitter can be installed globally as `Settings.node_parser`, or passed per index build through `transformations=[...]`. Explicit beats global: when both are present the explicitly passed splitter runs, and mixing the two is a common cause of "my splitter change did nothing". Because both are `TransformComponent`s, the same instance can be reused wherever a transformation list is accepted. ## How to justify the choice in an interview The strong answer is not a rule but a measurement: chunk a representative sample both ways, print the token-length distribution, and read a handful of nodes. If the `SentenceSplitter` output contains complete, self-contained passages with a reasonable size spread, keep it. If it produced a bimodal mess of 30-token fragments and 1,000-token blobs, the input lacks the structure that splitter assumes, and the uniform windows of `TokenTextSplitter` — or a domain-specific parser — will retrieve better.

  • What is the risk of switching splitters without restating chunk_overlap?
    The class defaults differ. In llama-index-core 0.14.x `SentenceSplitter` defaults to an overlap of 200 tokens while `TokenTextSplitter` defaults to 20, so moving from one to the other silently changes how much text is duplicated across nodes — index size, embedding cost and near-duplicate top-k results all shift. Pass `chunk_size` and `chunk_overlap` explicitly so the behaviour is stated rather than inherited.
  • Your SentenceSplitter output is full of 20-token nodes. What does that tell you?
    The source is not prose in the shape the splitter assumes — typically headings, bullet lists, table rows or navigation boilerplate converted to text, each separated by paragraph breaks. Fix it upstream by cleaning or merging short blocks, or switch to a splitter that packs uniform windows. Tiny nodes embed poorly because they carry too little context for a query vector to match.
  • How do you make a splitter choice reviewable rather than a matter of taste?
    Chunk a representative sample with each candidate, print the token-length distribution and read a random sample of nodes for readability and self-containment, then run the same evaluation query set through both indexes. The comparison is cheap, it is reproducible, and it turns an opinion into two numbers plus a handful of examples anyone can check.

saying these in an interview costs you the question

  • Claims SentenceSplitter never produces oversized nodes
  • Uses sentence splitting on punctuation-free transcripts
  • Swaps splitter classes and inherits a different default overlap
  • Thinks TokenTextSplitter understands sentence boundaries
  • Picks a splitter without ever looking at the produced nodes

context