skip to content

Semantic chunking yields chunks from 40 to 3,000 tokens — what breaks downstream?

level: seniorimportance: should knowfreq 40%

answer

  1. coherence was optimised, not uniformity
  2. long tail in both directions
  3. silent truncation at the embedder
  4. merge the small, re-split the large
  5. pack by token budget, not fixed k

basics

~20 s

Variable chunk sizes break the assumptions after retrieval: oversized chunks get truncated by the embedding model and swallow the prompt budget, tiny chunks retrieve without enough context to answer, and a fixed top-k returns wildly different amounts of text per query.

solid answer

~50 s

Semantic chunking optimises for coherence, not uniformity, so the size distribution is long-tailed by construction. Three concrete failures follow. **Oversized chunks** may exceed the embedding model's input limit and be silently truncated, so part of the chunk is indexed under a vector that does not describe it; they also dominate the prompt, and two of them can consume the whole context budget that top-k assumed. **Undersized chunks** — a single sentence left between two strong breakpoints — retrieve on a keyword-ish match but carry no surrounding context, so the generator answers from a fragment. **Variance itself** breaks capacity planning: with a fixed k, one query pulls 400 tokens of evidence and the next pulls 9,000, which makes latency and cost per query unpredictable. The standard fix is to bound the distribution after semantic splitting: merge chunks below a minimum, recursively re-split chunks above a maximum, and pack the prompt by token budget rather than by a fixed k.

go deeper

for a junior

Know that semantic chunking produces chunks of very different sizes on purpose, and that both very long and very short chunks cause problems when they are retrieved.

for a middle

Explain the concrete consequences: truncation at the embedding model's input limit, fragments that lack context, and unpredictable prompt size for a fixed top-k. Name the min-merge and max-resplit guardrails.

for a senior

Show that you would instrument the size distribution at ingest and alert on it, since truncation degrades recall silently. Argue for packing the prompt by token budget instead of a fixed k, and mention reranking as a mitigation for noisier first-stage scores.

for a principal

Frame it as a systems-contract problem: chunking made an implicit promise about piece size that retrieval, prompting and cost models all depend on. Own the decision about where those bounds sit and how they trade evidence volume against cost predictability.

## Where the variance comes from A semantic splitter cuts wherever the similarity curve says the topic moved. If a document holds one idea for four pages, that is one chunk; if a speaker fires off three unrelated one-liners, those are three chunks of a sentence each. Nothing in the algorithm bounds the output length, so the size distribution is long-tailed and heavily corpus-dependent. Field ecology notebooks, where entries change site and species mid-paragraph, produce exactly this shape — a mass of short entries plus a few very long uninterrupted observation passages. That is not a bug; coherence was the objective. But every stage after chunking was probably designed assuming roughly uniform pieces, and that mismatch is where production problems appear. ## Failure 1 — oversized chunks and silent truncation Embedding models have a maximum input length. Feed a chunk longer than that and it is typically truncated rather than rejected: the vector then describes the first portion of the chunk while the index stores the whole text. Queries that should match the tail never do, and nobody sees an error. This is the nastiest of the failures because it degrades recall quietly, and it only shows up if you instrument chunk token lengths at ingest. Oversized chunks also distort the prompt. If top-k is 5 and one hit is 3,000 tokens, that single result crowds out four others; on a long-context model it merely costs money, on a tighter budget it means real evidence gets dropped by the packer. Retrieval scores can even look *better* — a huge chunk is more likely to contain the answer somewhere — while answer quality falls because the generator must find a needle in it. ## Failure 2 — undersized chunks with no standalone meaning At the other tail, a 40-token chunk may be a single sentence stranded between two strong breakpoints. Two problems: its embedding is noisy and easily matched by superficially similar queries, and even when it is the *right* passage it lacks the surrounding sentences that make it interpretable. "That figure was revised down after the site visit" retrieves cleanly and answers nothing. Short chunks are also over-represented in results by some scoring schemes, because a short text about exactly one thing has a tight, well-aligned vector. ## Failure 3 — unpredictable capacity per query With a fixed k, evidence volume becomes a lottery. One query returns 400 tokens, another 9,000. That propagates into token cost per query, time-to-first-token, and the risk of overflowing the model's window on the tail queries — the ones you notice last, in production, on the users who ask the broadest questions. ## The fixes **Bound the distribution after splitting.** Post-process the semantic output with two guardrails: merge any chunk below a minimum token count into its neighbour (the neighbour on the side with the smaller boundary distance is the sensible choice), and re-split any chunk above a maximum, falling back to a cheaper length-based split inside it. This keeps the semantic boundaries you found while clipping both tails. It is a couple of dozen lines and it is the single highest-value addition to a naive semantic pipeline. **Never exceed the embedder's input limit.** Make the maximum chunk size strictly smaller than the embedding model's limit, and assert it at ingest rather than trusting the splitter. Log the distribution — median, p95, max, and the count of chunks under 100 tokens — as an ingest metric you can alert on when a new corpus arrives. **Pack the prompt by token budget, not by k.** Retrieve generously, then fill the context up to a token ceiling in rank order, stopping when the next chunk would not fit. This converts unpredictable evidence volume into a fixed budget and makes cost per query stable. **Consider a reranker.** With variable sizes, first-stage vector scores are noisier, and a cross-encoder rerank over a larger candidate set recovers much of the ordering quality before the packing step. **Score short chunks with care.** If tiny chunks keep surfacing and disappointing, either raise the minimum size or attach a little neighbouring context at retrieval time so what reaches the model is interpretable. ## What good looks like in an answer The strong version of this answer names the *silent* failure first — truncation at the embedder — because it is the one that does not announce itself, then gives the guardrail pair (min-merge, max-resplit) and the budget-based packing change. The weak version says only "chunks should be a consistent size", which misses that consistency was traded away deliberately and the job is to bound the tails, not to abandon the method.

  • When you merge an undersized chunk, how do you decide which neighbour it joins?
    Use the boundary distances you already computed: merge across the weaker of the two boundaries, since that is the seam the splitter was least confident about. Merging into the preceding chunk by default is a reasonable simpler rule for narrative text, where a short sentence usually continues the thought before it. Cap the merge so the combined chunk does not cross the maximum size.
  • How would you detect that oversized chunks are being truncated by the embedding model in production?
    Instrument the ingest path: record the token length of every chunk against the embedder's documented input limit and alert when any chunk approaches it. A distribution log — median, p95, max, and count over the limit — catches it the day a new corpus lands. Do not rely on retrieval metrics; truncation degrades recall gradually and looks like ordinary corpus difficulty.
  • Does bounding chunk sizes after semantic splitting defeat the purpose of the technique?
    No. The guardrails only clip the tails: chunks between the minimum and maximum keep exactly the boundaries the similarity curve found, which is the majority of them. What you lose is the extreme cases, where a 3,000-token chunk was going to be truncated anyway and a 40-token chunk was going to retrieve uninterpretably. Bounded semantic chunking still beats a pure length split on unstructured prose.

saying these in an interview costs you the question

  • Assuming the embedding model errors rather than truncates oversized input
  • Treating top-k as a fixed token budget
  • Believing short chunks retrieve poorly, when they often over-retrieve
  • Fixing variance by abandoning semantic chunking entirely
  • Never measuring the chunk-size distribution at ingest

context