skip to content

Semantic Chunking

Instead of cutting at a fixed length, semantic chunking looks for topic boundaries: embed consecutive sentences and split where similarity between neighbours drops off.

on this pageshow

questions

4

How does semantic chunking decide where to split a document into chunks?

level: middleimportance: must knowfreq 62%

answer

  1. boundaries from meaning, not length
  2. similarity between neighbouring sentences
  3. a buffer decides how local the cut is
  4. spike in consecutive-window distance
  5. threshold on a distance curve

basics

~20 s

Semantic chunking splits text into sentences, embeds each sentence together with a small window of its neighbours, and cuts wherever similarity between consecutive windows drops sharply — treating that dip as a topic boundary rather than counting characters.

solid answer

~50 s

The pipeline has four steps. First, split the document into sentences — these are the atomic units, and the splitter never cuts inside one. Second, build a **sliding window** around each sentence (typically the sentence plus one or two neighbours on each side) and embed that window, so each position gets a vector describing its local topic rather than one isolated line. Third, walk the sequence and compute a distance (usually cosine distance) between each pair of consecutive window embeddings; this produces a curve of "how much did the topic move here". Fourth, apply a breakpoint rule — split at every position whose distance clears a threshold — and glue the sentences between breakpoints into chunks. The key property is that boundaries come from *content*, not from a length counter, so a 90-minute council-meeting transcript that drifts from zoning to the library budget with no heading still gets cut at the real seam.

code

python · 9 lines
python
def windows(sentences, buffer=1):
    out = []
    for i in range(len(sentences)):
        lo = max(0, i - buffer)
        hi = min(len(sentences), i + buffer + 1)
        out.append(" ".join(sentences[lo:hi]))
    return out

print(windows(["Zoning item closed.", "Next, the library budget.", "Motion carried."]))

go deeper

for a junior

Be able to say that semantic chunking cuts where the meaning changes instead of at a fixed character or token count, and that it uses embeddings of nearby sentences to detect that change.

for a middle

Explain the four mechanical steps — sentence split, windowed embedding, consecutive-pair distance, threshold on the distance curve — and say what the buffer size controls. Interviewers expect you to describe it without hand-waving over how the boundary is actually detected.

for a senior

Show that you know where the heuristic breaks: interleaved topics, gradual drift with no spike, boilerplate that flattens the curve, and bad sentence segmentation on unpunctuated transcripts. Name the extra ingest pass as a real operational cost, not a footnote.

for a principal

Own the framing that chunking sets the ceiling on everything downstream, and that adopting semantic splitting is an evaluation decision rather than a taste one. Be ready to argue when the added ingest complexity is worth a few points of retrieval quality and when it is not.

## The problem it addresses Retrieval quality is capped by chunking. If a chunk mixes two unrelated topics, its embedding is an average of both and matches neither query well — the classic "muddy centroid". If a chunk is cut mid-argument, the retrieved passage answers half the question and the model fills in the rest by guessing. Length-based splitting has no idea where an idea ends; it cuts at whatever offset the counter reaches. Semantic chunking replaces that counter with a *content signal*: it looks for the places where the text stops being about one thing and starts being about another, and cuts there. This matters most for documents with no visible structure to exploit. A 90-minute city-council meeting transcript may run 20,000 words with no headings, no section markers and no blank lines, while the subject moves from a zoning variance to the library budget and back. Field ecology notebooks are similar: entries change site and species mid-paragraph. There is nothing to key on except meaning. ## The algorithm **1. Sentence segmentation.** The document is split into sentences (or another small atomic unit — a clause, a line of dialogue, a speaker turn). Sentences are never broken; every subsequent decision is about *which* sentence boundaries get promoted to chunk boundaries. This alone rules out the mid-sentence truncation that naive length splitting produces. **2. Windowed embedding.** Each sentence is embedded *in context*: the splitter concatenates the sentence with a buffer of neighbours (buffer = 1 means the previous sentence, the sentence and the next sentence) and embeds the concatenation. The buffer size is the knob that decides how *local* the boundary decision is. With no buffer, one short sentence — "Right." or "Any objections?" — is a wildly different vector from its neighbours and looks like a topic change. With a buffer, that noise is smoothed away by the surrounding text. **3. Consecutive-pair distance.** Walk the sequence of window embeddings and compute a similarity or distance between position *i* and position *i+1*. Cosine distance is the usual choice. The output is a one-dimensional series, one number per candidate boundary, that you can literally plot: flat stretches are coherent passages, spikes are candidate topic shifts. **4. Breakpoint selection.** Choose which spikes count. Every spike above the chosen threshold becomes a chunk boundary; the sentences between consecutive breakpoints are concatenated into a chunk. Common rules pick the threshold from the distribution of distances in *this* document rather than from an absolute number, because raw cosine distances are not comparable across models or corpora. ## Why the distances are relative, not absolute A fixed cutoff like "split when cosine distance > 0.25" is nearly useless in practice. Different embedding models place typical prose at very different distance ranges, and a densely technical document has smaller sentence-to-sentence distances than a chatty transcript. The reliable rules are distributional: split at the top *N*th percentile of distances in the document, or at distances more than *k* standard deviations above the document's mean. Both adapt to the corpus automatically; both change the resulting chunk count sharply as you tune them. ## What it produces, and what it costs Output chunks are variable-length by construction — that is the whole point, and also the main operational headache, because a chunk may come out at 40 tokens or at 3,000. Cost is the other consideration: boundary detection embeds every sentence window, and then the resulting chunks are embedded again for the index, so a semantic pipeline pays roughly twice the embedding calls of a length-based one, plus the latency of an extra pass over the corpus at ingest time. ## Failure modes worth naming - **Interleaved topics.** A discussion that alternates between two subjects sentence by sentence has high distance everywhere; the splitter either shatters it into tiny chunks or, at a loose threshold, merges the interleaving into one incoherent chunk. - **Gradual drift.** When a topic changes slowly over twenty sentences there is no spike to find, only a plateau shift, and consecutive-pair distance is blind to it. - **Boilerplate.** Repeated headers, disclaimers or speaker tags inflate similarity and mask real seams. - **Sentence segmentation errors.** Transcripts with no punctuation degrade the atomic units, and everything downstream inherits that. ## The honest framing for an interview Semantic chunking is a heuristic that converts a topic-boundary question into a threshold on a similarity curve. It is a real improvement on unstructured prose and a poor trade on documents that already carry explicit structure or on pipelines whose budget cannot absorb a second embedding pass. Say that plainly; interviewers are more interested in whether you know when it helps than in whether you can recite the four steps.

  • What changes if you set the buffer size to zero and embed each sentence on its own?
    Boundary decisions become maximally local and much noisier. A one-word acknowledgement, a speaker tag or a short aside is a very different vector from its neighbours, so it registers as a topic shift and triggers a split. You get many more chunks, several of them a single sentence long, and the real seams are buried among false ones. Raising the buffer averages the neighbourhood and makes the curve smoother.
  • Consecutive-pair distance misses a topic that drifts gradually over twenty sentences. How would you catch that?
    Compare against a running summary of the chunk so far rather than only the immediate neighbour — for example, distance between the next window and the mean embedding of the accumulated chunk, which drifts as the topic drifts. A max-chunk-size cap is the pragmatic backstop: it forces a cut even when the curve never spikes, which is usually good enough for retrieval.
  • Why not use a single absolute cosine-distance cutoff for all documents?
    Raw distance scales are model-specific and corpus-specific. Dense technical prose sits at much lower sentence-to-sentence distances than conversational transcripts, and a different embedding model shifts the whole distribution. A cutoff tuned on one corpus produces almost no splits on another and shreds a third. Distributional rules — percentile or standard deviations above this document's mean — normalize that away automatically.

It is like chaptering an unedited recording by watching a talk-about-the-same-thing meter: wherever the needle jumps, you drop a chapter marker.

saying these in an interview costs you the question

  • Claiming semantic chunking guarantees equal-sized chunks
  • Saying it cuts at paragraph or heading markers
  • Thinking it splits mid-sentence when similarity dips
  • Believing a fixed cosine cutoff transfers across models
  • Assuming it costs the same as length-based splitting

context

open as a page

Semantic chunking yields chunks from 40 to 3,000 tokens — what breaks downstream?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Variable chunk sizes break the assumptions after retrieval: oversized chunks get truncated by the embedding model and swallow the prompt budget, tiny chunks retrieve without enough context to answer, and a fixed top-k returns wildly different amounts of text per query.

open as a page

How do you choose the breakpoint threshold in semantic chunking, and what does it change?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The threshold is the one real knob: it decides which similarity dips count as boundaries. Pick it distributionally — a percentile or a number of standard deviations over the document's own distances — then sweep candidate values and score end-to-end retrieval, never trust a default.

open as a page

When does semantic chunking not pay for itself compared with cheaper splitting?

level: principalimportance: should knowfreq 38%

basics

~20 s

Semantic chunking earns its keep on long, unstructured prose with no markers to exploit. It usually does not pay on documents that already carry explicit structure, on short documents, on high-churn corpora where the extra ingest pass recurs, or when a measured sweep shows no retrieval gain.

open as a page