Why is fixed-size chunking still the baseline any smarter chunker must beat?
answer
- the control everything else is measured against
- deterministic, linear, no model calls
- re-indexing happens more often than you think
- tune the baseline before replacing it
- end-to-end quality on your own eval set
basics
~20 sBecause it is deterministic, linear-time, model-free and trivially reproducible, so its cost and behaviour are fully predictable across re-indexes. Any alternative adds inference cost, nondeterminism and an extra dependency, and must earn that with measured end-to-end gains.
solid answer
~50 sLength-based splitting has properties that are easy to undervalue until you operate a corpus at scale. It runs in linear time with no model calls, so re-chunking millions of documents costs CPU rather than inference spend. It is deterministic: the same input always yields the same chunks, which means diffable indexes, reproducible evals and cache hits. It has no dependency that can change under you. Anything smarter gives up some of that — per-document model calls, latency during ingestion, results that shift when the helper model is upgraded, and a re-index bill every time you change the method. None of that is disqualifying, but it is a real cost that has to be paid for with measured improvement. The discipline is to hold the fixed-size result as the control, build an eval set of real questions, and require an alternative to win on end-to-end answer quality on your own corpus — not on a published comparison from someone else's.
go deeper
Know that fixed-size and recursive splitting are the standard starting point, and that they are fast, deterministic and require no model calls.
Be able to name the concrete operational properties — linear time, reproducible output, no external dependency — and explain why they make the method a natural control in a comparison.
Describe the experiment you would run: tune the baseline first, build an eval set from real questions, score end to end, and check the margin against the noise before adopting anything more expensive.
Own the economics. Put re-index frequency and per-document ingestion cost on the table before the method decision, and hold the line that chunking changes compete for effort with reranking and query rewriting, which are cheaper to try and need no re-index.
## What the baseline actually gives you Fixed-size and recursive splitting are usually described as the naive options, which understates them. Their properties are the ones a platform team cares about: - **Linear time, no inference.** Chunking a ten-million-document corpus is a CPU job you can parallelise trivially. There is no per-document model call, so ingestion throughput is bounded by I/O rather than by a rate limit. - **Deterministic.** The same document produces byte-identical chunks every run. That makes indexes diffable, makes incremental re-indexing safe (only changed documents change chunks), and makes eval results attributable to the change you made rather than to sampling noise in a helper model. - **No extra dependency.** Nothing in the chunker can be deprecated, rate-limited, or silently upgraded. A method that calls a model inherits that model's lifecycle. - **Debuggable.** When a chunk looks wrong you can point at the exact rule that produced it — the counter ran out here, this separator fired there. Explaining why a model-driven boundary landed where it did is much harder, and the answer is often "it just did". - **Cheap to change.** Because there is no inference cost, sweeping four chunk sizes over an eval corpus is an afternoon, which means the baseline itself is tunable in a way that pays before any method change is considered. ## The costs an alternative brings Every step away from pure length-driven splitting buys quality with operational complexity, and the complexity is usually paid at the worst time — during a re-index. The re-index cost is the one most often missed. Chunking is not a one-off. You re-chunk when you change chunk size, when you change embedding model, when you onboard a new document format, and when a bug is found. If the chunker costs a fraction of a cent per document in inference, that is a rounding error on one document and a serious budget line across a corpus, repeated every time. A CPU-only chunker makes re-indexing a scheduling problem; a model-driven one makes it a procurement conversation. Nondeterminism is the second. If chunk boundaries can shift between runs, two indexes built from the same corpus differ, an A/B comparison has an extra source of variance, and reproducing a reported retrieval bug becomes hard. Fixing temperature and pinning versions mitigates this but does not remove the exposure: the helper model will eventually be deprecated and the replacement will not chunk identically. ## The evaluation discipline The reason this framing matters in an interview is that chunking is a subject where published claims travel further than evidence. The disciplined position is straightforward: 1. **Establish the control properly.** Tune the baseline first — size, overlap, separator ordering for each document format. A badly configured baseline makes any alternative look good, and a surprising share of reported wins are really wins over an untuned control. 2. **Build an eval set from your own corpus.** Real questions users ask, with known answer locations. A few hundred is enough to see effects that matter. 3. **Score end to end.** Retrieval recall is an intermediate metric and frequently disagrees with answer quality. Judge on whether the final answer is right. 4. **Compare against the other things you could have spent the effort on.** Chunking changes compete with reranking, query rewriting, hybrid keyword-plus-vector retrieval, and simply raising k. Those often deliver more per unit of effort, and unlike a chunking change they do not require a full re-index to try. 5. **Require the margin to exceed the noise.** With a small eval set, differences of a couple of points are not signal. ## When spending is justified There are corpora where length-driven splitting has a visible ceiling — where you can look at sampled chunks and see the damage, where the atomic units of meaning genuinely do not align with any length, and where you have already exhausted the cheaper retrieval-side levers. That is a legitimate reason to invest. The distinction is between adopting a more expensive method because it measurably fixed a diagnosed failure, and adopting it because it sounds more sophisticated than counting tokens. ## The organisational angle At scale, chunking policy is shared infrastructure across many document types and teams. The baseline's real advantage there is that it is a policy anyone can reason about, reproduce, and re-run without a budget request. A principal engineer's contribution is usually not choosing the cleverest chunker; it is insisting that whichever chunker is chosen is measured against the cheap one on the organisation's own data, and that the re-index cost of the choice is on the table before the decision, not after.
- A vendor reports a large retrieval gain from their chunking method. What do you ask before adopting it?What baseline it was measured against and on whose corpus. Gains over an untuned fixed-size control are common and usually shrink once size, overlap and separators are configured properly. Then ask for the end-to-end answer-quality delta rather than retrieval recall, and price the per-document ingestion cost multiplied by how often you re-index. Reproduce it on your own eval set before committing.
- Why does re-indexing frequency matter so much to this decision?Because chunking cost is not paid once. You re-chunk on every size change, every embedding-model upgrade, every new document format and every chunker bug fix. A method costing a fraction of a cent per document is negligible on one document and a major line item across millions, repeated several times a year. A CPU-only chunker keeps that a scheduling question rather than a budget one.
- What would convince you the baseline has genuinely hit its ceiling on a corpus?Direct evidence, not intuition: sampled chunks showing that units of meaning are routinely split regardless of size, failed eval questions whose gold passages sit across boundaries, and a size sweep that shows no setting fixes it. Combined with having already tried the cheaper retrieval-side levers — reranking, query rewriting, hybrid search — that is a diagnosed ceiling rather than a hunch.
saying these in an interview costs you the question
- Calls fixed-size chunking obsolete without measuring it
- Compares a new method against an untuned baseline
- Ignores that chunking cost is paid on every re-index
- Judges chunkers on retrieval recall rather than answer quality
- Treats a published benchmark as evidence for their own corpus