How do you chunk API reference docs so code blocks and tables stay whole?
answer
- some spans may never be cut
- half an example is not an example
- the header row carries the meaning
- parse the code, don't measure it
- caption stays inside the chunk text
basics
~20 sSplit on the document's own structure, not on length. Treat each fenced code block, each table together with its caption and header row, and each function or class in a source file as an atomic unit the splitter is forbidden to cut.
solid answer
~50 sTake a Python SDK's reference docs as the case. Split first on the `###` headings so each symbol gets its own section, then apply format-native rules inside: a fenced code block is atomic, a table is atomic with its caption and header row, and for the source files themselves you parse the AST and use function and class node spans as boundaries. Length-based splitting destroys exactly the parts a developer needs — half an example is not an example, and a table cut mid-body leaves rows whose columns have no header, so `0.7 | 40 | true` becomes unreadable to both the embedding model and the LLM. Then you need a policy for atomic units bigger than the budget: for a table, repeat the caption and header row on each row group; for a giant function, split at logical blocks and prepend the signature and enclosing class name to each piece.
code
python · 13 linesimport ast
def split_by_top_level_defs(source: str) -> list[str]:
"""Chunk a Python file at function and class boundaries."""
tree = ast.parse(source)
lines = source.splitlines(keepends=True)
chunks = []
for node in tree.body:
start = node.lineno - 1
end = node.end_lineno
chunks.append("".join(lines[start:end]))
return chunksgo deeper
Know that some parts of a document must not be cut — code blocks, tables with their headers — and that a splitter should detect them from the markup rather than guessing by length.
Explain the atomic-unit model and the parser-based alternative for source files, and give the fallback policies for a table or a function that is bigger than the chunk budget.
Show that you have run this over a real mixed corpus: how you sequenced heading splitting, atomic-unit detection and packing, what metadata the parser gave you, and how you measured that broken snippets stopped appearing in answers.
Own the build-versus-buy call. Format-native splitters mean maintaining parsers per format and re-running the whole index when the rules change; argue for the smallest rule set that removes the failures you can actually measure, rather than a splitter per file type.
## Atomic units The useful reframing is that structure-aware chunking is not just about where to cut but about where you are *forbidden* to cut. Some spans of a document only carry meaning whole. A fenced code block. A table with its header row. A function body with its signature. A numbered procedure. A JSON example. Declare those atomic, and let length-based logic operate only in the space between them. This inverts the usual mental model. Instead of "walk forward 800 tokens and look for a nearby separator", you parse the document into a tree of units, mark the atomic ones, and pack the rest around them. ## Code blocks in prose documents Markdown and HTML both mark code explicitly — fenced blocks or `<pre><code>` elements — so a splitter can detect them reliably without heuristics. A cut through one produces two chunks that are each individually wrong: the first ends mid-expression, the second starts with a dangling closing brace and no context about what is being demonstrated. The practical damage in a reference-docs corpus is specific. Developers ask questions whose ideal answer is a runnable snippet. Retrieval returns the half that happens to match the query terms, the generator patches over the missing half, and the user gets code that does not run — the worst kind of RAG failure, because it looks authoritative. Keep the code with its surrounding prose where you can: the sentence before a block usually says what it demonstrates, and the sentence after usually states the caveat. A chunk containing prose + block + caveat is a far better retrieval unit than the block alone. ## Source files: parse, don't split When the corpus contains actual source files rather than prose about them, the right splitter is a parser. Language parsers expose node spans — a function definition, a class body, a method — and those spans are your boundaries. Python's standard library exposes this directly through its `ast` module; every mainstream language has an equivalent, and general-purpose incremental parsers cover many languages with one interface. AST boundaries buy you three things. Chunks that are syntactically complete. Names you can attach as metadata — module, class, function, decorators — which are exactly the terms a developer searches for. And the ability to attach a stable prefix: `module payments.client, class BillingClient` on every method chunk, so a method named `create` is not indistinguishable from the twelve other `create` methods in the corpus. When a single function exceeds the budget, split at logical block boundaries inside it and prepend the signature, the enclosing class name and the relevant imports to each piece. The pieces stay ungrammatical as code, but they stay *identifiable*, which is what retrieval needs. ## Tables Tables are the format where length-based splitting fails most obviously. A table's meaning lives in the relationship between the header row and each body row; cut the header off and every subsequent row is a tuple of unlabelled values. The narrative around a financial statement can be split normally, but the statement itself should stay whole with its caption, because the caption carries the unit, the period and the entity. For tables that exceed the budget there are two workable policies. Row grouping: emit chunks of N rows, each repeating the caption and header row. Or row serialization: turn each row into a sentence-like record (`Q3 2025 — revenue: 412M; operating margin: 11.2%`) so that each unit is self-describing regardless of where it lands. Serialization retrieves better for lookup questions; keeping the grid retrieves better for questions about the shape of the data. Choose based on the questions you actually see. Also keep the caption *inside* the chunk text, not only in metadata — otherwise the numbers are indexed without the words anyone would search for. ## Sequencing the rules A workable pipeline for a docs corpus: strip boilerplate, split on headings, detect atomic units inside each section, pack non-atomic prose up to the budget without crossing an atomic boundary, apply the oversized-unit policies, then prefix every chunk with its heading path and source identifiers. Deterministic, no model calls, cheap to re-run. ## How you would know it is broken The diagnostic is direct: sample chunks and count how many contain a syntactically incomplete code block or a table fragment with no header. That number should be zero. Then track whether the answers your system gives contain runnable code, since that is the outcome the whole rule set exists to protect.
- A single table is larger than the chunk budget. How do you split it without destroying it?Two options. Emit row groups, repeating the caption and header row on every group so columns stay labelled. Or serialize each row into a self-describing record — "Q3 2025, revenue 412M, margin 11.2%" — which retrieves well for point lookups but loses the shape of the grid. Pick based on whether your users ask lookup questions or comparison questions.
- You have one 900-line function that exceeds the budget. What now?Split at logical blocks inside the body and prepend the same header to every piece: the module, the enclosing class, the function signature and the imports it depends on. The pieces are no longer valid code, but they remain identifiable and retrievable. In practice a function that size is also a signal that the ideal chunk is a summary of it plus a pointer to the file.
- Why not just wrap code blocks in a separator token and keep using a recursive length splitter?Because a length splitter still cuts when the block alone exceeds the window, and separators do not tell it that the block is atomic — only that it is a preferred break point. You also lose the metadata a parser gives you for free: module, class and function names, which are the terms developers actually search for.
saying these in an interview costs you the question
- Says code is just text, so a generic splitter is fine
- Splits tables mid-body and loses the header row
- Keeps a code block but drops the prose that explains it
- Assumes an AST splitter never produces oversized chunks
- Stores the table caption only as metadata, never in the chunk text