skip to content

Why split Markdown docs on their heading hierarchy instead of a fixed character count?

level: juniorimportance: must knowfreq 62%

answer

  1. use the structure the author already wrote
  2. boundaries come from the document, not the counter
  3. the heading path travels with the chunk
  4. one section, one subject
  5. still needs a length fallback

basics

~20 s

Headings mark where one topic ends and the next begins, so splitting on them yields chunks that each cover one subject. Fixed-length cuts merge unrelated sections and throw away the heading path that says what the chunk is about.

solid answer

~50 s

A fixed window has no idea where a topic starts, so it routinely produces a chunk that is the tail of one section plus the head of the next. That chunk's embedding is an average of two subjects and matches neither question cleanly. Splitting on the heading hierarchy instead uses boundaries the author already drew: one section, one subject. The second win is the heading path. When you split on `## Rate limits` you know the chunk sits under `Billing API > Rate limits`, and you can prepend that path into the chunk text so the embedding and any keyword index see it, as well as store it as a field for filtering and citations. Header splitting is not a complete strategy on its own — sections vary wildly in length, so you still need a rule for sections that overflow your budget and for merging tiny ones.

code

markdown · 15 lines
markdown
# Billing API

## Rate limits

Requests are capped at 100 per minute per key.

### Retry behaviour

On a 429 the response carries a Retry-After header in seconds.

<!-- resulting chunk text, with the heading path prepended -->

Billing API > Rate limits > Retry behaviour

On a 429 the response carries a Retry-After header in seconds.

go deeper

for a junior

Be ready to say plainly that headings are boundaries the author already drew, and that keeping the heading text with the chunk is as valuable as the boundary itself.

for a middle

Explain the two policies a header splitter needs on top of the split: a length fallback for oversized sections and merging for tiny ones, and why the heading path gets both embedded and stored as a field.

for a senior

Show you have handled messy real corpora — HTML boilerplate, OCR'd PDFs, docs where the heading levels are used inconsistently — and describe how you measured whether the new boundaries actually improved top-k recall.

for a principal

Frame chunking as a pipeline decision with a rebuild cost: every strategy change re-embeds the corpus. Argue for deterministic structure-based splitting as the default because it is cheap, reproducible and auditable, and reserve model-driven strategies for corpora that demonstrably need them.

## What structure-aware splitting means Structure-aware chunking uses boundaries the document's author already created — headings, sections, list blocks, table rows, fenced code — instead of counting characters or tokens. In Markdown that means splitting on `#`, `##`, `###` levels; in HTML on `<h1>`–`<h3>`, `<section>` and `<article>` elements; in a wiki or docs-site export on whatever section markup it emits. The splitter parses the markup, walks its tree, and emits one chunk per section together with the chain of headings above that section. It is deterministic and cheap: no model calls, no embeddings, reproducible across runs. That matters because you re-run chunking every time the corpus changes. ## Why length-based cuts hurt retrieval A fixed 800-character window is blind to topic boundaries, and three failure shapes follow. First, straddling. One chunk holds the last two paragraphs of "Rate limits" and the first paragraph of "Error codes". A single embedding vector has to represent both, so it lands somewhere between the two subjects and ranks poorly for either question. Second, splitting a unit that only makes sense whole. A definition ends up in chunk 7 and the example that illustrates it in chunk 8. Retrieval returns one of them and the model answers half the question, confidently. Third, orphaned references. The heading is the sentence that says what the text is about. Cut it away and the chunk reads "this endpoint returns a 429 with a `Retry-After` header" — full of pronouns with no antecedent, and containing none of the words a user would actually search for. ## The heading path is the real prize The boundary is useful; the path is more useful. A header splitter knows, for every chunk, the full chain of ancestors: `Billing API > Rate limits > Retry behaviour`. Do two things with it. Prepend it into the chunk text, so it is part of what gets embedded and part of what a keyword index matches. A chunk whose body never repeats the word "billing" becomes findable by a query about billing. Also store it as a structured field alongside the vector. That lets you filter ("only sections under the v2 API"), route, and — most importantly — show a real citation with a deep link back to the source section rather than an anonymous fragment. ## Where it needs help Headings are irregular, so a header splitter alone will not produce chunks that fit your budget. Two policies close the gap. Oversized sections: when one section exceeds the budget, fall back to a length-based split inside that section only, and repeat the heading path on every resulting piece so none of them becomes an orphan. Undersized sections: a page of one-line `###` entries produces dozens of near-empty chunks whose embeddings are noisy. Merge adjacent siblings under the same parent until they reach a useful size, keeping the shared parent path. And some documents have no usable structure at all — OCR'd scans, meeting transcripts, minified HTML, plain text dumps. There the honest answer is either to recover structure first (layout detection, speaker turns, a heading-inference pass) or to fall back to length-based splitting and accept the cost. ## HTML has an extra trap HTML carries structure and noise in the same tree. Navigation menus, cookie banners, footers and sidebars will be dutifully chunked alongside the content if you split naively, and they are near-identical across every page in the corpus — which means near-duplicate embeddings that crowd out real answers. Strip boilerplate before splitting, and keep tables and code elements intact as units rather than letting a text extractor flatten them into unreadable runs. ## How you would know it is working Measure retrieval, not vibes. Build a small set of real questions with the section that actually answers each one, and track whether that section appears in the top-k. Then eyeball the boundaries: print thirty random chunks and ask whether each one reads as a coherent, self-contained passage about a single subject, and whether you can tell what document and section it came from without being told. If the answer to the second question is no, your prefix is missing, not your splitter. ## The mental model Length-based splitting asks "how much text fits?" Structure-aware splitting asks "what is one idea?" The first is about the embedding model's window; the second is about the document. You need both, in that order: cut on structure, then enforce length inside whatever structure gives you.

  • What do you do when one section under a heading is far larger than your chunk budget?
    Fall back to length-based splitting inside that section only, and repeat the heading path on every piece so each one stays self-describing. Optionally add a marker like "part 2 of 4" so a reader knows the passage continues. The key point is that the fallback is scoped to the offending section rather than applied to the whole document.
  • A page has forty one-line subsections. How would you chunk it?
    Merge adjacent siblings under the same parent until the merged block reaches a useful size, keeping the shared parent path as the prefix. Very short chunks produce noisy, near-duplicate embeddings and inflate the index for no retrieval gain. Merging is the mirror-image policy to splitting oversized sections; a production splitter needs both.
  • Where does header splitting simply not apply?
    Documents with no markup: scanned PDFs, transcripts, plain-text logs, or HTML stripped to raw text. You either recover structure first — layout or heading detection on the PDF, speaker and topic turns on a transcript — or accept length-based splitting there. Pretending the structure exists gives you one giant chunk per file, which is worse than a plain window.

saying these in an interview costs you the question

  • Claims heading splits alone guarantee chunks fit the embedding window
  • Discards the heading text once the boundary is found
  • Assumes every document has usable headings, including scanned PDFs
  • Treats a heading split as also solving retrieval relevance
  • Chunks raw HTML including nav bars and footers as content

context