In RAGFlow, what does choosing a chunk method per document actually change?
answer
- a template per document genre
- not one global chunk size
- set on the KB, overridable per file
- takes effect only on re-parse
- re-parse rebuilds chunks from scratch
basics
~20 sThe chunk method is a per-document-type parsing template — General, Q&A, Manual, Table, Paper, Book, Laws, Presentation or One — that decides where the document is cut and what a chunk represents. It is a knowledge-base default you can override per file, and changing it requires re-parsing.
solid answer
~50 sRAGFlow does not ask you for one global chunk size; it asks which **chunk method** fits each document. The method is a template that knows the shape of a document type: `Paper` splits an academic PDF at its sections, `Book` and `Laws` split at chapter and article boundaries, `Manual` follows the heading hierarchy, `Presentation` makes each slide a chunk, `Table` makes each spreadsheet row a chunk, `Q&A` makes each question-answer pair a chunk, `One` puts the whole document in a single chunk, and `General` is the fallback that splits by token count and delimiters using DeepDoc's layout output. You set a default method when you create the knowledge base, and you can override it on individual documents in the file list. Changing the method on an already-parsed document does nothing until you parse it again — that run deletes the existing chunks and rebuilds them, so it also discards manual chunk edits.
go deeper
Know that RAGFlow makes you pick a chunk method per document type rather than one global chunk size, and be able to name a few templates and what a chunk becomes under each.
Explain that the method determines what a chunk represents — a row, a slide, an article, a section — and that it is a knowledge-base default overridable per file, applied only when parsing runs.
Show the operational discipline: sample-parse representative documents per genre, read the chunks before bulk upload, and remember that a template change on a large corpus means re-parsing everything.
Argue the design tradeoff — genre templates encode assumptions that fail on off-format documents, so decide where template-driven ingestion beats a single tunable splitter for your corpus mix.
## The design choice behind chunk methods Most RAG stacks expose splitting as parameters: a chunk size, an overlap, maybe a separator list. RAGFlow exposes it as a **choice of template per document type**, and only then as parameters within that template. The reasoning is that the right cut points are a property of the document's genre, not of a number: a statute is cut at articles, a slide deck at slides, a spreadsheet at rows, a paper at sections. Picking the genre gets you most of the way; tuning the token count is the residual. ## The templates and what a chunk becomes - **General** — the default and the broadest in file-type support (office documents, PDF, text, Markdown, HTML, images). It uses DeepDoc's layout output, then combines pieces up to a configured token count, respecting delimiters. This is the closest thing to conventional splitting. - **Q&A** — the source already contains question-answer pairs; each pair becomes one chunk, with the question as the part that matches a user query. Built for FAQ and support-ticket corpora. - **Manual** — assumes a hierarchical, titled user manual and cuts at the section level, so a chunk is a coherent procedure rather than an arbitrary window. - **Table** — for spreadsheets and delimited files: the first row supplies the field names, and each subsequent row becomes a chunk carrying those field names with its values. - **Paper** — academic PDFs, split by section, which keeps abstract, method and results separable rather than smeared across arbitrary windows. - **Book** — long-form documents split at chapter/section granularity with a token cap so one chapter does not become one enormous chunk. - **Laws** — legal texts, cut at the article/clause headings that legal documents reliably carry. - **Presentation** — one chunk per slide, which matches how slide content is authored and read. - **One** — the entire document becomes a single chunk. Useful for short documents that only make sense whole, and for cases where you want the model to see the complete text. The exact set grows between RAGFlow releases, so check the dropdown in the version you are running rather than memorizing a list. ## Where the setting lives A chunk method is chosen when you create a knowledge base and becomes the default for everything uploaded into it. Each document in the file list also carries its own chunk-method column, so a knowledge base holding a mix of PDFs, slide decks and spreadsheets can apply the appropriate template per file. The chunk method also determines which configuration fields you are shown — token count and delimiters for General, and which of the enrichment toggles (auto-keyword, auto-question, RAPTOR) are applicable. ## Changing the method later This is the operational detail interviewers probe. Selecting a different chunk method does not retroactively re-cut anything: parsing is a job, and the new method only takes effect when you run parsing again for that document. That run replaces the document's chunks wholesale. Two consequences follow. First, any manual chunk edits, added keywords or disabled chunks for that document are lost — hand-fixing chunks and then re-parsing is wasted work, so fix the template first and edit only afterwards. Second, re-parsing a large scanned corpus is not free; DeepDoc runs again per page, so a template change on a big knowledge base is an hours-long job, which is a strong argument for sample-parsing a few representative documents before bulk upload. ## Choosing well The practical procedure: group your corpus by genre, not by folder. For each genre, upload two or three representative files, parse with the template that matches the genre, and read the resulting chunks. You are checking that chunks are self-contained (a chunk answers a question without its neighbours), that tables and captions are intact, and that no chunk is a fragment of two unrelated sections. If General with a token count is producing better chunks than the genre template, that is a legitimate finding — the genre templates assume the documents actually follow their genre's conventions, and a "paper" exported from a slide deck does not. ## The common misconception A chunk method is not merely a splitter setting. It changes what a chunk *is* — a row, a slide, a Q&A pair, an article, a section — and therefore what retrieval can return as a unit and what a citation points at. That is why RAGFlow keeps it a first-class per-document choice rather than a hidden parameter.
- You change a document's chunk method after it has already been parsed. What must you do, and what do you lose?You must run parsing again for that document — the new method has no effect until you do. That run discards the document's existing chunks and rebuilds them, so any manual chunk edits, added keywords and disabled chunks for that file are gone. The practical order is: settle the template first, review the output, then make manual corrections.
- A knowledge base holds PDFs, slide decks and spreadsheets together. How do you configure it?Set a sensible default chunk method on the knowledge base, then override per document in the file list: Presentation for the decks, Table for the spreadsheets, General or the matching genre template for the PDFs. The chunk method is a per-document property, so one knowledge base can mix them; you do not need to split the corpus into separate knowledge bases just to vary parsing.
- When is the One chunk method the right choice rather than a defect?When the document is short and only coherent whole — a one-page policy, a short contract, a product blurb — splitting it destroys the context that makes it answerable, and the entire text still fits comfortably in the prompt. It is a poor choice for anything long, because retrieval then returns the whole document and floods the context window regardless of which part was relevant.
saying these in an interview costs you the question
- Treats chunk method as just a chunk-size number
- Assumes changing the method re-cuts existing chunks automatically
- Thinks one method must apply to the whole knowledge base
- Believes manual chunk edits survive a re-parse
- Picks Paper or Laws for documents that lack that structure