In RAGFlow, a parsed chunk has cut a table in half — what do you do next?
answer
- look before you edit
- trace the chunk to its page region
- parsing fault or cutting fault?
- disable is the fast mitigation
- re-parse wipes manual edits
basics
~20 sOpen the document's chunk list and inspect the chunk against the page region it came from, then decide scope: edit or disable that one chunk by hand, or, if the same break repeats across documents, fix the chunk method and re-parse — a re-parse discards manual edits.
solid answer
~50 sRAGFlow deliberately exposes parsed chunks for human review, and that is the right first stop. In the document's chunk view you can read each chunk, and for PDFs see the region of the page it was extracted from, which tells you immediately whether the fault is parsing (DeepDoc mis-detected the table region) or segmentation (the chunk method cut across it). From there you have two kinds of fix. **Local**: edit the chunk's text directly, attach keywords so it matches the queries it should, or disable the chunk so retrieval stops returning it — appropriate for a handful of one-off defects. **Systemic**: change the chunk method or its settings and re-parse the document, appropriate when the same break repeats. The ordering matters, because re-parsing regenerates chunks from scratch and throws away every manual edit and keyword on that document. Fix the template first, then hand-correct the residue.
go deeper
Know that RAGFlow lets you open a parsed document's chunks and read, edit, disable or delete them, and that chunks are what retrieval actually returns.
Explain how to use the source-region view to separate a DeepDoc extraction error from a chunk-method segmentation error, and which fix each one calls for.
Demonstrate the operating discipline: sample-parse and review before bulk ingest, disable bad chunks as immediate mitigation, and know that re-parsing wipes every manual correction.
Decide how much human review the corpus warrants at all — where manual curation is worth its ongoing cost versus where the template or the source documents must be fixed upstream.
## Why a review surface exists at all Most RAG stacks treat ingestion as a black box: you point a loader at files and hope. RAGFlow's position is that document understanding is imperfect and its errors are *silent* — a mis-parsed table produces a plausible chunk containing wrong numbers, not an exception — so the pipeline must be inspectable. The chunk list per document is that inspection point, and using it is a routine operational habit, not an emergency measure. ## What the chunk view tells you For each parsed document you get its chunks as they will be retrieved. For PDFs, a chunk can be traced back to the region of the page image it came from, which is the single most useful diagnostic available: it separates two failures that look identical in the text. - **A parsing failure.** DeepDoc's layout detection missed the table region, or table-structure recognition inferred the wrong grid. The chunk's source region will show the table but the text will be mangled — cells out of order, headers detached. Editing text here is treating a symptom. - **A segmentation failure.** DeepDoc extracted the table correctly, but the chunk method cut it across a boundary, so half the rows are in one chunk and half in the next. The source regions of the two chunks will show adjacent halves of the same table. The fix is the template, not the text. That distinction drives everything else, so make it before you touch anything. ## The local repairs Within the chunk view you can: - **Edit a chunk's content.** Paste the corrected table, rejoin split sentences, delete leaked headers or footers. The edited text is what gets embedded and what will be shown as the citation, so this genuinely fixes retrieval and grounding for that chunk. - **Add keywords to a chunk.** This gives the chunk extra matchable text beyond its body, which is how you rescue a chunk that is correct but phrased nothing like the queries users actually type. - **Enable or disable a chunk.** Disabling takes a chunk out of retrieval without deleting the document. This is the fastest mitigation for a chunk that is confidently wrong and keeps getting cited — you stop the bleeding immediately and fix the cause afterwards. - **Add or delete chunks.** For a small, high-value document, hand-authoring a chunk that states the fact cleanly is legitimate. All of these are per-chunk, manual work. They scale to dozens of fixes, not thousands. ## The systemic repair, and the trap When the same defect appears across many documents, the answer is upstream: a different chunk method (a genre template that respects the document's real boundaries), different settings within the method, or fixing the source file itself. Then re-parse. The trap that catches people: **re-parsing rebuilds a document's chunks from scratch**, so every manual edit, added keyword and disabled flag on that document disappears. Anyone who spends an afternoon hand-correcting chunks and then decides to try a different template loses the afternoon. The correct sequence is always template first, review second, hand-fixes last — and if you do need to change the template afterwards, expect to redo the manual work. ## Building it into a workflow A workable ingestion routine for a real corpus: 1. Upload a small representative sample per document genre. 2. Parse with the candidate chunk method. 3. Review chunks: are they self-contained, are tables and their captions intact, is reading order sane, is page furniture gone? 4. Iterate on the template until the sample is clean. 5. Bulk-parse the rest. 6. Spot-check, and use the local repairs for the remaining one-off defects. The review pass also produces information you cannot get any other way: it tells you which document genres your pipeline genuinely handles and which ones need a different template or a preprocessing step. That is a far better basis for a rollout decision than an end-to-end answer quality impression. ## Related but distinct Chunk review answers "is the indexed content correct". It does not answer "does retrieval find the right chunk for a query" — that is a separate tuning surface. Keep the two diagnoses apart: if the chunk is wrong, no amount of retrieval tuning helps; if the chunk is right and never surfaces, editing it will not help either.
- How do you tell a DeepDoc parsing error from a chunk-method segmentation error?Trace the chunk back to the page region it came from. If the region shows the whole table but the text is scrambled — cells out of order, headers detached — table-structure recognition failed. If two adjacent chunks map to the top and bottom halves of one table and each half reads correctly, extraction was fine and the chunk method cut across it. The first calls for different parsing settings, the second for a different template.
- A single chunk is confidently wrong and keeps getting cited in answers. What is the fastest mitigation?Disable that chunk in the document's chunk view. It stops being retrievable immediately without deleting the document or waiting for a re-parse, which is the difference between minutes and hours on a large file. Then diagnose the cause: if the same defect appears elsewhere, fix the chunk method and re-parse; if it is genuinely a one-off, edit the chunk text and re-enable it.
- Why should manual chunk editing come after settling the chunk method, never before?Because re-parsing a document regenerates its chunks from scratch and discards every manual edit, added keyword and disabled flag on that file. Hand-correcting first and changing the template afterwards throws the work away. Settle the template on a small sample, confirm the chunks are structurally right, and reserve manual edits for the residual defects that no template setting fixes.
saying these in an interview costs you the question
- Hand-edits dozens of chunks before fixing the template
- Assumes manual chunk edits survive a re-parse
- Never opens the chunk view and tunes retrieval instead
- Deletes the whole document to remove one bad chunk
- Treats a systemic parsing defect as a series of one-offs