In RAGFlow, what does DeepDoc parsing do that plain PDF text extraction cannot?
answer
- treats the page as an image
- three vision stages, not one
- knows what a region is
- tables come out as HTML
- works on scans with no text layer
basics
~20 sDeepDoc runs vision models over each page: OCR for scanned text, layout analysis that labels titles, paragraphs, figures, captions and headers, and table-structure recognition that rebuilds tables as HTML. Plain extraction returns only a flat text stream.
solid answer
~50 sDeepDoc is RAGFlow's document-understanding front end, and it treats a page as an **image**, not as a text stream. It runs three vision stages: OCR (text detection plus recognition, so a scanned or photographed page still yields text), layout recognition (a detector that labels regions as text, title, figure, figure caption, table, table caption, header, footer, reference or equation), and table-structure recognition, which infers rows, columns and spans and serializes the table into HTML that survives into the chunk. A plain extractor pulls whatever text objects the PDF happens to contain, in whatever order they were drawn — so two-column papers interleave, running headers and page numbers pollute every chunk, and a table collapses into loose numbers. Because DeepDoc knows *what* each region is, RAGFlow can drop furniture, keep reading order, keep a table with its caption, and hand the chunker structured input. The cost is real: it is neural inference per page, so parsing is minutes-per-document, not milliseconds.
go deeper
Be able to say that RAGFlow parses documents with DeepDoc before indexing them, and that it handles scans through OCR instead of relying on a PDF's text layer.
Name the three stages — OCR, layout recognition, table-structure recognition — and explain concretely what each fixes: unreadable scans, interleaved columns and page furniture, and mangled tables.
Show you have operated it: parse time scales per page, output is silently wrong rather than erroring, and the right habit is to sample-parse and review chunks before loading a whole corpus.
Own the tradeoff of putting a heavy vision pipeline in the ingestion path — one-time cost and hardware footprint against retrieval quality — and know which document classes justify it versus plain extraction.
## Where parsing sits in RAGFlow In RAGFlow, uploading a file to a knowledge base does not index it. Parsing is a separate, explicitly triggered background job, and DeepDoc is what runs during that job before any chunking or embedding happens. The output of parsing is a set of chunks; everything downstream — embeddings, retrieval, the citations shown in chat — is only ever as good as what DeepDoc produced here. This is why RAGFlow markets itself as a "deep document understanding" engine rather than a chain library: the differentiator is the extraction layer, not the orchestration. ## Why plain extraction fails on real documents A PDF is a drawing program, not a document model. Text is a list of glyph-placement instructions with coordinates; there is no notion of "paragraph", "column", or "table". A generic text extractor emits those instructions in the order they appear in the content stream, which produces three recurring failures: - **Reading order** — a two-column paper or a magazine layout comes out with the left and right columns interleaved line by line, so sentences are spliced from unrelated columns. - **Furniture** — running headers, footers, page numbers and watermarks are text objects like any other, so they land in the middle of chunks and get embedded as if they were content. - **Tables** — the cells are just positioned text runs. Extraction gives you the values with no row/column association, so "Region: EMEA, Q3 revenue: 4.1M" becomes a bag of tokens that answers no question correctly. And a scanned document contains no text objects at all. A plain extractor returns an empty string, which silently produces a knowledge base with zero useful chunks. ## The three DeepDoc vision stages **OCR.** DeepDoc detects text regions in the rendered page image and recognizes the characters inside them. This is what makes scanned PDFs, photographs of documents and image files usable at all. It also means DeepDoc does not depend on the file having a text layer, so it behaves the same whether the PDF was exported from a word processor or produced by a scanner. **Layout recognition.** A detection model classifies each region of the page into layout categories — text, title, figure, figure caption, table, table caption, header, footer, reference, equation. Two things follow. First, RAGFlow can *discard* categories that are noise for retrieval (headers, footers, page numbers). Second, it can establish a sensible reading order and keep a caption attached to the figure or table it describes, so the chunk that contains a table also contains the sentence that explains what the table is. **Table-structure recognition (TSR).** For regions labelled as tables, a further model infers the grid: row and column boundaries, header rows, merged/spanning cells. RAGFlow serializes the result as HTML and stores it in the chunk. That is important twice over — an LLM reads an HTML table far more reliably than a whitespace-mangled one, and the table stays a single retrievable unit instead of being cut in half by a character-count splitter. ## What the chunker receives Because of this, RAGFlow's chunk methods are not operating on raw text. They receive typed, ordered regions, which is precisely what lets templates like Paper split at section titles, Presentation split at slides, or Manual split at heading hierarchy. Layout information is the input that makes per-document-type templates possible in the first place. ## What it costs, and when to skip it DeepDoc is neural inference on every page, so parse time scales with page count and is orders of magnitude slower than reading a text layer — a large scanned corpus is an hours-long, one-time job, and it is CPU/GPU-bound rather than I/O-bound. RAGFlow exposes a layout-recognition setting on the General chunk method precisely so you can turn it off and fall back to plain text extraction for clean, single-column, born-digital documents where DeepDoc buys you nothing. ## Failure modes to expect OCR misreads low-resolution scans, unusual fonts and handwriting; layout detection struggles with dense multi-column journals, sidebars and forms; TSR mis-splits borderless tables and deeply nested headers. None of these are announced as errors — you get a plausible-looking chunk with wrong content. That is exactly why RAGFlow pairs parsing with a chunk-review UI: the correct operational habit is to parse a sample, read the chunks, and only then bulk-load the corpus.
- How would you tell, from the parsed chunks alone, that layout recognition mis-handled a two-column PDF?Read a few chunks end to end: interleaved columns show up as sentences that switch subject mid-line, or clauses that never complete. Repeated running headers appearing inside body chunks is the other tell. RAGFlow's chunk view shows the source region on the page image for PDFs, so you can compare a suspicious chunk against where it actually came from before deciding to re-parse or change template.
- When would you deliberately turn layout recognition off in a RAGFlow knowledge base?When documents are born-digital, single-column and clean — plain Markdown, exported text, simple Word files — the text layer is already correct and DeepDoc only adds parse time. The General chunk method lets you switch layout recognition to plain-text extraction for that case. The moment the corpus contains scans, multi-column layouts or tables that matter, turn it back on.
- Why does DeepDoc emit tables as HTML rather than as plain text?HTML preserves the row/column association and spans that table-structure recognition inferred, so a cell's value stays bound to its row and column header. An LLM reading that chunk can answer "what was EMEA in Q3" correctly, whereas a whitespace-flattened table gives it a bag of numbers. It also keeps the table as one coherent unit rather than something a character-count splitter cuts across.
saying these in an interview costs you the question
- Says DeepDoc is just a faster PDF text extractor
- Assumes scanned PDFs work without any OCR step
- Thinks tables survive fine as plain extracted text
- Claims parsing is instant and free per page
- Confuses layout recognition with the chunking template