Parsing 500 scanned PDFs in RAGFlow takes hours — how would you cut that time?
answer
- cost is per page, not per file
- find out which files really need OCR
- enrichment adds an LLM call per chunk
- parsing is background workers you can scale
- sample first, bulk-parse once
basics
~20 sParse cost scales per page with DeepDoc's OCR, layout and table models, so cut work rather than waiting: keep DeepDoc only for documents that need it, switch clean born-digital files to plain-text extraction, turn off per-chunk LLM enrichment, and run more parsing workers on suitable hardware.
solid answer
~50 sFirst establish where the time goes. DeepDoc runs neural inference on **every page** — OCR, layout detection, table-structure recognition — so parse time tracks page count, not file count, and scanned pages are the expensive case because OCR cannot be skipped. Then reduce work in three directions. **Scope**: segment the corpus, and only keep full DeepDoc for documents that actually need it; clean single-column born-digital PDFs can use the General method with layout recognition switched to plain-text extraction, which is orders of magnitude faster. **Enrichment**: the auto-keyword and auto-question options add an LLM call per chunk, and RAPTOR and knowledge-graph extraction add further LLM passes over the whole corpus — leave them off unless you have measured the retrieval gain. **Capacity**: parsing runs as background task-executor jobs, so add workers and give the vision models better hardware. Finally, treat it as a one-time batch: sample-parse a few documents, verify the chunks, and only then run the bulk job.
go deeper
Know that RAGFlow parsing is a background job that can take minutes per document, and that scanned pages are slow because every page goes through OCR.
Explain that cost scales per page across DeepDoc's vision models, and that optional enrichment such as auto-keyword extraction adds an LLM call per chunk on top.
Show a diagnosis-first approach: measure where time goes, segment the corpus so only documents that need DeepDoc get it, scale parse workers, and validate on a sample before the bulk run.
Own the ingestion budget as a design decision — one-time heavy parsing versus ongoing retrieval quality, steady-state latency for new documents, and what hardware that commits you to.
## Understand the cost model before optimizing RAGFlow's ingestion is a background pipeline: parse (DeepDoc) → chunk (the selected chunk method) → optional enrichment → embed → index. For a scanned corpus the dominant term is almost always the first one. DeepDoc's per-page work is neural inference — text detection and recognition, layout region detection, and table-structure recognition for regions labelled as tables — so a 500-document corpus that averages 40 pages is 20,000 inference passes, several models deep. Nothing about the chunk method or the embedding model changes that number. The practical implication is that "tune the chunk size" is the wrong first move. Measure the page count, watch the per-document progress that RAGFlow reports during parsing, and confirm whether documents are stuck in parsing or in embedding before changing anything. ## Lever 1 — do not run DeepDoc where it buys nothing The biggest win is scoping. Split the corpus by how the files were produced: - **Scanned or photographed pages** genuinely need OCR. There is no shortcut; the text does not exist otherwise. - **Born-digital, single-column, table-free PDFs, Word files, Markdown and plain text** already carry a correct text layer. The General chunk method exposes a layout-recognition setting you can switch to plain-text extraction, skipping the vision stack entirely for these files. - **Born-digital but complex** — multi-column papers, documents whose value is in tables — should keep DeepDoc, because the failure mode of plain extraction on those is silent garbage rather than a visible error. Because chunk method and its settings are per-document in RAGFlow, you can apply this split inside one knowledge base rather than building separate ones. ## Lever 2 — turn off enrichment you have not justified Several optional features add work *on top of* parsing, and each is charged per chunk or per corpus: - **Auto-keyword** and **auto-question** extraction call the configured LLM for every chunk. On a corpus of hundreds of thousands of chunks this dwarfs parsing in both time and money. - **RAPTOR** clusters chunks and generates summaries with the LLM, recursively, after chunking. - **Knowledge-graph extraction** runs entity and relationship extraction with the LLM across chunks. All three can improve retrieval for some corpora. None should be on during a first bulk ingest that you have not yet validated, because if the chunk method turns out to be wrong you pay for them again on the re-parse. ## Lever 3 — capacity and hardware Parsing is executed by background task-executor workers, so throughput scales with how many are running and what hardware backs them. Raising worker count converts a serial queue into a parallel one; the constraint is CPU (or accelerator) contention, because every worker is running the same heavy vision models. Give the vision models accelerated hardware if you have it available — OCR and layout detection are exactly the workloads that benefit. Watch memory too: several concurrent parsers each holding model weights and page images is a real footprint. ## Lever 4 — sequence the work The most expensive mistake is parsing 500 documents with the wrong settings and doing it twice. Re-parsing is not incremental; it re-runs DeepDoc for the document from scratch. So: 1. Pick five to ten documents that represent the genres in the corpus. 2. Parse them with the intended chunk method and settings. 3. Read the resulting chunks in the chunk-review UI — are tables intact, is reading order sane, do chunks stand alone? 4. Only then queue the bulk job, and let it run as an overnight batch. Also decide up front what the steady state looks like. If new documents arrive daily, the relevant metric is not the initial backlog but per-document latency from upload to searchable, and the backlog can be drained once at low priority. ## What not to do Do not disable layout recognition globally to make the numbers look better; on scans that produces empty or near-empty chunks and a knowledge base that silently answers nothing. Do not shrink the chunk token count to "reduce work" — chunking is cheap and the token count is a retrieval-quality decision, not a throughput one. And do not conclude the system is broken because parsing takes minutes per document; that is the deliberate trade RAGFlow makes — heavy one-time ingestion in exchange for structure-aware chunks and citable, layout-grounded retrieval.
- How do you decide which documents can safely skip DeepDoc layout recognition?Check how the file was produced and what it contains. Born-digital, single-column documents with no meaningful tables already have a correct text layer, so plain-text extraction gives the same chunks far faster. Keep DeepDoc for anything scanned, multi-column, or table-heavy, where plain extraction fails silently rather than loudly. Validate by parsing one file each way and comparing the chunks.
- Why is turning on auto-keyword and auto-question extraction during a first bulk ingest a bad idea?Each adds an LLM call per chunk, so on a large corpus the cost and time exceed parsing itself — and you pay it again if the chunk method turns out to be wrong and you re-parse. Validate chunk quality with enrichment off, confirm retrieval is limited by matching rather than by parsing, and only then enable it and measure whether it actually improves results.
- Parsing is slow but CPU sits idle. Where would you look?At concurrency and queueing rather than at model cost: how many task-executor workers are consuming the parse queue, and whether documents are actually queued behind one another. Also confirm the stall is in parsing at all — the per-document progress reporting distinguishes parsing from embedding, and a slow or rate-limited embedding endpoint produces the same end-to-end symptom with idle local CPU.
saying these in an interview costs you the question
- Reduces chunk token size hoping parsing gets faster
- Disables layout recognition globally, including for scans
- Enables RAPTOR and auto-keywords during first bulk ingest
- Assumes parse time scales with file count, not page count
- Re-parses the whole corpus before validating on a sample