skip to content

RAGFlow

You will learn the RAG engine built around deep document understanding: layout-aware DeepDoc parsing, chunk templates chosen per document type, retrieval tuning, grounded chat with citations, and deployment through its APIs. Interviewers ask because messy real documents — tables, scans, multi-column PDFs — are where naive text extraction collapses, and citation grounding is what makes an answer auditable.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

explore

questions

18

What services does the RAGFlow Docker Compose stack run, and what does each store?

level: juniorimportance: must knowfreq 68%

answer

  1. Five moving parts, not one process
  2. One container, two processes inside
  3. Chunks, metadata, blobs, queue
  4. Elasticsearch or Infinity holds the vectors
  5. vm.max_map_count 262144 or it loops

basics

~20 s

RAGFlow ships as a Compose stack: the ragflow server container (web UI plus API, with a task executor beside it), a document engine (Elasticsearch by default, or Infinity) for chunks and vectors, MySQL for metadata, MinIO for the original files, and Redis for the task queue.

solid answer

~50 s

`docker compose -f docker/docker-compose.yml up -d` brings up five moving parts. The **ragflow** container serves the web UI and the `/api/v1` HTTP API behind nginx, and runs the **task executor** that does parsing and embedding. The **document engine** — Elasticsearch by default, Infinity as the lighter alternative, selected by `DOC_ENGINE` in `docker/.env` — holds the chunks, their embedding vectors and the full-text index. **MySQL** holds metadata: tenants and users, knowledge bases, document records, chat assistants, agent definitions and sessions. **MinIO** holds the uploaded originals as blobs. **Redis** carries the parsing task queue and cache between the API server and the executor. The practical consequence is that RAGFlow is not one process you can drop on a small box: the docs ask for roughly 4 CPUs, 16 GB RAM and 50 GB of disk, and Elasticsearch needs `vm.max_map_count` raised to at least 262144.

go deeper

for a junior

Be able to name the four backing services and what each one holds — chunks and vectors in the doc engine, metadata in MySQL, original files in MinIO, task queue in Redis — and know the stack starts from a Compose file in the docker/ directory.

for a middle

Explain the split between the API server and the task executor inside the ragflow container, why parsing failures show up in the executor, and which prerequisites (memory, disk, vm.max_map_count) make the doc engine refuse to start.

for a senior

Show you have operated it: which volumes you back up and how often, what a down -v actually destroys, how you would upgrade image tags without losing state, and how you diagnose a stack where the UI is up but ingestion has stalled.

for a principal

Own the tradeoff of adopting a stack with four stateful dependencies. Be ready to argue whether to run the bundled engines or point RAGFlow at managed MySQL, object storage and search, and what that does to your backup, upgrade and on-call story.

## The shape of a RAGFlow deployment RAGFlow is distributed as a Docker Compose stack rather than a library. The documented start is to clone the repo and run `docker compose -f docker/docker-compose.yml up -d` from the `docker/` directory, with `docker/.env` supplying image tag, ports and passwords. Understanding the stack matters because every operational question about RAGFlow — backups, scaling, upgrades, "where did my data go" — resolves to *which of these four stores holds the thing you care about*. ## The application container The `ragflow` service is the application: a Python backend serving the REST API and an nginx front end serving the built web UI. Its entrypoint starts **two** things — the API server and a **task executor** process. That split matters: the API server is what your HTTP calls hit, while the executor is what pulls parsing jobs off the queue and does the expensive work (layout analysis, OCR, embedding). If parsing appears stuck while the UI is responsive, the executor, not the API, is the thing to look at. By default the container publishes ports 80 and 443, so you reach the UI at `http://<host>`. The backend API port is `SVR_HTTP_PORT` in `.env`, 9380 by default, and SDK clients typically use `http://<host>:9380` as `base_url`. ## The document engine Chunks, their embedding vectors and the full-text index live in the document engine. Elasticsearch is the default; Infinity is the alternative, chosen with `DOC_ENGINE` in `docker/.env`. This is the store the retrieval path actually queries, and it is the memory-hungry one — the `.env` file carries a memory limit for it, and on Linux Elasticsearch refuses to start unless the host kernel setting `vm.max_map_count` is at least 262144 (`sysctl -w vm.max_map_count=262144`, made permanent in `/etc/sysctl.conf`). A first-run failure where the engine container restarts in a loop is almost always this. ## MySQL: metadata MySQL is the system of record for everything that is *not* a chunk or a file: tenants and users, API keys, knowledge-base (dataset) configuration, one row per document with its parsing state and progress, chat assistants and their prompt configuration, agent definitions, and chat sessions with their message history. When an API client passes only a `session_id` and a new question, the prior turns are read from here. ## MinIO: the originals Uploaded files are stored as blobs in MinIO. RAGFlow keeps them because the UI lets you preview a source document and jump to the cited passage, and because re-parsing a document with a different chunking configuration needs the original bytes again. Losing MinIO means losing the ability to re-parse anything. ## Redis: the queue Redis carries the parsing task queue between the API server and the task executor, plus caching. It is the one store whose contents are transient — an in-flight queue, not durable knowledge. ## Image variants and resources The image tag in `.env` picks between the full image and a `-slim` image. The slim image is far smaller because it ships without bundled embedding models, which means you must register an external or self-hosted embedding provider before any dataset can be parsed. The full image bundles embedding models so the box can embed locally. Resource guidance in the docs is roughly 4 CPU cores, 16 GB RAM and 50 GB disk. That is not a formality: Elasticsearch alone will claim several gigabytes, and DeepDoc parsing is CPU-intensive. ## Why interviewers ask The question separates people who ran the quickstart from people who operated it. The follow-ups are the interesting part: which volumes must be backed up (MySQL and MinIO always; the doc-engine index is rebuildable but only by re-parsing everything), what happens on `docker compose down -v` (the `-v` removes volumes and therefore all four stores), and where you would put load if ingestion throughput became the bottleneck (more task-executor capacity, not more API replicas).

  • Which of those stores must a backup cover, and which can you afford to lose?
    MySQL and MinIO are the irreplaceable ones — MySQL holds tenants, datasets, document records, assistants and session history; MinIO holds the original files, without which you cannot re-parse. The document-engine index is derived data: losing it costs a full re-parse of every knowledge base, which is expensive but recoverable. Redis holds an in-flight task queue and can be lost outright; the worst case is re-triggering parsing for jobs that were queued.
  • A colleague runs docker compose down -v to clean up. What did they just destroy?
    `-v` removes the named volumes along with the containers, so it wipes MySQL metadata, MinIO originals, the document-engine index and Redis in one command — every dataset, document, assistant and user on that instance. A plain `docker compose down` stops and removes containers but leaves the volumes, so an `up -d` afterwards comes back with data intact. The `-v` form is only appropriate when you deliberately want a clean slate, such as when switching document engines.
  • Parsing jobs sit at 0% while the UI responds normally. Where do you look first?
    The task executor, not the API server. Both run inside the ragflow container, so check that container's logs for the executor process, then check that Redis is reachable — the queue is how the API hands work to the executor. Also confirm an embedding model is configured, since parsing needs one to embed chunks, and confirm the document engine is healthy, because the executor writes finished chunks there.

saying these in an interview costs you the question

  • Calling RAGFlow a single self-contained container
  • Thinking MinIO stores the embedding vectors
  • Assuming MySQL holds the chunks and the index
  • Believing the web UI process also does the parsing
  • Ignoring the vm.max_map_count prerequisite on Linux

context

open as a page

In RAGFlow, what does choosing a chunk method per document actually change?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The chunk method is a per-document-type parsing template — General, Q&A, Manual, Table, Paper, Book, Laws, Presentation or One — that decides where the document is cut and what a chunk represents. It is a knowledge-base default you can override per file, and changing it requires re-parsing.

open as a page

In a RAGFlow chat assistant's system prompt, what does the {knowledge} variable do?

level: juniorimportance: must knowfreq 68%

basics

~20 s

{knowledge} is the placeholder RAGFlow fills with the retrieved chunks before calling the model. Remove it from the assistant's system prompt and retrieval still runs, but no document text ever reaches the LLM, so answers stop being grounded.

open as a page

In RAGFlow's HTTP API, how do you authenticate and take a document from upload to searchable chunks?

level: middleimportance: must knowfreq 72%

basics

~20 s

Generate an API key in the RAGFlow UI and send it as Authorization: Bearer <key> to /api/v1. Create a dataset, upload documents into it, then trigger parsing explicitly — parsing is asynchronous, and a document is only searchable once its run state reaches completion.

open as a page

In RAGFlow, what does DeepDoc parsing do that plain PDF text extraction cannot?

level: middleimportance: must knowfreq 70%

basics

~20 s

DeepDoc runs vision models over each page: OCR for scanned text, layout analysis that labels titles, paragraphs, figures, captions and headers, and table-structure recognition that rebuilds tables as HTML. Plain extraction returns only a flat text stream.

open as a page

In RAGFlow, what do the similarity threshold and keyword similarity weight control?

level: middleimportance: must knowfreq 75%

basics

~20 s

RAGFlow scores each candidate chunk as a weighted blend of keyword term similarity and vector cosine similarity. The keyword similarity weight sets that blend; the similarity threshold is the minimum blended score a chunk must reach to be kept.

open as a page

A RAGFlow assistant answers from model knowledge, not your documents — how do you debug it?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Bisect retrieval from generation. Replay the question in the knowledge base's retrieval-testing panel using the assistant's exact threshold and weight: if good chunks come back there, the fault is prompt-side; if nothing comes back, it is retrieval-side.

open as a page

What does RAGFlow's DOC_ENGINE setting select, and what does switching it cost?

level: middleimportance: should knowfreq 44%

basics

~20 s

DOC_ENGINE in RAGFlow's docker/.env picks the store that holds chunks, embedding vectors and the full-text index — Elasticsearch by default, or Infinity. It is not a live-migratable choice: switching means bringing the stack down with volumes removed and re-parsing every knowledge base.

open as a page

In RAGFlow, how do the Table and Q&A chunk methods treat a spreadsheet differently?

level: middleimportance: should knowfreq 48%

basics

~20 s

Table treats the first row as field names and turns every data row into one chunk carrying those field names with its values. Q&A expects a question column followed by an answer column and turns each pair into one chunk whose question is what a user query matches.

open as a page

In RAGFlow, how do top_k and top_n differ in a chat assistant's retrieval config?

level: middleimportance: should knowfreq 52%

basics

~20 s

top_k bounds the candidate pool RAGFlow pulls from the indexes and scores, defaulting to 1024. top_n bounds how many surviving chunks are actually injected into the prompt, defaulting to 6. One controls recall and cost, the other context size.

open as a page

How does an external app call a RAGFlow chat assistant or Agent over HTTP?

level: seniorimportance: should knowfreq 48%

basics

~20 s

RAGFlow exposes three query surfaces: native chat-assistant completions under /api/v1/chats/{chat_id}, Agent completions under /api/v1/agents/{agent_id} for a canvas-built workflow, and OpenAI-compatible endpoints so an existing OpenAI client works by changing base_url. Sessions are server-side, so clients send only the new question.

open as a page

How do you register a model provider in RAGFlow, and what breaks with a self-hosted model in Docker?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Models are registered per tenant on RAGFlow's model-providers page — an API key for a hosted provider, or a base URL for a self-hosted one — and then promoted to defaults in system model settings. The usual failure is pointing a container at localhost, which resolves to the container itself, not the host running the model server.

open as a page

In RAGFlow, a parsed chunk has cut a table in half — what do you do next?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Open the document's chunk list and inspect the chunk against the page region it came from, then decide scope: edit or disable that one chunk by hand, or, if the same break repeats across documents, fix the chunk method and re-parse — a re-parse discards manual edits.

open as a page

Parsing 500 scanned PDFs in RAGFlow takes hours — how would you cut that time?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Parse cost scales per page with DeepDoc's OCR, layout and table models, so cut work rather than waiting: keep DeepDoc only for documents that need it, switch clean born-digital files to plain-text extraction, turn off per-chunk LLM enrichment, and run more parsing workers on suitable hardware.

open as a page

In RAGFlow, what constrains a chat assistant that spans multiple knowledge bases?

level: seniorimportance: should knowfreq 42%

basics

~20 s

All knowledge bases attached to one RAGFlow assistant must share an embedding model, since their vectors are compared in a single ranking. Beyond that hard rule, one threshold, one weight and one top-N are applied across corpora whose score distributions differ.

open as a page

When is enabling RAGFlow's RAPTOR or knowledge-graph extraction worth the indexing cost?

level: principalimportance: should knowfreq 32%

basics

~20 s

Both are opt-in LLM passes over your chunks at index time. Enable them when queries need information no single chunk holds — corpus-level summaries or multi-hop links across documents. For fact-lookup corpora they add cost, latency and a new source of wrong content for no gain.

open as a page

Would you run RAGFlow's Compose stack as-is in production, and what would you change?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

The stock Compose stack is a fine single-team deployment but it is four stateful services on one host with no redundancy. Before production, separate ingestion capacity from query serving, back up MySQL and object storage, size the document engine deliberately, and put real credentials and TLS in front of it.

open as a page

How would you turn RAGFlow's retrieval-testing panel into a real evaluation loop?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Replace ad-hoc queries with a fixed question set whose correct chunks are known, replay it after every configuration or parsing change, change one variable at a time, and record the settings alongside the results so a regression is attributable.

open as a page