skip to content

Deployment & APIs

You will learn what running RAGFlow actually costs you: a Docker Compose stack with a doc engine, database, object store, and cache, plus the HTTP and Python APIs that let another service drive datasets and chats. Interviewers ask about that ops footprint before they ask about features.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

questions

6

What services does the RAGFlow Docker Compose stack run, and what does each store?

level: juniorimportance: must knowfreq 68%

answer

  1. Five moving parts, not one process
  2. One container, two processes inside
  3. Chunks, metadata, blobs, queue
  4. Elasticsearch or Infinity holds the vectors
  5. vm.max_map_count 262144 or it loops

basics

~20 s

RAGFlow ships as a Compose stack: the ragflow server container (web UI plus API, with a task executor beside it), a document engine (Elasticsearch by default, or Infinity) for chunks and vectors, MySQL for metadata, MinIO for the original files, and Redis for the task queue.

solid answer

~50 s

`docker compose -f docker/docker-compose.yml up -d` brings up five moving parts. The **ragflow** container serves the web UI and the `/api/v1` HTTP API behind nginx, and runs the **task executor** that does parsing and embedding. The **document engine** — Elasticsearch by default, Infinity as the lighter alternative, selected by `DOC_ENGINE` in `docker/.env` — holds the chunks, their embedding vectors and the full-text index. **MySQL** holds metadata: tenants and users, knowledge bases, document records, chat assistants, agent definitions and sessions. **MinIO** holds the uploaded originals as blobs. **Redis** carries the parsing task queue and cache between the API server and the executor. The practical consequence is that RAGFlow is not one process you can drop on a small box: the docs ask for roughly 4 CPUs, 16 GB RAM and 50 GB of disk, and Elasticsearch needs `vm.max_map_count` raised to at least 262144.

go deeper

for a junior

Be able to name the four backing services and what each one holds — chunks and vectors in the doc engine, metadata in MySQL, original files in MinIO, task queue in Redis — and know the stack starts from a Compose file in the docker/ directory.

for a middle

Explain the split between the API server and the task executor inside the ragflow container, why parsing failures show up in the executor, and which prerequisites (memory, disk, vm.max_map_count) make the doc engine refuse to start.

for a senior

Show you have operated it: which volumes you back up and how often, what a down -v actually destroys, how you would upgrade image tags without losing state, and how you diagnose a stack where the UI is up but ingestion has stalled.

for a principal

Own the tradeoff of adopting a stack with four stateful dependencies. Be ready to argue whether to run the bundled engines or point RAGFlow at managed MySQL, object storage and search, and what that does to your backup, upgrade and on-call story.

## The shape of a RAGFlow deployment RAGFlow is distributed as a Docker Compose stack rather than a library. The documented start is to clone the repo and run `docker compose -f docker/docker-compose.yml up -d` from the `docker/` directory, with `docker/.env` supplying image tag, ports and passwords. Understanding the stack matters because every operational question about RAGFlow — backups, scaling, upgrades, "where did my data go" — resolves to *which of these four stores holds the thing you care about*. ## The application container The `ragflow` service is the application: a Python backend serving the REST API and an nginx front end serving the built web UI. Its entrypoint starts **two** things — the API server and a **task executor** process. That split matters: the API server is what your HTTP calls hit, while the executor is what pulls parsing jobs off the queue and does the expensive work (layout analysis, OCR, embedding). If parsing appears stuck while the UI is responsive, the executor, not the API, is the thing to look at. By default the container publishes ports 80 and 443, so you reach the UI at `http://<host>`. The backend API port is `SVR_HTTP_PORT` in `.env`, 9380 by default, and SDK clients typically use `http://<host>:9380` as `base_url`. ## The document engine Chunks, their embedding vectors and the full-text index live in the document engine. Elasticsearch is the default; Infinity is the alternative, chosen with `DOC_ENGINE` in `docker/.env`. This is the store the retrieval path actually queries, and it is the memory-hungry one — the `.env` file carries a memory limit for it, and on Linux Elasticsearch refuses to start unless the host kernel setting `vm.max_map_count` is at least 262144 (`sysctl -w vm.max_map_count=262144`, made permanent in `/etc/sysctl.conf`). A first-run failure where the engine container restarts in a loop is almost always this. ## MySQL: metadata MySQL is the system of record for everything that is *not* a chunk or a file: tenants and users, API keys, knowledge-base (dataset) configuration, one row per document with its parsing state and progress, chat assistants and their prompt configuration, agent definitions, and chat sessions with their message history. When an API client passes only a `session_id` and a new question, the prior turns are read from here. ## MinIO: the originals Uploaded files are stored as blobs in MinIO. RAGFlow keeps them because the UI lets you preview a source document and jump to the cited passage, and because re-parsing a document with a different chunking configuration needs the original bytes again. Losing MinIO means losing the ability to re-parse anything. ## Redis: the queue Redis carries the parsing task queue between the API server and the task executor, plus caching. It is the one store whose contents are transient — an in-flight queue, not durable knowledge. ## Image variants and resources The image tag in `.env` picks between the full image and a `-slim` image. The slim image is far smaller because it ships without bundled embedding models, which means you must register an external or self-hosted embedding provider before any dataset can be parsed. The full image bundles embedding models so the box can embed locally. Resource guidance in the docs is roughly 4 CPU cores, 16 GB RAM and 50 GB disk. That is not a formality: Elasticsearch alone will claim several gigabytes, and DeepDoc parsing is CPU-intensive. ## Why interviewers ask The question separates people who ran the quickstart from people who operated it. The follow-ups are the interesting part: which volumes must be backed up (MySQL and MinIO always; the doc-engine index is rebuildable but only by re-parsing everything), what happens on `docker compose down -v` (the `-v` removes volumes and therefore all four stores), and where you would put load if ingestion throughput became the bottleneck (more task-executor capacity, not more API replicas).

  • Which of those stores must a backup cover, and which can you afford to lose?
    MySQL and MinIO are the irreplaceable ones — MySQL holds tenants, datasets, document records, assistants and session history; MinIO holds the original files, without which you cannot re-parse. The document-engine index is derived data: losing it costs a full re-parse of every knowledge base, which is expensive but recoverable. Redis holds an in-flight task queue and can be lost outright; the worst case is re-triggering parsing for jobs that were queued.
  • A colleague runs docker compose down -v to clean up. What did they just destroy?
    `-v` removes the named volumes along with the containers, so it wipes MySQL metadata, MinIO originals, the document-engine index and Redis in one command — every dataset, document, assistant and user on that instance. A plain `docker compose down` stops and removes containers but leaves the volumes, so an `up -d` afterwards comes back with data intact. The `-v` form is only appropriate when you deliberately want a clean slate, such as when switching document engines.
  • Parsing jobs sit at 0% while the UI responds normally. Where do you look first?
    The task executor, not the API server. Both run inside the ragflow container, so check that container's logs for the executor process, then check that Redis is reachable — the queue is how the API hands work to the executor. Also confirm an embedding model is configured, since parsing needs one to embed chunks, and confirm the document engine is healthy, because the executor writes finished chunks there.

saying these in an interview costs you the question

  • Calling RAGFlow a single self-contained container
  • Thinking MinIO stores the embedding vectors
  • Assuming MySQL holds the chunks and the index
  • Believing the web UI process also does the parsing
  • Ignoring the vm.max_map_count prerequisite on Linux

context

open as a page

In RAGFlow's HTTP API, how do you authenticate and take a document from upload to searchable chunks?

level: middleimportance: must knowfreq 72%

basics

~20 s

Generate an API key in the RAGFlow UI and send it as Authorization: Bearer <key> to /api/v1. Create a dataset, upload documents into it, then trigger parsing explicitly — parsing is asynchronous, and a document is only searchable once its run state reaches completion.

open as a page

What does RAGFlow's DOC_ENGINE setting select, and what does switching it cost?

level: middleimportance: should knowfreq 44%

basics

~20 s

DOC_ENGINE in RAGFlow's docker/.env picks the store that holds chunks, embedding vectors and the full-text index — Elasticsearch by default, or Infinity. It is not a live-migratable choice: switching means bringing the stack down with volumes removed and re-parsing every knowledge base.

open as a page

How does an external app call a RAGFlow chat assistant or Agent over HTTP?

level: seniorimportance: should knowfreq 48%

basics

~20 s

RAGFlow exposes three query surfaces: native chat-assistant completions under /api/v1/chats/{chat_id}, Agent completions under /api/v1/agents/{agent_id} for a canvas-built workflow, and OpenAI-compatible endpoints so an existing OpenAI client works by changing base_url. Sessions are server-side, so clients send only the new question.

open as a page

How do you register a model provider in RAGFlow, and what breaks with a self-hosted model in Docker?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Models are registered per tenant on RAGFlow's model-providers page — an API key for a hosted provider, or a base URL for a self-hosted one — and then promoted to defaults in system model settings. The usual failure is pointing a container at localhost, which resolves to the container itself, not the host running the model server.

open as a page

Would you run RAGFlow's Compose stack as-is in production, and what would you change?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

The stock Compose stack is a fine single-team deployment but it is four stateful services on one host with no redundancy. Before production, separate ingestion capacity from query serving, back up MySQL and object storage, size the document engine deliberately, and put real credentials and TLS in front of it.

open as a page