Would you run RAGFlow's Compose stack as-is in production, and what would you change?
answer
- Convenient quickstart, not a production topology
- Two workloads competing for one box
- Four stores, three durability answers
- Rebuildable is not the same as cheap
- Managed services buy failover, not speed
basics
~20 sThe stock Compose stack is a fine single-team deployment but it is four stateful services on one host with no redundancy. Before production, separate ingestion capacity from query serving, back up MySQL and object storage, size the document engine deliberately, and put real credentials and TLS in front of it.
solid answer
~60 sAs shipped, `docker compose up -d` gives you an application container, a document engine, MySQL, MinIO and Redis on one box, with default passwords in `docker/.env` and the UI on port 80. That is a good internal deployment and a poor production one. The changes I would argue for: put TLS and real authentication in front, rotate every default credential out of `.env` into a secret store; back up MySQL and MinIO on a schedule, because the doc-engine index is rebuildable only by re-parsing the whole corpus; size the document engine explicitly, since Elasticsearch is the memory hog and needs `vm.max_map_count` raised; and separate the ingestion path from the query path — DeepDoc parsing is CPU-bound and a bulk ingest on a shared box will show up directly in query latency. Beyond a certain scale, the honest answer is to stop running the bundled data services and point RAGFlow at managed MySQL, managed object storage and a managed search cluster, so failover and backups become someone else's operational problem.
go deeper
Know that the stock Compose stack is a single-host deployment with default credentials in docker/.env, and that production needs real secrets, TLS and backups before anyone outside the team uses it.
Explain which stores hold irreplaceable state, why the document-engine index is rebuildable only by re-parsing, and why bulk ingestion competes with query latency on a shared host.
Show you would operate it: separate executor capacity from query serving, size and tune the document engine, pin image tags and rehearse upgrades on staging, take and test restores of MySQL and object storage, and scope API keys per integration.
Own the topology decision end to end — the recovery objective that justifies moving to managed data services, the one-way doors (document engine, embedding model) that must be settled before ingest, and how the ingestion cost profile shapes budget and capacity planning.
## What "as-is" actually gives you The quickstart stack is deliberately convenient: one command, five services, defaults that work on a laptop. Read as a production topology, it is a single-host, single-replica deployment of four stateful systems with shared-secret defaults committed in a file. Nothing about that is wrong for its purpose — it is wrong as an unexamined production choice, and the interview question is whether you can tell the difference. ## Capacity: the two very different workloads RAGFlow has two load profiles that compete for the same box. **Ingestion** is bursty and CPU-heavy. DeepDoc layout analysis, OCR and table-structure recognition are real computer vision work, and the task executor is where it runs. Embedding adds either more CPU or an external API round trip per chunk. A one-off load of ten thousand PDFs will saturate a machine for hours. **Query** is latency-sensitive and comparatively light: a vector plus keyword search in the document engine, optional reranking, then generation — the last of which is usually dominated by the model provider, not by RAGFlow. Run them together on one host and a bulk ingest becomes a visible query-latency incident. The structural fix is to give ingestion its own capacity — additional executor capacity, ideally on separate hosts, drawing from the same Redis queue — so the two scale independently. Where GPUs are available, a GPU-enabled Compose variant accelerates the parsing/OCR path, which is the single biggest ingestion win. ## Memory and the document engine Elasticsearch is the heaviest container in the stack. `docker/.env` exposes a memory limit for it, and the documented host baseline is roughly 4 CPUs and 16 GB RAM — a floor, not a target. On Linux, `vm.max_map_count` must be at least 262144 or the node will not start. If memory is the binding constraint and the corpus is modest, Infinity is the lighter engine, but that choice must be made before ingestion because switching engines means re-parsing everything. Index growth is driven by chunk count and embedding dimensionality, not by the size of the source PDFs, so capacity planning starts from an estimate of chunks per document times documents, not gigabytes of input. ## Durability: what a backup must cover Four stores, three different answers: - **MySQL** — irreplaceable. Tenants, API keys, datasets and their configuration, one record per document, assistants, agent definitions, session history. Take real backups and test a restore. - **MinIO** — irreplaceable in practice. It holds the originals, which are the only input a re-parse can use. Lose it and the corpus cannot be rebuilt from RAGFlow at all. - **Document engine** — derived. Technically rebuildable by re-parsing, which is true and expensive: for a large corpus that is hours of CPU and, on a hosted embedding provider, a real invoice. Snapshot it anyway, and treat "we can rebuild" as a recovery plan with an RTO measured in hours. - **Redis** — transient. Losing it costs re-queuing in-flight parse jobs. And the operational landmine worth saying out loud: `docker compose down -v` removes the volumes and therefore all of the above. It should never be in a production runbook without a very deliberate reason. ## Security posture `docker/.env` ships with default passwords for MySQL, MinIO, Redis and the search engine; those are laptop conveniences. In a real deployment they come from a secret manager, the data services are not published to any interface beyond the Compose network, and the application is fronted by TLS with authentication that matches your organisation's identity story rather than local accounts alone. Tenant API keys are full-authority credentials: issue a separate one per integration so a leak is revocable without breaking ingestion, and keep them out of browser-side code. ## Availability Single replica means planned downtime for upgrades and unplanned downtime for host failure. Whether that is acceptable is a product question — an internal knowledge assistant can tolerate it; a customer-facing support answerer cannot. The path from tolerable to not-tolerable is where you stop running bundled data services: managed MySQL, managed object storage and a managed search cluster convert your redundancy and backup problem into a vendor's, leaving you to run the stateless-ish application container behind a load balancer. That is a real cost increase and the reason to do it is failover, not fashion. ## Upgrades Image tags in `.env` are the upgrade lever. Pin them — never track a moving tag in production — read the release notes for schema and index changes, snapshot MySQL before an upgrade, and rehearse on a staging stack built from the same Compose files. Any release that changes chunking or embedding behaviour raises the question of whether existing datasets need re-parsing to benefit, which is a scheduling decision, not an automatic one. ## How I would answer in the room "For an internal deployment serving one team, the stock stack plus rotated credentials, TLS, and nightly MySQL and object-store backups is proportionate. For anything customer-facing, I would separate ingestion capacity from query serving, pin image tags, and move the four stateful services to managed equivalents so that a host failure is a page rather than an outage — and I would make the document-engine and embedding-model choices before the first real ingest, because both are effectively one-way doors."
- Ingestion of a large backlog is tanking query latency. What is your first structural move?Separate the two workloads. Parsing runs in the task executor, pulling jobs from Redis, so give it its own capacity — additional executor capacity on hosts that are not serving queries — rather than adding API replicas, which are not the bottleneck. Throttle the backlog into batches so it drains over off-peak hours instead of all at once, and if GPUs are available, use the GPU-enabled deployment variant to accelerate the OCR and layout path, which is the dominant cost.
- At what point would you stop running the bundled MySQL, MinIO and search engine?When the availability requirement outgrows a single host. The bundled services have no redundancy and no managed backup story, so a host failure is an outage and every upgrade is planned downtime. Moving to managed database, object storage and search converts failover, patching and point-in-time recovery into a vendor's problem and leaves you operating the application container, which is much easier to run redundantly. It costs more; the justification is the recovery objective, not scale alone.
- How do you plan capacity for the document engine as the corpus grows?Estimate from chunks, not bytes. Index size scales with the number of chunks times embedding dimensionality plus the text and full-text postings, so multiply expected documents by average chunks per document under your chunking configuration and measure the per-chunk footprint on a representative sample. Then size memory for the query working set, since search latency degrades sharply once the index no longer sits comfortably in RAM, and keep headroom for the re-parse scenario where you temporarily hold both an old and a new dataset.
saying these in an interview costs you the question
- Treating the quickstart Compose file as production-ready
- Leaving the default .env passwords in place
- Calling the document-engine index disposable because it is derived
- Adding API replicas to fix slow document parsing
- Tracking a moving image tag in production