What does RAGFlow's DOC_ENGINE setting select, and what does switching it cost?
answer
- An .env value, not a runtime flag
- It picks where chunks and vectors live
- No migration path between the two
- The real bill is re-parsing
- down -v takes everything, not just the index
basics
~20 sDOC_ENGINE in RAGFlow's docker/.env picks the store that holds chunks, embedding vectors and the full-text index — Elasticsearch by default, or Infinity. It is not a live-migratable choice: switching means bringing the stack down with volumes removed and re-parsing every knowledge base.
solid answer
~50 s`DOC_ENGINE` lives in `docker/.env` and is read into the generated `service_conf.yaml`, selecting which backend RAGFlow uses for chunk storage, vector search and full-text search. `elasticsearch` is the default and best-tested path; `infinity` is the lighter-weight alternative aimed at lower memory use and faster hybrid search. What trips people is that this is a *storage* choice, not a runtime flag. There is no migration path: the documented switch is to stop the stack, remove the volumes (`docker compose down -v`), set the new value, and start again — which destroys every dataset's chunks along with the rest of the stack's state. In practice that means you decide the engine before you ingest anything real, or you accept re-parsing your entire corpus, which for DeepDoc-parsed documents is the expensive operation. Feature parity also lags: Elasticsearch is the default for a reason, so check the release notes for your tag before committing to Infinity.
code
ini · 8 lines# docker/.env
DOC_ENGINE=elasticsearch
ES_PORT=1200
ELASTIC_PASSWORD=infini_rag_flow
MYSQL_PASSWORD=infini_rag_flow
MINIO_PASSWORD=infini_rag_flow
SVR_HTTP_PORT=9380
RAGFLOW_IMAGE=infiniflow/ragflow:v0.20.5-slimgo deeper
Know that DOC_ENGINE lives in docker/.env, defaults to elasticsearch, and chooses where parsed chunks and their vectors are stored — not the chunking method and not the embedding model.
Explain that the setting is a storage commitment: the two engines have incompatible formats, the documented switch removes volumes, and every knowledge base has to be parsed again afterwards.
Show you would treat this as a one-way door on a populated instance — test Infinity in a parallel stack, protect MinIO because re-parsing is the recovery path, and re-validate retrieval behaviour after any engine change.
Own the decision at deployment-design time: weigh Elasticsearch's maturity and operational familiarity against Infinity's lower footprint, and account for the fact that reversing the choice costs a full re-ingest of the corpus.
## What the setting is `DOC_ENGINE` is an environment variable in `docker/.env`, the file that parameterises RAGFlow's Compose stack. On container start the entrypoint renders `service_conf.yaml` from its template, substituting environment values, and the rendered config tells the backend which document-engine block to use. Compose also uses the value to decide which engine container to bring up, so setting it changes both the client and the server side in one place. The accepted values are `elasticsearch` (the default) and `infinity`; recent builds have added further options, so read the `.env` comments for your tag rather than trusting a list from memory. ## What the engine actually stores This is the crux. The document engine is where a parsed document *ends up*: one record per chunk, containing the chunk text, its embedding vector, and the fields used for keyword matching. Every retrieval call — the retrieval-testing panel, a chat assistant answering a question, an agent's retrieval step — queries this store. MySQL knows a document exists and what state it is in; MinIO has the original bytes; but the searchable content only exists in the document engine. That is why the switch is destructive. The two engines have entirely different on-disk formats and indices. Nothing copies your Elasticsearch indices into Infinity. ## Why the two options exist Elasticsearch is mature, ubiquitous, and gives RAGFlow a well-understood hybrid of BM25 keyword scoring and vector similarity. Its cost is memory: it is the heaviest container in the stack, it wants gigabytes of heap, and on Linux the host must have `vm.max_map_count` at 262144 or higher or the node refuses to boot. Infinity is InfiniFlow's own database, built for exactly this workload — dense vectors plus sparse/full-text retrieval in one engine — and its pitch is lower resource consumption and faster hybrid queries. The tradeoff is maturity and reach: it is a younger project, platform support is narrower, and features occasionally land on the Elasticsearch path first. For a production deployment, "Elasticsearch unless you have measured a reason" is a defensible default; for a laptop demo or a memory-constrained VM, Infinity is a real win. ## The switching procedure and its cost The documented sequence is: stop the stack; remove the volumes so the old engine's data and the stale references to it are gone; change `DOC_ENGINE`; start again. `docker compose down -v` is the command usually cited, and the `-v` is the dangerous half — it removes MySQL and MinIO volumes too, so you lose datasets, documents, assistants and users, not just the index. That is worth stating plainly in an interview, because the cost is not the downtime. It is the **re-parse**. RAGFlow's value proposition is layout-aware DeepDoc parsing with OCR and table-structure recognition; running that over a corpus of thousands of PDFs is hours of CPU, and if you are calling a hosted embedding API it is a real bill as well. So the engine choice is effectively a one-way door for a populated instance. ## Operational implications A few things follow: - **Decide early.** Make the engine choice part of the initial deployment decision, alongside the image variant and the embedding model — the other two settings that are also painful to change after ingestion. - **Keep the originals.** Because a re-parse is your recovery path for any doc-engine loss, MinIO's contents are the thing that actually makes the corpus reconstructible. Back it up even if you consider the index disposable. - **Don't treat the index as disposable in practice.** "It's derived data, we can rebuild it" is true and useless if rebuilding takes eight hours and blocks the product. - **Separate the config change from the data wipe.** If you are experimenting, run a second stack with its own project name and volumes rather than flipping the value on the instance that holds real data. - **Verify after the switch.** After re-parsing, use the retrieval-testing panel to confirm hybrid retrieval behaves as it did before; scoring characteristics between two different engines are not identical, so previously tuned thresholds may need revisiting. ## The interview signal A weak answer treats `DOC_ENGINE` as a swappable driver. A strong one identifies it as a storage-format commitment, names re-parsing as the true cost, and notes that `docker compose down -v` takes the rest of the stack with it — plus the instinct to make this decision before the first real ingest rather than after.
- Why can't RAGFlow just migrate the existing index when you flip the value?Because the two engines are different databases with different index structures, scoring and APIs — there is no common export format for chunk records plus vectors plus full-text postings. RAGFlow instead treats the document engine as derived data whose source of truth is the original file in MinIO and the document record in MySQL, so the sanctioned rebuild path is to re-parse. That is simple and correct, but it puts the whole DeepDoc cost back on you.
- You want to try Infinity without risking the production corpus. How?Run a second, isolated stack rather than flipping the setting in place: a separate copy of the docker directory with its own Compose project name, its own `.env` with `DOC_ENGINE=infinity`, and different published ports so volumes and containers do not collide. Ingest a representative slice of documents, then compare retrieval quality and latency against the production instance before deciding. Never test this by editing the live `.env`.
- Which other early settings are as expensive to change after ingestion?The embedding model of a knowledge base — changing it invalidates every stored vector, so the dataset must be re-parsed, and RAGFlow constrains assistants to datasets that share an embedding model. The image variant matters too: moving from the slim image to the full one (or the reverse) changes whether embedding runs locally or against an external provider. Both belong in the same up-front decision as DOC_ENGINE.
saying these in an interview costs you the question
- Calling it a hot-swappable driver you can flip anytime
- Assuming RAGFlow migrates indices between engines
- Thinking only the vector index is lost, not the corpus state
- Believing DOC_ENGINE selects the embedding model
- Running down -v on a populated instance to 'just try' Infinity