skip to content

How do you register a model provider in RAGFlow, and what breaks with a self-hosted model in Docker?

level: seniorimportance: should knowfreq 55%

answer

  1. Per-tenant registry, not per-request keys
  2. Register, then promote to system defaults
  3. Slim image ships no embedding models
  4. Embedding choice is sticky; chat model is not
  5. localhost inside a container means the container

basics

~20 s

Models are registered per tenant on RAGFlow's model-providers page — an API key for a hosted provider, or a base URL for a self-hosted one — and then promoted to defaults in system model settings. The usual failure is pointing a container at localhost, which resolves to the container itself, not the host running the model server.

solid answer

~50 s

RAGFlow does not read model credentials from the chat request; it keeps a per-tenant provider registry. You add a provider in the UI (OpenAI, DeepSeek and similar hosted vendors take an API key; Ollama, Xinference and other local servers take a base URL), then set the defaults — chat model, embedding model, rerank model — in system model settings, since a dataset needs an embedding model before it can parse anything. Two deployment facts bite. First, the `-slim` image ships without bundled embedding models, so on slim you *must* register an external or self-hosted embedding model or every parse fails. Second, `http://localhost:11434` entered into the UI is resolved inside the RAGFlow container, where it means the container itself: point at `host.docker.internal` on Docker Desktop or the host's LAN address on Linux, and make sure the model server listens on all interfaces rather than loopback. For scripted deployments, `service_conf.yaml` can carry a default LLM factory and key so new tenants start pre-configured.

code

bash · 5 lines
bash
# On the host: make Ollama reachable from other network namespaces
OLLAMA_HOST=0.0.0.0 ollama serve

# Verify from inside the RAGFlow container before touching the UI
docker exec -it ragflow-server curl -s http://host.docker.internal:11434/api/tags

go deeper

for a junior

Know that models are configured once in RAGFlow's provider settings — an API key for hosted vendors, a base URL for local servers — and that a default chat and embedding model must be selected before a dataset can be used.

for a middle

Explain the three model roles and why the embedding model is the sticky one, and diagnose the localhost base-URL failure by reasoning that the URL is resolved from inside the RAGFlow container.

for a senior

Show operational ownership: pick the embedding model and image variant before ingest, pre-seed provider configuration for reproducible deployments, verify reachability from inside the container, and recognise a bulk-parse failure caused by provider rate limits.

for a principal

Own the model-sourcing strategy — hosted versus self-hosted for cost, latency and data residency, the migration cost of ever changing the embedding model, and how provider credentials are stored, scoped and rotated across environments.

## The registry model RAGFlow separates *which models exist* from *which models a knowledge base or assistant uses*. Credentials are registered once per tenant on the model-providers page. For a hosted vendor you paste an API key; for a self-hosted server you give a base URL and a model name. RAGFlow ships integrations for a long list of providers — OpenAI-compatible endpoints, major Chinese vendors, and local runtimes such as Ollama and Xinference — but the shape is always the same: a provider entry, then one or more usable models under it. After registration comes the step people forget: **system model settings**, where you nominate the tenant's default chat model, default embedding model, and default rerank model (plus image-to-text and speech-to-text where relevant). Registering a provider does not make it the default; a dataset created before a default embedding model exists cannot parse. ## Model roles, and why embedding is special Three roles matter operationally: - **Chat model** — generates the grounded answer. Swappable at any time; you can change an assistant's model between questions and nothing else breaks. - **Embedding model** — turns chunks into vectors at parse time and turns the query into a vector at retrieval time. This one is *sticky*. A knowledge base's vectors were produced by a specific model with a specific dimensionality; changing the model invalidates them, so the dataset must be re-parsed. RAGFlow also requires that datasets combined in one assistant agree on their embedding model, because otherwise their vectors are not in a comparable space. - **Rerank model** — an optional second-stage scorer over retrieved candidates. Cheap to add or remove, but it costs latency per query. The sticky one is the deployment decision: pick your embedding model with the same care as `DOC_ENGINE`, before you ingest a real corpus. ## The slim image trap RAGFlow publishes a full image and a `-slim` image, and `docker/.env` defaults to slim in recent releases because the full image is several times larger. The difference is that the full image bundles embedding models so RAGFlow can embed locally with no external dependency; the slim image does not. On slim, until you register an external embedding provider and set it as the system default, every parse job fails — and the symptom is documents stuck in a failed state, not an obvious "no model configured" error at upload time. Knowing this converts a confusing first-run failure into a thirty-second fix. ## Networking a self-hosted model This is the highest-frequency real-world failure. Someone runs Ollama on their laptop, opens RAGFlow in a browser on the same laptop, enters `http://localhost:11434`, and gets a connection error. The base URL is not fetched by the browser — it is fetched by the RAGFlow **server process, inside its container**. Inside that container, `localhost` is the container's own loopback interface, where nothing is listening. The fixes: - **Docker Desktop (macOS/Windows):** use `http://host.docker.internal:11434`, the special name that resolves to the host. - **Linux:** use the host's LAN or bridge address (`http://192.168.x.x:11434`), or add a host-gateway mapping. - **Either way:** the model server must listen on all interfaces, not loopback. For Ollama that means `OLLAMA_HOST=0.0.0.0` on the host; a server bound to 127.0.0.1 is unreachable from a container regardless of the address you use. - **Same-stack alternative:** if the model server runs as another service in the same Compose project, address it by its service name on the shared network, which sidesteps host networking entirely. The general rule worth stating in an interview: *every URL you type into a containerised app's configuration is resolved from inside the container.* The same reasoning applies to a self-hosted OpenAI-compatible gateway or a corporate proxy. ## Pre-seeding defaults for automated deployments Clicking through the UI does not fit a reproducible deployment. `service_conf.yaml` — rendered from its template at container start — can carry a default LLM factory, API key and base URL that newly registered tenants inherit, so a freshly provisioned instance comes up with working models instead of a checklist for whoever logs in first. That belongs in your secret management alongside the MySQL and MinIO passwords in `docker/.env`, not in the image. ## Failure modes to be able to name - Parsing fails on a slim image because no embedding model is set. - A self-hosted base URL points at `localhost` and cannot be reached from the container. - Someone changes a dataset's embedding model and retrieval quality collapses, because old vectors and new query vectors are from different models. - An assistant spans two datasets embedded with different models and is rejected or answers poorly. - A hosted provider key hits a rate limit during a bulk parse, and documents fail in batches — visible only in per-document state. ## The interview signal A weak answer stops at "you paste an API key in the settings". A strong one covers the three model roles and which is sticky, the slim-image dependency, and the container-networking reasoning behind the `localhost` failure — that last point is the one that shows the candidate has actually deployed this rather than watched a demo.

  • A team wants to swap the embedding model on a knowledge base that already has 5,000 parsed documents. What do you tell them?
    That the stored vectors were produced by the old model and are not comparable with queries embedded by the new one, so the change requires re-parsing every document in that dataset — hours of DeepDoc work and, on a hosted provider, a real embedding bill. Plan it as a migration with a window, keep the old dataset serving until the new one is parsed and validated in the retrieval-testing panel, then cut assistants over. Chat models can be swapped freely; embedding models cannot.
  • How would you provision model configuration for ten RAGFlow instances without clicking through each UI?
    Pre-seed it. The rendered service_conf.yaml can carry a default LLM factory, API key and base URL that new tenants inherit, and the surrounding values come from docker/.env, so both files are what your deployment tooling templates from a secret store. Anything left to the UI is a manual step that will drift between instances; anything baked into the image is a leaked credential. Keep the keys in the deployment secret manager and render the config at start.
  • Documents fail to parse right after a fresh install and the logs mention no embedding model. What is the likely cause?
    The stack is running the -slim image, which ships without bundled embedding models, and no external embedding provider has been registered and promoted in system model settings. Either register a hosted or self-hosted embedding model and set it as the default, or switch the image tag in docker/.env to the full image, which bundles embedding models locally. The tell is that upload succeeds and only parsing fails, because embedding happens in the task executor.

saying these in an interview costs you the question

  • Passing an OpenAI key in the chat request instead of registering it
  • Assuming registration alone makes a model the default
  • Entering localhost as the base URL for a host-run model server
  • Thinking the embedding model can be swapped without re-parsing
  • Not knowing the slim image lacks bundled embedding models

context