How does an external app call a RAGFlow chat assistant or Agent over HTTP?
answer
- Three surfaces: chats, agents, openai-compatible
- Sessions live on the server
- Send the question, not the transcript
- References ride along with the native answer
- Compatibility path costs you the citations
basics
~20 sRAGFlow exposes three query surfaces: native chat-assistant completions under /api/v1/chats/{chat_id}, Agent completions under /api/v1/agents/{agent_id} for a canvas-built workflow, and OpenAI-compatible endpoints so an existing OpenAI client works by changing base_url. Sessions are server-side, so clients send only the new question.
solid answer
~50 sThree surfaces, chosen by what you need. **Native chat assistant**: create the assistant with `POST /api/v1/chats` bound to dataset ids, open a session with `POST /api/v1/chats/{chat_id}/sessions`, then `POST /api/v1/chats/{chat_id}/completions`, which streams by default and returns the answer *plus a reference payload* naming the chunks it cited. **Agent**: the same shape one level up — `POST /api/v1/agents/{agent_id}/completions` runs a workflow built on the Agent canvas, so retrieval, branching and tool steps live in RAGFlow rather than in your code. **OpenAI-compatible**: `chats_openai` and `agents_openai` paths accept the standard chat-completions body, which is how you drop RAGFlow into a codebase already written against an OpenAI SDK. The architectural point is that conversation state is RAGFlow's, not yours: with a `session_id`, history lives in MySQL and you post only the new question — the opposite of the stateless OpenAI contract, and the reason a naive port that resends the whole message array behaves oddly.
code
python · 9 linesfrom ragflow_sdk import RAGFlow
rag = RAGFlow(api_key="ragflow-XXXXXXXXXXXX", base_url="http://localhost:9380")
assistant = rag.create_chat("support", dataset_ids=["<dataset_id>"])
session = assistant.create_session()
for msg in session.ask("What is the refund window?", stream=True):
print(msg.content, end="")go deeper
Know that a chat assistant is created and bound to datasets first, then queried through a session, and that RAGFlow also offers OpenAI-compatible endpoints so an existing client works by changing base_url.
Explain the stateful session contract — history stored server-side, client sends only the new question — and that the native response carries a reference payload identifying the chunks behind the answer.
Choose deliberately between the three surfaces for a real integration: citations and streaming through a proxy, separate API keys per integration, session retention and isolation between users, and the latency budget of retrieval plus rerank plus generation.
Own where orchestration lives. Argue the governance tradeoff of canvas-authored agents versus code in your repository, and set the policy for versioning, reviewing and testing flows that a product team can change without a deploy.
## Why there are several surfaces RAGFlow can be consumed at three levels of abstraction, and the choice shapes where your product's logic lives. ### 1. Native chat-assistant API A **chat assistant** is a configured object: it is bound to one or more datasets, carries a prompt configuration, and holds retrieval settings. You create it with `POST /api/v1/chats` (dataset ids in the body), and it lives in MySQL. Conversation goes through a **session**: `POST /api/v1/chats/{chat_id}/sessions` returns a session id, and `POST /api/v1/chats/{chat_id}/completions` takes `question` and `session_id`. Streaming is the default, delivered as server-sent chunks; you can turn it off for a single JSON response. The distinctive part of the response is the **reference** payload: alongside the answer text, RAGFlow returns the chunks that grounded it, with the document they came from. That is what makes citation-backed UI possible in your own front end rather than only in RAGFlow's — and it is a genuine reason to prefer the native API over the compatibility layer. ### 2. Agent API The **Agent canvas** is RAGFlow's visual workflow builder: a DAG of components (a begin step, retrieval, LLM generation, categorisation and branching, message output, iteration, code and tool steps) that you wire up in the UI. Anything built there is addressable over HTTP — `POST /api/v1/agents/{agent_id}/completions`, with a companion sessions endpoint, and an `agents_openai` compatibility path. The tradeoff is authorship. With the chat API, your service owns orchestration and RAGFlow is a retrieval-plus-generation backend. With the Agent API, the orchestration is a canvas artifact edited by whoever has UI access — fast to iterate, and genuinely useful when a non-engineer owns the flow, but it is logic outside your repository, outside code review, and outside your test suite. Teams that adopt agents in production usually end up exporting the agent definition into version control precisely to get that back. Agent begin steps can declare parameters, so an integration supplies typed inputs when it opens a run rather than smuggling everything into the question text. ### 3. OpenAI-compatible endpoints `POST /api/v1/chats_openai/{chat_id}/chat/completions` (and the agent equivalent) accepts the standard chat-completions request shape. Point an existing OpenAI client's `base_url` at it, pass the RAGFlow API key as the key, and an application written against OpenAI now answers from your knowledge base with no client rewrite. This is the fastest path to adoption and the right choice when a third-party tool only speaks the OpenAI protocol. What you give up is the RAGFlow-specific envelope — most importantly the structured reference payload — because the OpenAI response schema has nowhere to put it. If citations are a product requirement, use the native API. ## Session semantics: the thing people get wrong OpenAI's chat completions are stateless: the client owns the transcript and resends it every turn. RAGFlow's native API is the opposite. A session is a server-side object; history is stored in MySQL against the session id, and each call sends only the new question. Consequences: - **The client is thin.** You persist a session id per conversation, not a message array. - **The server is stateful.** Sessions accumulate; a long-lived deployment needs a retention policy, and multi-tenant products need to be sure a session id from user A can never be replayed by user B. - **Ports go wrong quietly.** Code migrated from an OpenAI integration that keeps resending the full history will duplicate context, inflate tokens, and confuse the assistant. - **Context growth is RAGFlow's problem, not yours** — which is convenient until a conversation runs long and you discover you have less control over what stays in the window than you would writing the prompt yourself. ## Choosing between them - Want citations rendered in your own UI, and orchestration in your code? **Native chat API.** - Want the flow authored visually, with branching and tools, by people who are not shipping your service? **Agent API.** - Have an existing OpenAI-shaped client or third-party tool you cannot modify? **OpenAI-compatible endpoint.** ## Operational notes All of these use the same tenant API key as the ingestion endpoints. Give the query integration its own key so it can be revoked without breaking ingestion. Streaming responses need your proxy configured to not buffer, or the token-by-token experience collapses into one slow response. And remember that the answer's latency is retrieval plus optional reranking plus generation — a rerank model configured for quality shows up directly in your p95. ## The interview signal Strong answers name all three surfaces, explain that the OpenAI-compatible path trades the reference payload for drop-in compatibility, and articulate the stateful-session contract as an actual design consequence rather than a detail.
- What do you lose by using the OpenAI-compatible endpoint instead of the native completions API?The RAGFlow-specific response envelope, and with it the structured reference payload that names the chunks and documents behind the answer — the OpenAI response schema has no field for it. You also lose the native session semantics, so the client goes back to owning the transcript. Use the compatibility path when you need an unmodified OpenAI client or a third-party tool to work; use the native path when citations or server-side sessions are product requirements.
- Your team is deciding between building the flow in your service against the chat API and building it on the Agent canvas. How do you frame it?It is a question of where orchestration lives. The canvas is fast to iterate and lets a non-engineer own the flow, but the logic sits outside your repository, outside code review, and outside your test suite, and it can change under you between deploys. Code against the chat API is versioned, reviewable and testable, at the cost of building branching yourself. A common compromise is to use the canvas for exploration and export the agent definition into version control once it stabilises.
- A conversation goes 40 turns and answers start drifting. Where do you look with the native session API?At the session, because history is server-side. You are not resending a transcript, so you have less direct control over what enters the context window — the assistant is accumulating turns in MySQL. Practical remedies are to start a fresh session at natural task boundaries rather than keeping one session per user forever, keep a retention policy so sessions do not grow unbounded, and check the assistant's prompt configuration and retrieval settings, since drift is often retrieval pulling in weakly related chunks as the topic wanders.
saying these in an interview costs you the question
- Resending the whole message history to the native session API
- Assuming the OpenAI-compatible endpoint returns citation references
- Treating an Agent canvas flow as version-controlled application logic
- Thinking each request must carry dataset ids and model settings
- Expecting a non-streaming default from the completions endpoint