In RAGFlow's HTTP API, how do you authenticate and take a document from upload to searchable chunks?
answer
- Bearer token on /api/v1
- Three calls, not one
- Upload stores; it does not index
- The parse trigger returns before the work does
- Poll the per-document run state
basics
~20 sGenerate an API key in the RAGFlow UI and send it as Authorization: Bearer <key> to /api/v1. Create a dataset, upload documents into it, then trigger parsing explicitly — parsing is asynchronous, and a document is only searchable once its run state reaches completion.
solid answer
~50 sRAGFlow exposes a REST API under `/api/v1` on the backend port (9380 by default), authenticated with a tenant API key created in the UI and sent as `Authorization: Bearer <key>`. The lifecycle is three distinct calls, and conflating them is the classic bug. `POST /api/v1/datasets` creates the knowledge base. `POST /api/v1/datasets/{dataset_id}/documents` uploads files as multipart — this only *stores* them. Parsing is a separate trigger, `POST /api/v1/datasets/{dataset_id}/chunks` with the document ids, and it is asynchronous: the request returns immediately while the task executor does layout analysis and embedding in the background. You then poll the document list and watch each document's run state and progress until it completes; only then do retrieval and chat see the content. The Python SDK (`pip install ragflow-sdk`) mirrors this exactly — `create_dataset`, `upload_documents`, `async_parse_documents` — and the `async_` prefix on that last call is the API telling you what it is.
code
python · 11 linesfrom ragflow_sdk import RAGFlow
rag = RAGFlow(api_key="ragflow-XXXXXXXXXXXX", base_url="http://localhost:9380")
dataset = rag.create_dataset(name="handbook")
with open("policy.pdf", "rb") as f:
dataset.upload_documents([{"display_name": "policy.pdf", "blob": f.read()}])
docs = dataset.list_documents()
dataset.async_parse_documents([d.id for d in docs]) # returns immediatelygo deeper
Know the request shape: an API key from the UI sent as an Authorization bearer header against /api/v1, and the order create dataset, upload documents, trigger parsing.
Explain that parsing is a separate asynchronous step handled by the task executor, that the parse endpoint returns before any chunk exists, and that you poll per-document run state and progress before querying.
Demonstrate ingestion as an engineered job: idempotent re-runs keyed on stored document ids, batched parse triggers, backoff polling with timeouts, alerting on documents stuck in a failure state, and API keys scoped and rotatable per integration.
Frame it as a pipeline contract between systems — who owns the source of truth for documents, how deletions propagate, what SLA an ingest has to meet, and how you detect that a chunk of the corpus silently stopped being indexed.
## Authentication RAGFlow's API is keyed per tenant. You create a key from the UI (the API page in your account settings), and every request carries it as a standard bearer token: `Authorization: Bearer ragflow-XXXXXXXXXXXX`. There is no separate handshake, no OAuth dance, no per-request signing. Because the key carries the tenant's full authority over datasets and assistants, it is a secret on the level of a database password: keep it out of source, out of front-end code, and out of logs, and give an external integration its own key so it can be revoked independently. The base URL is the backend HTTP port — `SVR_HTTP_PORT` in `docker/.env`, 9380 by default — so `http://<host>:9380/api/v1/...`. Requests can also arrive through the nginx front end on port 80. ## The three-step lifecycle ### 1. Create the dataset `POST /api/v1/datasets` with a JSON body carrying at minimum a name. A dataset ("knowledge base" in the UI) is the container that owns a chunking configuration and an embedding model. Both of those are effectively fixed once documents are parsed into it, so an integration that creates datasets programmatically should set them at creation time rather than expecting to change them later. ### 2. Upload the files `POST /api/v1/datasets/{dataset_id}/documents` as `multipart/form-data`. This writes the bytes into object storage and creates a document record in MySQL. **It does not index anything.** A document that has been uploaded and never parsed is invisible to every retrieval path — no chunks exist for it yet. This is the single most common integration mistake: a script uploads a hundred PDFs, immediately fires a question at the assistant, gets "I don't know", and the author concludes retrieval is broken. ### 3. Trigger parsing, then wait `POST /api/v1/datasets/{dataset_id}/chunks` with the list of `document_ids` starts parsing. The endpoint name is the giveaway: you are asking RAGFlow to *produce chunks* for those documents. The HTTP call returns as soon as the jobs are queued; the actual work happens in the task executor process, which pulls from Redis, runs layout analysis and OCR, splits according to the dataset's chunking configuration, embeds each chunk, and writes the results into the document engine. Because it is asynchronous, your client must poll. Listing documents in the dataset returns, per document, a run state and a progress value. A robust integration loops with a backoff until every document reaches the completed state, and treats the failure state as a real outcome to surface — a corrupt PDF or an unconfigured embedding model shows up here, not in the response to the parse call. ## The Python SDK `pip install ragflow-sdk` gives a thin typed wrapper over the same endpoints: - `RAGFlow(api_key=..., base_url=...)` — the client - `rag.create_dataset(name=...)` / `rag.list_datasets()` - `dataset.upload_documents([{"display_name": ..., "blob": ...}])` - `dataset.list_documents()` - `dataset.async_parse_documents([doc_id, ...])` The SDK offers no synchronous parse that blocks until chunks exist — the naming is honest about the model. Downstream, `rag.create_chat(...)`, `chat.create_session()` and `session.ask(...)` give you the query side. ## Chunk-level access Beyond the document level, the API exposes individual chunks: you can list a document's chunks, add a chunk by hand, update a chunk's content or keywords, and delete chunks. That is how you build a curation workflow — a reviewer fixes a badly split table, or injects a canonical answer as a chunk — without going through the UI. It is also how you verify programmatically that parsing produced something sane before you let an assistant serve from that dataset. ## Designing an ingestion integration Practical guidance for a service that feeds RAGFlow: - **Make ingestion a job, not a request.** Upload plus parse plus poll can take minutes for a large PDF. Do not do it inside a user-facing HTTP handler. - **Store the returned document id.** It is your handle for polling, re-parsing after a config change, chunk inspection and deletion; deriving it later by matching filenames is fragile. - **Be idempotent.** Re-running an ingest should not silently create a second copy of the same document; check the dataset's document list first. - **Batch the parse trigger.** One call with many document ids queues them together rather than issuing one request per file. - **Alert on the failure state.** Documents that never finish parsing are a silent quality regression: retrieval keeps working, just without that content. ## The interview signal What separates a good answer is the asynchrony. Anyone can read an endpoint list; the person who has shipped this says "upload does not index, parsing is a separate asynchronous trigger, and you poll per-document state" — and then names the failure mode of ignoring it.
- Your script uploads files and immediately asks a question, and the assistant answers 'I don't know'. What happened?The documents were stored but never parsed, or parsing was still running. Upload only writes the file and creates a document record; chunks — the only thing retrieval can see — are produced by the separate parse trigger, which is asynchronous. The fix is to call the parse endpoint with the document ids and then poll each document's run state and progress until it completes before issuing any query. Treat a failed parse state as an error your ingest job reports, not something to ignore.
- How would you make a nightly sync of a document folder into RAGFlow robust?Run it as a background job, not a request handler. Keep a local mapping from source file to RAGFlow document id so re-runs update rather than duplicate; upload only changed files; batch the parse trigger into one call with many document ids; poll with backoff and a timeout; and surface documents that end in the failure state as an alert. Delete document ids whose source file disappeared, so the index and the source stay in step.
- Why does the API expose chunk-level create, update and delete?So curation can be automated. Real corpora produce bad chunks — a table split down the middle, a header glued to the wrong section — and the chunk endpoints let a reviewer or a script fix the text or keywords in place without re-parsing the whole document. They also let you inject a canonical, hand-written chunk for a question the source documents answer badly, and they give an integration a way to assert that parsing produced sensible content before that dataset serves traffic.
saying these in an interview costs you the question
- Assuming uploading a document makes it searchable
- Expecting the parse call to block until chunks exist
- Querying immediately after upload and blaming retrieval
- Putting a full ingest inside a user-facing request handler
- Embedding the tenant API key in front-end code