skip to content

AI agent & RAG frameworks

5 roadmaps258 questionsupdated

You will learn the framework layer that sits above raw model APIs: agent orchestration, multi-agent workflows, and retrieval-augmented generation pipelines, and how the major libraries differ in the abstractions they impose. Interviewers ask here because choosing a framework commits you to a control-flow model, and they want you to compare options rather than recite one vendor's docs.

on this pageshow

guide

overview

~1 min

Agent and RAG frameworks are the libraries that sit between a raw model API and a working application: they wrap model calls, turn your functions into tools, run the loop that lets a model act, and build the pipelines that feed it your documents. Interviewers ask about them because picking one commits a team to a control-flow model. They want to hear you compare options on the axes that matter (who decides the next step, where state lives, what a run costs, how you would test it) and say where you would drop down to a plain client instead. Reciting one framework's API is the weaker answer. The sections fall into three groups. [LangChain](/topics/ai-langchain) supplies the shared vocabulary of prompts, model wrappers, tools and retrievers, and [LangGraph](/topics/ai-langgraph) adds the stateful, looping orchestration a linear chain cannot express. The multi-agent frameworks, [CrewAI](/topics/ai-crewai) and [AutoGen](/topics/ai-autogen), make the questions of delegation, termination and runaway cost concrete, while [Google ADK](/topics/ai-google-adk) and [Koog](/topics/ai-koog) bring the same ideas to a vendor toolkit with built-in evaluation and to the Kotlin and JVM stack. The retrieval-first frameworks, [LlamaIndex](/topics/ai-llamaindex), [Haystack](/topics/ai-haystack) and [RAGFlow](/topics/ai-ragflow), are organised around the data pipeline: parsing, chunking, indexing, retrieval and grounded answers. Start with LangChain's building blocks, because every other section reuses the same nouns. Then learn one RAG pipeline end to end, then graph orchestration, and only then multi-agent designs, which combine all three. Junior questions ask what a component does; senior and principal ones ask which framework you would choose, what it hides, and what breaks in production. You rarely need every framework in depth: know the one your target role uses well, and one contrasting design well enough to argue the difference.

primer

### Every agent runs the same loop Under the branding, each agent framework drives one cycle: send messages and tool definitions to the model, execute whatever tool calls come back, return the results, and repeat until the model answers or a limit stops it. Frameworks differ in who owns that loop: a built-in agent, a graph you draw, a crew or team, or a pipeline you wire by hand. Recognising the loop lets you compare products instead of memorising them. ### A tool is a schema the model reads The model never sees your code. It sees a name, a description and a typed parameter list, and chooses tools from that text alone. Docstrings and parameter descriptions therefore behave like code: a vague one produces wrong calls, and a missing registration makes the tool invisible. ### RAG is two pipelines, and the first one decides quality Indexing runs ahead of time: load, parse, split, embed, store. Querying runs per request: retrieve, optionally rerank, then generate from what was found. Many bad answers are born in indexing (broken table extraction, chunks cut mid-thought, a mismatched embedding model), so a good diagnosis inspects the retrieved text before blaming the model. ### Prompt text shapes behaviour; code enforces it Roles, goals, backstories, instructions and descriptions all end up as prompt text, which the model may ignore. Guarantees come from mechanisms: schema-enforced structured output, which tools an agent is given, fixed edges in a graph, a step limit, a sandbox. Interviewers listen for whether you know which of the two you are relying on. ### State is a deployment decision Conversation history, graph state, session data and vector indexes often default to process memory. Persisting them is what makes a run resumable after a crash, pausable for human approval, and cheap to restart without re-embedding a corpus. ### Every extra agent is extra model calls Delegation, supervision and group chat each add calls, latency, tokens and nondeterminism. A multi-agent design has to earn that cost with genuinely separate roles or context, and needs an explicit way to stop.

Agent loop
The repeated cycle of model call, tool execution and feeding results back, which ends when the model produces a final answer or a limit halts it.
Tool schema
The name, description and typed parameters a framework sends to the model for each tool; the only information the model has when choosing and filling a call.
Structured output
A response constrained to a declared schema, ideally enforced by the provider's API rather than requested in prompt text and parsed afterwards.
Retrieval-augmented generation
Answering from documents fetched at query time and placed in the prompt, rather than from what the model memorised during training.
Chunking
Splitting parsed documents into retrievable pieces; size, overlap and boundaries decide whether a retrieved piece carries enough context to answer.
Embedding
A vector representation of text produced by a model, used to find passages whose meaning is close to a query's.
Vector store
A store that holds embeddings with their text and metadata and returns the nearest neighbours of a query vector.
Hybrid retrieval
Combining lexical keyword scoring with vector similarity, so exact terms such as names and codes are found alongside paraphrases.
Reranker
A second-stage model that rescores a retriever's candidates against the query, trading extra latency for a more precise top few.
State graph
An orchestration model where nodes transform a shared typed state and edges, fixed or conditional, decide what runs next, loops included.
Checkpoint
A persisted snapshot of a run's state, keyed to a conversation or thread, that lets the run resume, be inspected or be rewound.
Human-in-the-loop
A deliberate pause in an agent run for a person to approve, edit or reject a step before execution continues.
Termination condition
An explicit rule, such as a keyword, a message count or a token budget, that ends a multi-agent conversation.
Grounding
Tying each claim in an answer to retrieved source text, usually with citations, so a reader can audit where it came from.
Model Context Protocol
An open protocol through which an agent discovers and calls tools served by a separate process, so one tool server can serve several frameworks.

Read the sections as layers of one stack rather than nine rival products; most frameworks cover several layers and differ in which one they treat as central. ### The layers - **Model access.** A wrapper or executor that hides provider differences: LangChain's chat models, Koog's prompt executors, ADK's model integration, AutoGen's model clients. Chat-style message lists, tool calling and structured output all live here. - **Components.** Tools, retrievers, splitters, embedders, rankers and generators. LangChain, LlamaIndex and Haystack each ship a large catalogue; what differs is how explicitly they are connected. - **Orchestration.** How steps are sequenced: LangChain's composable runnables for linear flows, LangGraph and Koog's strategy graphs for loops and branches, Haystack's validated component wiring, CrewAI's crews and flows, AutoGen's teams, ADK's model-driven agents beside its fixed workflow agents. - **State and runtime.** Checkpointers, session services, memory stores and persisted indexes, which decide what survives a restart and whether a run can pause for a person. - **Evaluation and deployment.** Tracing, streamed events, retrieval metrics and serving. ADK builds this into its tooling, LlamaIndex has an evaluation section, and RAGFlow ships as a deployable engine with its own services, UI and HTTP API rather than a library you import. ### The recurring split Across every orchestration layer the same choice appears: let the model decide the next step, or fix it in code. LangGraph's conditional edges, CrewAI's sequential process and Flows, ADK's workflow agents, Koog's strategy graphs and Haystack pipelines are the fixed-flow side; supervisors, hierarchical managers, model-selected speakers and free-running agents are the model-driven side. Many production designs mix them, with a deterministic outer flow and model decisions only at the points that need judgement. ### How RAG and agents meet Retrieval usually enters an agent as one more tool, or as a fixed step before generation. The retrieval-first frameworks own that step in depth; the orchestration frameworks mainly decide when it runs and what happens with its results.

  1. LangChain →

    Models, prompts, tools and retrievers are the nouns every other framework reuses; learn them once here.

  2. LlamaIndex →

    Walks the whole retrieval pipeline from loading to evaluation, so chunking and index choice stop being abstract.

  3. LangGraph →

    Adds what a linear chain lacks: shared state, loops, conditional routing, checkpoints and pauses for human approval.

  4. CrewAI →

    Makes multi-agent design concrete: roles, task contracts, sequential versus managed delegation, and what each extra agent costs.

  5. AutoGen →

    Covers conversational teams, termination rules and sandboxed code execution, the safety questions agents raise.

  6. Haystack →

    Shows explicitly wired, type-checked pipelines, the production counterpoint to convenience-first composition.

  • Treating an agent's role, backstory or instruction as a guarantee; it is prompt text the model can ignore, so enforce limits with tools, schemas and code.

  • Shipping an in-memory default for sessions, checkpoints or vectors, then discovering that every restart wipes conversations and forces a full re-embed.

  • Running a multi-agent team with no termination condition or step limit; agents talking to agents can loop and keep billing until someone notices.

  • Executing model-generated code directly on the host instead of in a container or sandbox with no access to your files, secrets or network.

  • Blaming the model for a bad RAG answer without reading the retrieved chunks; parsing, chunking and retrieval are where many of these failures start.

  • Writing a vague tool description or untyped parameters; the model selects and fills tools from that text alone.

  • Asking for JSON in the prompt and parsing hopefully when the provider offers schema-enforced structured output.

  • Reaching for a supervisor or hierarchical manager when the step order is already known; a fixed flow is cheaper, faster and testable.

  • Changing the embedding model or chunk settings without rebuilding the index, leaving stored vectors that no longer match new queries.

  • Answering “which framework?” with one vendor's feature list instead of comparing control flow, state handling, cost and when you would use a raw client.

The same handful of decisions sits under most senior questions in this hub; naming the one you are making is usually the answer being looked for. - **Framework versus raw client.** A framework gives you the tool loop, retries, persistence and integrations for free, and costs you abstraction layers to debug and upgrades to track. For a single model call with one tool, a plain client is often clearer; for loops, pauses and persistence, a framework usually pays for itself. - **Model-driven versus fixed control flow.** Letting the model route is flexible for open-ended tasks and costs calls, latency and repeatability. Fixed edges, sequential processes and workflow agents are cheap and testable but only work when you already know the order. - **One agent with many tools versus several agents.** Splitting helps when roles need different context, instructions or permissions. It hurts when agents mostly relay messages, since every hop adds a call and a chance to drift. - **Convenience versus explicit wiring.** Automatic coercion and defaults make prototypes short; explicit, validated connections catch mistakes before a model is ever called and make the graph inspectable. - **Native tool calling versus text-parsed actions.** Native calling needs a capable model and is robust; parsing reasoning written as text works with more models, at a cost in tokens and reliability. - **Keyword, vector or hybrid retrieval.** Vectors find paraphrases, keywords find exact identifiers; blending them and reranking raises precision at the price of latency and tuning. - **Library versus deployable engine.** An importable library fits into your service; an engine with its own services and UI gives more out of the box and becomes infrastructure you operate.

A few shapes recur across the frameworks under different names. Spotting them is how you transfer what you know from one section to the next. - **The tool-calling loop.** Model, tools, results, model again, until a final answer or a limit. Built-in agents, strategy graphs and agent components are all packagings of it. - **Supervisor and workers.** One coordinating model chooses which specialist acts next and receives control back after each turn; it appears as a supervisor node, a hierarchical manager or an agent that transfers control. - **Shared-conversation team.** Several agents take turns on one message history, chosen in a fixed order or by a model, and stop on an explicit condition. - **Index, then query.** An offline pipeline builds the store; an online pipeline retrieves from it and generates. Separating them is what makes re-indexing and evaluation possible. - **Retrieve wide, rerank narrow.** A cheap retriever fetches many candidates and a stronger scorer keeps the best few for the prompt. - **Pause, persist, resume.** A run checkpoints its state, waits for a person or an external event, and continues later on the same thread or session. - **Stream the events.** Consumers read intermediate messages, tool calls and state deltas as they happen instead of waiting for the final result, which is how progress UIs and traces are built.

explore

report an issue with this guide →

questions

258 · 9 sections

In LangChain, what does the @tool decorator expose to the model from a function?

level: juniorimportance: must knowfreq 80%
basics
~20 s

The @tool decorator wraps a Python function as a Tool: the function name becomes the tool name, the docstring becomes the description, and the type hints become a JSON argument schema. The model sees only those three things.

open as a page

In LangChain, what do a Runnable's invoke, batch and stream methods each do?

level: juniorimportance: must knowfreq 78%
basics
~20 s

invoke runs the runnable once on one input and returns the final output. batch runs many inputs, concurrently by default. stream yields output in chunks as it is produced. Each has an async twin: ainvoke, abatch, astream.

open as a page

In LangChain, what does a BaseChatMessageHistory store, and when does in-memory storage break?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A chat message history holds one conversation's ordered messages, keyed by a session id, through the messages property plus add_messages and clear. InMemoryChatMessageHistory keeps that list in process RAM, so a restart or a second replica loses the conversation.

open as a page

In LangChain, how does a ChatModel differ from the older text-completion LLM interface?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A ChatModel takes a list of role-tagged messages and returns an AIMessage object; the LLM interface takes a plain string and returns a plain string. Tool calling, structured output and multimodal content exist only on the ChatModel path.

open as a page

In LangChain, when do you use PromptTemplate versus ChatPromptTemplate?

level: juniorimportance: must knowfreq 78%
basics
~20 s

PromptTemplate formats one string with declared variables and suits plain text-completion models. ChatPromptTemplate formats an ordered list of role-tagged messages (system, human, ai) and is what you use with chat models, which is nearly always today.

open as a page

In LangGraph, what steps build a minimal StateGraph from schema to invoke?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Define a state schema, usually a TypedDict; create StateGraph(State); register functions with add_node; wire them with add_edge from START to END; then call compile() to get a runnable graph you invoke with an initial state dict.

open as a page

In LangGraph, what does graph.stream() give you that graph.invoke() does not?

level: juniorimportance: must knowfreq 70%
basics
~10 s

graph.invoke() blocks and returns only the final state. graph.stream() yields a chunk as the graph advances, one per superstep, so a UI can show progress long before the run finishes.

open as a page

In LangGraph, what does compiling with a checkpointer add, and what is thread_id for?

level: middleimportance: must knowfreq 78%
basics
~20 s

A checkpointer makes LangGraph save a snapshot of the graph's state after every super-step. The thread_id you pass in config names the conversation those snapshots belong to, so the next invoke resumes that thread's accumulated state instead of starting empty.

open as a page

In LangGraph, what does interrupt() do inside a node and how do you resume?

level: middleimportance: must knowfreq 72%
basics
~20 s

interrupt() pauses the graph in the middle of a node, persists a checkpoint, and surfaces a payload to the caller under the interrupt key. You resume by invoking the same thread_id with Command(resume=value); the interrupt() call then returns that value.

open as a page

In LangGraph, what does returning Command(goto=...) from a node give you over a plain state update?

level: middleimportance: must knowfreq 68%
basics
~20 s

A Command return carries both a state update and a goto, so one node writes state and names its successor in a single value. That is the handoff idiom in multi-agent graphs; a plain dict return leaves routing entirely to the declared edges.

open as a page

How do you run a CrewAI Flow and pass initial values into its state?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Instantiate the Flow subclass and call kickoff(), or kickoff_async() to await it. Pass a dict as the inputs argument and those values are applied to the flow state before the start methods run. kickoff returns the last executed step's output.

open as a page

In CrewAI, what does allow_delegation=True add to an Agent, and what does it cost?

level: middleimportance: must knowfreq 64%
basics
~20 s

It appends coworker tools — delegate work and ask a question — to that agent's toolset, built from the crew's other agents and addressed by role string. The cost is extra LLM turns, longer prompts, and lossy handoffs, since only text crosses between agents.

open as a page

In CrewAI, what do an Agent's role, goal and backstory actually change at runtime?

level: middleimportance: must knowfreq 80%
basics
~20 s

Role, goal and backstory are prompt text: CrewAI interpolates them into the system prompt sent to the model for that agent. Role additionally names the agent for delegation and logs. None of the three enforces anything — tools and task assignment do that.

open as a page

What must you configure on a CrewAI Crew before Process.hierarchical will run?

level: middleimportance: must knowfreq 62%
basics
~20 s

A hierarchical CrewAI crew needs a manager: either manager_llm, from which CrewAI builds a default manager agent, or manager_agent, your own Agent. Supplying neither fails validation, and the manager agent must not also appear in the crew's agents list.

open as a page

In CrewAI, what changes when a Crew runs Process.hierarchical instead of Process.sequential?

level: middleimportance: must knowfreq 85%
basics
~20 s

Process.sequential executes the Crew's tasks in the order you listed them, each task seeing earlier output. Process.hierarchical inserts a manager agent that decides which worker gets which task and reviews the result, adding LLM calls, latency and nondeterminism.

open as a page

In AutoGen, what is the difference between LocalCommandLineCodeExecutor and DockerCommandLineCodeExecutor?

level: juniorimportance: must knowfreq 80%
basics
~20 s

LocalCommandLineCodeExecutor runs model-generated code as a subprocess on the host machine, with your files, environment variables and network. DockerCommandLineCodeExecutor runs the same code inside a container instead, isolating it from the host. Local is for trusted prototyping only.

open as a page

In AutoGen, how does RoundRobinGroupChat pick the next speaker and what does each agent see?

level: juniorimportance: must knowfreq 68%
basics
~20 s

RoundRobinGroupChat cycles participants in the exact order passed to its constructor, one response per turn, with no model deciding anything. Every response is broadcast to all participants, so each agent sees the whole shared conversation.

open as a page

How do you configure AutoGen's OpenAIChatCompletionClient for an unknown model?

level: middleimportance: must knowfreq 56%
basics
~20 s

Pass an explicit model_info dictionary alongside model and base_url. The client keeps a capability table for models it recognizes; for anything outside it — a self-hosted or newly released model behind an OpenAI-compatible endpoint — it cannot infer capabilities and raises unless you declare them.

open as a page

In AutoGen, what happens in one AssistantAgent turn when the model calls tools?

level: middleimportance: must knowfreq 70%
basics
~20 s

AssistantAgent makes one model call with its system message plus its context, executes every requested tool, then by default returns a ToolCallSummaryMessage holding the raw tool results. It only sends those results back to the model when reflect_on_tool_use is True.

open as a page

In AutoGen v0.4+, what do autogen-core, autogen-agentchat and autogen-ext each provide?

level: middleimportance: must knowfreq 72%
basics
~20 s

autogen-core is the event-driven actor runtime plus base abstractions. autogen-agentchat is the opinionated agent-and-team API built on top of it. autogen-ext holds the concrete integrations: model clients, tool adapters, code executors and the gRPC runtime.

open as a page

In Google ADK, what is the difference between an agent's instruction and description?

level: juniorimportance: must knowfreq 64%
basics
~20 s

instruction is the agent's own prompt — it tells that agent how to behave. description is a one-line summary other agents read when deciding whether to hand work over, so it drives delegation rather than behaviour.

open as a page

In Google ADK, what do adk web, adk run, and adk api_server each give you?

level: juniorimportance: must knowfreq 68%
basics
~20 s

adk web starts a local dev UI over your agent folder, with chat plus Events, Trace and Eval tabs. adk run drives the same agent from the terminal. adk api_server serves it as a local HTTP API.

open as a page

In Google ADK, what survives a restart with InMemorySessionService vs DatabaseSessionService?

level: juniorimportance: must knowfreq 72%
basics
~10 s

Nothing survives with InMemorySessionService — sessions, event history and state live in the process and vanish on restart. DatabaseSessionService writes them to a SQL database given by db_url, so the same session_id resumes afterwards.

open as a page

In Google ADK, how does a plain Python function become a callable tool?

level: juniorimportance: must knowfreq 80%
basics
~20 s

Google ADK wraps any callable placed in an agent's tools list into a FunctionTool. The function name, its type-hinted parameters and its docstring become the declaration the model sees, so the docstring is functional code rather than a comment.

open as a page

In Google ADK, when do you use a workflow agent instead of an LlmAgent?

level: middleimportance: must knowfreq 72%
basics
~20 s

Use LlmAgent when the model must decide what happens next. Use ADK's SequentialAgent, ParallelAgent or LoopAgent when the order is already known: they run their sub_agents on fixed control flow and call no LLM of their own.

open as a page

In Koog, when do you need a custom strategy graph instead of a plain AIAgent?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A plain Koog AIAgent already runs a built-in loop: send the prompt, execute any tool calls, return the final message. You pass a strategy when you need explicit control flow — branching, custom Kotlin steps, history compression, or subgraphs.

open as a page

How do you build a prompt with Koog's prompt("id") { } DSL, and what is the id for?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Koog's prompt("id") { } builder returns an immutable Prompt made of typed messages added by system(), user() and assistant() calls. The id is a stable label for that prompt, used to identify it in logs, traces and tooling — not sent as content.

open as a page

In Koog, what does ToolRegistry do and what breaks if a tool is missing?

level: juniorimportance: must knowfreq 70%
basics
~20 s

ToolRegistry is the single list of tools a Koog agent may call. Koog sends each registered tool's name, description and parameter schema with the prompt, then resolves incoming calls back to Kotlin. An unregistered tool is invisible and unresolvable.

open as a page

In a Koog strategy graph, which edges close the LLM tool-calling loop?

level: middleimportance: must knowfreq 66%
basics
~10 s

Send nodeLLMRequest to nodeExecuteTool on onToolCall, nodeExecuteTool straight to nodeLLMSendToolResult, and nodeLLMSendToolResult back to nodeExecuteTool on onToolCall. Both LLM nodes also need an onAssistantMessage edge to nodeFinish, or the loop never exits.

open as a page

In Koog, what does a PromptExecutor do, and how does MultiLLMPromptExecutor differ?

level: middleimportance: must knowfreq 72%
basics
~20 s

A PromptExecutor is Koog's single suspending entry point for running a Prompt against an LLModel, hiding provider HTTP details behind an LLMClient. SingleLLMPromptExecutor wraps one client; MultiLLMPromptExecutor holds several and routes each call by the model's provider.

open as a page

In LlamaIndex, what does QueryEngineTool.from_defaults do and why does its description matter?

level: juniorimportance: must knowfreq 66%
basics
~10 s

QueryEngineTool.from_defaults wraps an existing query engine in LlamaIndex's tool interface with a name and description. That description is the only thing the model reads when deciding whether to route a question to that corpus.

open as a page

In LlamaIndex's SentenceSplitter, what do chunk_size and chunk_overlap measure?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Both are token counts, not characters. chunk_size (default 1024) caps how many tokens a node may hold, measured with the tokenizer LlamaIndex is configured with; chunk_overlap (default 200) repeats that many trailing tokens at the start of the next node.

open as a page

In LlamaIndex, what does VectorStoreIndex.from_documents() actually do?

level: juniorimportance: must knowfreq 78%
basics
~20 s

VectorStoreIndex.from_documents() splits each document into nodes, calls the embedding model once per node, and writes those vectors plus the node text into a vector store — by default an in-memory SimpleVectorStore that vanishes when the process exits.

open as a page

What does SimpleDirectoryReader do in LlamaIndex, and how do you control which files it loads?

level: juniorimportance: must knowfreq 68%
basics
~10 s

SimpleDirectoryReader walks a folder, picks a reader based on each file's extension, and returns Document objects with file metadata attached. You narrow what it touches with input_files, required_exts, exclude, exclude_hidden, recursive and num_files_limit.

open as a page

In LlamaIndex, what is the difference between index.as_retriever() and index.as_query_engine()?

level: juniorimportance: must knowfreq 76%
basics
~20 s

In LlamaIndex, as_retriever() returns an object whose retrieve() call gives back scored nodes and makes no LLM call. as_query_engine() wraps a retriever with a response synthesizer, so query() also calls the LLM and returns an answer plus its source nodes.

open as a page

What does Haystack's Agent component add over calling a ChatGenerator yourself?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Haystack's Agent packages the whole tool-calling loop into one component: it sends messages to the chat generator, executes requested tools through an internal ToolInvoker, feeds the results back, and stops on its exit conditions or step limit.

open as a page

In Haystack, what does ChatPromptBuilder do inside a RAG pipeline?

level: juniorimportance: must knowfreq 62%
basics
~20 s

ChatPromptBuilder renders a Jinja2 template into a list of ChatMessage objects. Every placeholder in the template becomes an input it expects at run time, and the rendered messages leave on its prompt output, ready for a chat generator.

open as a page

Which components make up a Haystack indexing pipeline, and what does each do?

level: juniorimportance: must knowfreq 72%
basics
~10 s

A Haystack indexing pipeline runs a converter (file bytes to Document), then DocumentCleaner (strip whitespace and boilerplate), DocumentSplitter (cut into smaller Documents), a document embedder (attach vectors), and DocumentWriter (persist to a DocumentStore).

open as a page

How do you shape the input dict for Haystack's Pipeline.run(), and what does it return?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Pipeline.run() takes a dict keyed by component name, whose value is that component's run() keyword arguments. It returns a dict, also keyed by component name, holding the outputs no other component consumed. Use include_outputs_from to also get intermediate outputs.

open as a page

In Haystack, how does wiring InMemoryBM25Retriever differ from InMemoryEmbeddingRetriever?

level: juniorimportance: must knowfreq 74%
basics
~20 s

InMemoryBM25Retriever takes the raw query string on its query input and scores lexically over document text. InMemoryEmbeddingRetriever takes a query_embedding, so a text embedder must run first in the pipeline and feed its vector into that socket.

open as a page

What services does the RAGFlow Docker Compose stack run, and what does each store?

level: juniorimportance: must knowfreq 68%
basics
~20 s

RAGFlow ships as a Compose stack: the ragflow server container (web UI plus API, with a task executor beside it), a document engine (Elasticsearch by default, or Infinity) for chunks and vectors, MySQL for metadata, MinIO for the original files, and Redis for the task queue.

open as a page

In RAGFlow, what does choosing a chunk method per document actually change?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The chunk method is a per-document-type parsing template — General, Q&A, Manual, Table, Paper, Book, Laws, Presentation or One — that decides where the document is cut and what a chunk represents. It is a knowledge-base default you can override per file, and changing it requires re-parsing.

open as a page

In a RAGFlow chat assistant's system prompt, what does the {knowledge} variable do?

level: juniorimportance: must knowfreq 68%
basics
~20 s

{knowledge} is the placeholder RAGFlow fills with the retrieved chunks before calling the model. Remove it from the assistant's system prompt and retrieval still runs, but no document text ever reaches the LLM, so answers stop being grounded.

open as a page

In RAGFlow's HTTP API, how do you authenticate and take a document from upload to searchable chunks?

level: middleimportance: must knowfreq 72%
basics
~20 s

Generate an API key in the RAGFlow UI and send it as Authorization: Bearer <key> to /api/v1. Create a dataset, upload documents into it, then trigger parsing explicitly — parsing is asynchronous, and a document is only searchable once its run state reaches completion.

open as a page

In RAGFlow, what does DeepDoc parsing do that plain PDF text extraction cannot?

level: middleimportance: must knowfreq 70%
basics
~20 s

DeepDoc runs vision models over each page: OCR for scanned text, layout analysis that labels titles, paragraphs, figures, captions and headers, and table-structure recognition that rebuilds tables as HTML. Plain extraction returns only a flat text stream.

open as a page