skip to content

What is RAG in Spring AI, and what does the QuestionAnswerAdvisor do?

level: juniorimportance: must knowfreq 70%

answer

  1. Retrieve chunks -> inject into prompt -> ground the answer
  2. VectorStore similarity search
  3. QuestionAnswerAdvisor on ChatClient
  4. {question_answer_context} placeholder
  5. No retraining, reduces hallucination

basics

~10 s

RAG (Retrieval-Augmented Generation) fetches relevant documents from a vector store and adds them to the prompt so the LLM answers from your data. Spring AI's QuestionAnswerAdvisor does this automatically on each ChatClient call.

solid answer

~40 s

RAG means Retrieval-Augmented Generation: instead of relying only on the model's training data, you retrieve relevant chunks from a knowledge base (a VectorStore) and inject them into the prompt as context, so the LLM grounds its answer in your data. In Spring AI you wire this with QuestionAnswerAdvisor, an Advisor you add to a ChatClient. On every request it runs a similarity search against the configured VectorStore using the user's question, takes the top matching Documents, and appends their text to the prompt via a template placeholder. The model then answers using that context. This reduces hallucination and lets you answer questions about private or recent data the model never saw, without retraining or fine-tuning.

code

java · 19 lines
java
@RestController
class DocsController {
    private final ChatClient chatClient;

    DocsController(ChatClient.Builder builder, VectorStore vectorStore) {
        this.chatClient = builder
            // Adds RAG to every call: similarity search + context injection
            .defaultAdvisors(new QuestionAnswerAdvisor(vectorStore))
            .build();
    }

    @GetMapping("/ask")
    String ask(@RequestParam String q) {
        return chatClient.prompt()
            .user(q)
            .call()
            .content();
    }
}

go deeper

for a junior

Know the definition: retrieve relevant docs, add to prompt, model answers from them; QuestionAnswerAdvisor wires it to a ChatClient.

for a middle

Should explain the VectorStore + EmbeddingModel + similarity search chain and that ingestion must happen first.

for a senior

Discuss context-window limits, retrieval quality driving answer quality, and instructing the model to stay within context.

for a principal

Frame RAG vs fine-tuning trade-offs, freshness, cost per token, and evaluation of retrieval quality as a system concern.

**RAG (Retrieval-Augmented Generation)** is a pattern for making an LLM answer from data it was never trained on. A large language model only "knows" what was in its training set, which is frozen and public. RAG adds a retrieval step: at query time you look up relevant text in your own knowledge base and paste it into the prompt as *context*, so the model reads your data and answers from it. This is cheaper and faster than fine-tuning and keeps answers current. **The building blocks in Spring AI:** - `Document` — the core content unit: text plus a `Map<String,Object>` of metadata (source, page, etc.). - `VectorStore` — stores Documents as *embeddings* (numeric vectors capturing meaning) and supports *similarity search*: given a query, it returns the semantically closest Documents. Implementations include `PgVectorStore`, `RedisVectorStore`, `ChromaVectorStore`, and the in-memory `SimpleVectorStore`. - `EmbeddingModel` — turns text into vectors; the VectorStore uses it under the hood. - `ChatClient` — the fluent API to call the LLM. - `Advisor` — an interceptor in the ChatClient call chain that can modify the request/response. **`QuestionAnswerAdvisor`** is the simplest RAG advisor. You add it via `ChatClient.builder(...).defaultAdvisors(new QuestionAnswerAdvisor(vectorStore))` or per-call with `.advisors(...)`. On each request it: (1) takes the user's question, (2) runs a similarity search on the `VectorStore`, (3) formats the retrieved Documents into the prompt using a template that has a `{question_answer_context}` placeholder, and (4) sends the augmented prompt to the model. **Why it matters / when to use:** Use RAG for question-answering over private docs, product manuals, tickets, or any corpus. Its limits: retrieval quality caps answer quality (garbage retrieved = garbage answer), the context window bounds how much you can inject, and you must first *ingest* your data into the VectorStore (the ETL step). **Gotchas:** the model can still hallucinate if the retrieved context is irrelevant; you should instruct it to answer only from context. Retrieval uses semantic similarity, not keyword match, so embedding quality matters. And you pay token cost for every injected chunk.

  • How does the retrieved data actually get into the prompt?
    QuestionAnswerAdvisor formats the retrieved Documents' text and substitutes it into a prompt template at the {question_answer_context} placeholder, appending it to the user message before the call reaches the model.
  • What has to happen before RAG can retrieve anything?
    You must ingest your data first: read documents, split/transform them, embed them, and write them into the VectorStore (the ETL pipeline). Retrieval can only find what was previously stored.

saying these in an interview costs you the question

  • Thinking RAG fine-tunes or retrains the model (it does not — it only changes the prompt at runtime)
  • Believing the model searches the internet or a database itself (retrieval is a separate step done by the advisor/VectorStore)
  • Assuming QuestionAnswerAdvisor also ingests data (it only retrieves; ingestion is the ETL pipeline)

context