skip to content

RAG Overview

Retrieval-augmented generation retrieves relevant documents and injects them into the prompt, with an ETL pipeline preparing the store. Interviewers ask when RAG beats fine-tuning, and 'when the knowledge changes' is the short answer.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

What is RAG in Spring AI, and what does the QuestionAnswerAdvisor do?

level: juniorimportance: must knowfreq 70%

answer

  1. Retrieve chunks -> inject into prompt -> ground the answer
  2. VectorStore similarity search
  3. QuestionAnswerAdvisor on ChatClient
  4. {question_answer_context} placeholder
  5. No retraining, reduces hallucination

basics

~10 s

RAG (Retrieval-Augmented Generation) fetches relevant documents from a vector store and adds them to the prompt so the LLM answers from your data. Spring AI's QuestionAnswerAdvisor does this automatically on each ChatClient call.

solid answer

~40 s

RAG means Retrieval-Augmented Generation: instead of relying only on the model's training data, you retrieve relevant chunks from a knowledge base (a VectorStore) and inject them into the prompt as context, so the LLM grounds its answer in your data. In Spring AI you wire this with QuestionAnswerAdvisor, an Advisor you add to a ChatClient. On every request it runs a similarity search against the configured VectorStore using the user's question, takes the top matching Documents, and appends their text to the prompt via a template placeholder. The model then answers using that context. This reduces hallucination and lets you answer questions about private or recent data the model never saw, without retraining or fine-tuning.

code

java · 19 lines
java
@RestController
class DocsController {
    private final ChatClient chatClient;

    DocsController(ChatClient.Builder builder, VectorStore vectorStore) {
        this.chatClient = builder
            // Adds RAG to every call: similarity search + context injection
            .defaultAdvisors(new QuestionAnswerAdvisor(vectorStore))
            .build();
    }

    @GetMapping("/ask")
    String ask(@RequestParam String q) {
        return chatClient.prompt()
            .user(q)
            .call()
            .content();
    }
}

go deeper

for a junior

Know the definition: retrieve relevant docs, add to prompt, model answers from them; QuestionAnswerAdvisor wires it to a ChatClient.

for a middle

Should explain the VectorStore + EmbeddingModel + similarity search chain and that ingestion must happen first.

for a senior

Discuss context-window limits, retrieval quality driving answer quality, and instructing the model to stay within context.

for a principal

Frame RAG vs fine-tuning trade-offs, freshness, cost per token, and evaluation of retrieval quality as a system concern.

**RAG (Retrieval-Augmented Generation)** is a pattern for making an LLM answer from data it was never trained on. A large language model only "knows" what was in its training set, which is frozen and public. RAG adds a retrieval step: at query time you look up relevant text in your own knowledge base and paste it into the prompt as *context*, so the model reads your data and answers from it. This is cheaper and faster than fine-tuning and keeps answers current. **The building blocks in Spring AI:** - `Document` — the core content unit: text plus a `Map<String,Object>` of metadata (source, page, etc.). - `VectorStore` — stores Documents as *embeddings* (numeric vectors capturing meaning) and supports *similarity search*: given a query, it returns the semantically closest Documents. Implementations include `PgVectorStore`, `RedisVectorStore`, `ChromaVectorStore`, and the in-memory `SimpleVectorStore`. - `EmbeddingModel` — turns text into vectors; the VectorStore uses it under the hood. - `ChatClient` — the fluent API to call the LLM. - `Advisor` — an interceptor in the ChatClient call chain that can modify the request/response. **`QuestionAnswerAdvisor`** is the simplest RAG advisor. You add it via `ChatClient.builder(...).defaultAdvisors(new QuestionAnswerAdvisor(vectorStore))` or per-call with `.advisors(...)`. On each request it: (1) takes the user's question, (2) runs a similarity search on the `VectorStore`, (3) formats the retrieved Documents into the prompt using a template that has a `{question_answer_context}` placeholder, and (4) sends the augmented prompt to the model. **Why it matters / when to use:** Use RAG for question-answering over private docs, product manuals, tickets, or any corpus. Its limits: retrieval quality caps answer quality (garbage retrieved = garbage answer), the context window bounds how much you can inject, and you must first *ingest* your data into the VectorStore (the ETL step). **Gotchas:** the model can still hallucinate if the retrieved context is irrelevant; you should instruct it to answer only from context. Retrieval uses semantic similarity, not keyword match, so embedding quality matters. And you pay token cost for every injected chunk.

  • How does the retrieved data actually get into the prompt?
    QuestionAnswerAdvisor formats the retrieved Documents' text and substitutes it into a prompt template at the {question_answer_context} placeholder, appending it to the user message before the call reaches the model.
  • What has to happen before RAG can retrieve anything?
    You must ingest your data first: read documents, split/transform them, embed them, and write them into the VectorStore (the ETL pipeline). Retrieval can only find what was previously stored.

saying these in an interview costs you the question

  • Thinking RAG fine-tunes or retrains the model (it does not — it only changes the prompt at runtime)
  • Believing the model searches the internet or a database itself (retrieval is a separate step done by the advisor/VectorStore)
  • Assuming QuestionAnswerAdvisor also ingests data (it only retrieves; ingestion is the ETL pipeline)

context

open as a page

Describe Spring AI's ETL pipeline for ingesting documents into a vector store (Reader, Transformer, Writer).

level: middleimportance: must knowfreq 65%

basics

~10 s

You read source files into Documents with a DocumentReader, split/enrich them with a DocumentTransformer (e.g., TokenTextSplitter), then write them into the VectorStore with a DocumentWriter, which embeds and stores them for later retrieval.

open as a page

How does QuestionAnswerAdvisor perform context injection, and how do you control what it retrieves?

level: seniorimportance: should knowfreq 55%

basics

~20 s

It runs a similarity search using a SearchRequest (topK, similarityThreshold, filter), formats the returned Documents' text, and substitutes them into a prompt template at the {question_answer_context} placeholder before the model call. You tune retrieval via the SearchRequest and override the template.

open as a page

How does RetrievalAugmentationAdvisor differ from QuestionAnswerAdvisor, and what modular stages does it support?

level: seniorimportance: should knowfreq 40%

basics

~10 s

RetrievalAugmentationAdvisor is Spring AI's modular RAG advisor. Unlike the simple QuestionAnswerAdvisor, it lets you compose pre-retrieval query transformation/expansion, a pluggable DocumentRetriever, post-retrieval processing, and a QueryAugmenter that injects context and can handle empty results.

open as a page

As an architect, how do you design and evaluate a production RAG pipeline in Spring AI end to end?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Design two paths: an offline ETL (read, chunk, enrich, embed, write to VectorStore) and an online query path (retrieve with SearchRequest, augment, generate). Tune chunking, topK, threshold, and metadata filters; measure retrieval quality and guard against empty-context hallucination.

open as a page