How does RetrievalAugmentationAdvisor differ from QuestionAnswerAdvisor, and what modular stages does it support?
answer
- Modular RAG: pre-retrieval / retrieval / post-retrieval / generation
- QueryTransformer (rewrite/translate/compress) + QueryExpander (multi-query)
- VectorStoreDocumentRetriever + DocumentJoiner
- ContextualQueryAugmenter.allowEmptyContext guards hallucination
- QAA = simple one-shot; RAA = composable
basics
~10 sRetrievalAugmentationAdvisor is Spring AI's modular RAG advisor. Unlike the simple QuestionAnswerAdvisor, it lets you compose pre-retrieval query transformation/expansion, a pluggable DocumentRetriever, post-retrieval processing, and a QueryAugmenter that injects context and can handle empty results.
solid answer
~40 s`QuestionAnswerAdvisor` is one-shot: search then inject. `RetrievalAugmentationAdvisor` implements *modular RAG* — a configurable pipeline of stages you assemble from components in `org.springframework.ai.rag`. **Pre-retrieval:** `QueryTransformer`s (`RewriteQueryTransformer`, `TranslationQueryTransformer`, `CompressionQueryTransformer`) clean/rewrite the query, and `QueryExpander`s (`MultiQueryExpander`) generate query variants. **Retrieval:** a `DocumentRetriever` (`VectorStoreDocumentRetriever`) fetches Documents, with a `DocumentJoiner` merging results from multiple sub-queries. **Post-retrieval:** `DocumentPostProcessor`s can rank/compress/select. **Generation:** a `QueryAugmenter` (`ContextualQueryAugmenter`) injects the documents into the prompt and, crucially, can decide behavior when context is empty (return a canned answer vs. allow the model to answer). You wire it with `RetrievalAugmentationAdvisor.builder()`. Use it when a single similarity search isn't enough — multilingual queries, conversational rewriting, multi-query recall, or explicit no-context guarding.
code
java · 21 linesAdvisor rag = RetrievalAugmentationAdvisor.builder()
// Pre-retrieval: rewrite conversational query into a search query
.queryTransformers(RewriteQueryTransformer.builder()
.chatClientBuilder(chatClientBuilder).build())
// Retrieval: pull from the vector store
.documentRetriever(VectorStoreDocumentRetriever.builder()
.vectorStore(vectorStore)
.topK(6)
.similarityThreshold(0.7)
.build())
// Generation: inject context; refuse safely when nothing is found
.queryAugmenter(ContextualQueryAugmenter.builder()
.allowEmptyContext(false)
.build())
.build();
String answer = chatClient.prompt()
.advisors(rag)
.user("Remind me what the refund window was?")
.call()
.content();go deeper
Just know a more advanced, configurable RAG advisor exists beyond QuestionAnswerAdvisor.
Name the four stages and the main components (transformer, retriever, augmenter).
Explain when to choose RAA over QAA, the allowEmptyContext guard, and multi-query recall trade-offs.
Weigh added LLM-call latency/cost per stage, dedup/joining strategy, reranking, and evolving-API stability when standardizing a RAG platform.
`RetrievalAugmentationAdvisor` (package `org.springframework.ai.rag.advisor`) is Spring AI's **Advanced/Modular RAG** advisor, modeled on the Modular RAG architecture. Where `QuestionAnswerAdvisor` bakes in a single 'search then inject' flow, this advisor exposes each RAG stage as a swappable component, so you can build sophisticated pipelines declaratively. **The four stages and their component types (all under `org.springframework.ai.rag`):** 1. **Pre-Retrieval — reshape the query before searching.** - `QueryTransformer` (`org.springframework.ai.rag.preretrieval.query.transformation`): - `RewriteQueryTransformer` — uses the LLM to rewrite a verbose/conversational question into a search-optimized query. - `TranslationQueryTransformer` — translates the query into the language of your indexed content (multilingual corpora). - `CompressionQueryTransformer` — compresses conversation history + follow-up into a standalone query (good for chat). - `QueryExpander` (`...preretrieval.query.expansion`): - `MultiQueryExpander` — produces several query variants to widen recall; each is retrieved independently. 2. **Retrieval — fetch candidate Documents.** - `DocumentRetriever` (`...retrieval.search`): `VectorStoreDocumentRetriever` wraps a `VectorStore` with its own topK, similarityThreshold, and (optionally per-request) filter expression. - `DocumentJoiner` (`...retrieval.join`): `ConcatenationDocumentJoiner` merges/dedups results when multiple expanded queries each return documents. 3. **Post-Retrieval — refine the candidate set.** - `DocumentPostProcessor` (`...postretrieval.document`): rank, compress, or select documents (e.g., drop redundant chunks, re-order by relevance) before augmentation. 4. **Generation/Augmentation — inject context into the prompt.** - `QueryAugmenter` (`...generation.augmentation`): `ContextualQueryAugmenter` formats the retrieved Documents into the prompt. Its key options: - `allowEmptyContext` — when **false** (default) and nothing was retrieved, it augments with an *empty-context* prompt so the model returns a safe 'I don't have information' style answer instead of hallucinating; when **true**, it lets the model answer from its own knowledge. - Customizable prompt templates for both the normal and empty-context cases. **Wiring:** ```java var advisor = RetrievalAugmentationAdvisor.builder() .queryTransformers(RewriteQueryTransformer.builder().chatClientBuilder(builder).build()) .documentRetriever(VectorStoreDocumentRetriever.builder() .vectorStore(vectorStore).topK(6).similarityThreshold(0.7).build()) .queryAugmenter(ContextualQueryAugmenter.builder().allowEmptyContext(false).build()) .build(); chatClient.prompt().advisors(advisor).user(q).call().content(); ``` **QAA vs RAA — choosing:** - Use **QuestionAnswerAdvisor** for the common case: single query, single store, simple injection. Less config, fewer LLM calls. - Use **RetrievalAugmentationAdvisor** when you need any of: query rewriting/compression for chat, translation for multilingual data, multi-query expansion for recall, post-retrieval reranking/compression, or explicit empty-context handling. **Gotchas:** - Query transformers and expanders **cost extra LLM calls** (latency + tokens) per request — measure before enabling. - `MultiQueryExpander` multiplies retrieval calls; the `DocumentJoiner` must dedup or you inject redundant context. - Some components in the RAG module were incubating in early 1.0 releases; verify exact builder names against your Spring AI version. - Empty-context behavior is opt-in via `ContextualQueryAugmenter` — the default guards against hallucination, which is a behavior change to be aware of if you expected the model to always answer.
- What does ContextualQueryAugmenter's allowEmptyContext flag control?When false (default), if retrieval returns no documents it augments with an empty-context prompt so the model responds with a safe 'no information' answer instead of hallucinating; when true, it permits the model to answer from its own parametric knowledge.
- Why would you add a MultiQueryExpander, and what's its cost?It generates several query variants to improve recall for ambiguous questions. Cost: an extra LLM call to produce variants plus one retrieval per variant, so latency, tokens, and the need for a DocumentJoiner to dedup results all increase.
saying these in an interview costs you the question
- Treating RetrievalAugmentationAdvisor as just a renamed QuestionAnswerAdvisor with no added stages
- Thinking query rewriting/expansion is free (each is an extra LLM call)
- Assuming empty retrieval always yields a refusal — it depends on allowEmptyContext
- Forgetting a DocumentJoiner is needed when multiple expanded queries each return documents