skip to content

How does RetrievalAugmentationAdvisor differ from QuestionAnswerAdvisor, and what modular stages does it support?

level: seniorimportance: should knowfreq 40%

answer

  1. Modular RAG: pre-retrieval / retrieval / post-retrieval / generation
  2. QueryTransformer (rewrite/translate/compress) + QueryExpander (multi-query)
  3. VectorStoreDocumentRetriever + DocumentJoiner
  4. ContextualQueryAugmenter.allowEmptyContext guards hallucination
  5. QAA = simple one-shot; RAA = composable

basics

~10 s

RetrievalAugmentationAdvisor is Spring AI's modular RAG advisor. Unlike the simple QuestionAnswerAdvisor, it lets you compose pre-retrieval query transformation/expansion, a pluggable DocumentRetriever, post-retrieval processing, and a QueryAugmenter that injects context and can handle empty results.

solid answer

~40 s

`QuestionAnswerAdvisor` is one-shot: search then inject. `RetrievalAugmentationAdvisor` implements *modular RAG* — a configurable pipeline of stages you assemble from components in `org.springframework.ai.rag`. **Pre-retrieval:** `QueryTransformer`s (`RewriteQueryTransformer`, `TranslationQueryTransformer`, `CompressionQueryTransformer`) clean/rewrite the query, and `QueryExpander`s (`MultiQueryExpander`) generate query variants. **Retrieval:** a `DocumentRetriever` (`VectorStoreDocumentRetriever`) fetches Documents, with a `DocumentJoiner` merging results from multiple sub-queries. **Post-retrieval:** `DocumentPostProcessor`s can rank/compress/select. **Generation:** a `QueryAugmenter` (`ContextualQueryAugmenter`) injects the documents into the prompt and, crucially, can decide behavior when context is empty (return a canned answer vs. allow the model to answer). You wire it with `RetrievalAugmentationAdvisor.builder()`. Use it when a single similarity search isn't enough — multilingual queries, conversational rewriting, multi-query recall, or explicit no-context guarding.

code

java · 21 lines
java
Advisor rag = RetrievalAugmentationAdvisor.builder()
    // Pre-retrieval: rewrite conversational query into a search query
    .queryTransformers(RewriteQueryTransformer.builder()
        .chatClientBuilder(chatClientBuilder).build())
    // Retrieval: pull from the vector store
    .documentRetriever(VectorStoreDocumentRetriever.builder()
        .vectorStore(vectorStore)
        .topK(6)
        .similarityThreshold(0.7)
        .build())
    // Generation: inject context; refuse safely when nothing is found
    .queryAugmenter(ContextualQueryAugmenter.builder()
        .allowEmptyContext(false)
        .build())
    .build();

String answer = chatClient.prompt()
    .advisors(rag)
    .user("Remind me what the refund window was?")
    .call()
    .content();

go deeper

for a junior

Just know a more advanced, configurable RAG advisor exists beyond QuestionAnswerAdvisor.

for a middle

Name the four stages and the main components (transformer, retriever, augmenter).

for a senior

Explain when to choose RAA over QAA, the allowEmptyContext guard, and multi-query recall trade-offs.

for a principal

Weigh added LLM-call latency/cost per stage, dedup/joining strategy, reranking, and evolving-API stability when standardizing a RAG platform.

`RetrievalAugmentationAdvisor` (package `org.springframework.ai.rag.advisor`) is Spring AI's **Advanced/Modular RAG** advisor, modeled on the Modular RAG architecture. Where `QuestionAnswerAdvisor` bakes in a single 'search then inject' flow, this advisor exposes each RAG stage as a swappable component, so you can build sophisticated pipelines declaratively. **The four stages and their component types (all under `org.springframework.ai.rag`):** 1. **Pre-Retrieval — reshape the query before searching.** - `QueryTransformer` (`org.springframework.ai.rag.preretrieval.query.transformation`): - `RewriteQueryTransformer` — uses the LLM to rewrite a verbose/conversational question into a search-optimized query. - `TranslationQueryTransformer` — translates the query into the language of your indexed content (multilingual corpora). - `CompressionQueryTransformer` — compresses conversation history + follow-up into a standalone query (good for chat). - `QueryExpander` (`...preretrieval.query.expansion`): - `MultiQueryExpander` — produces several query variants to widen recall; each is retrieved independently. 2. **Retrieval — fetch candidate Documents.** - `DocumentRetriever` (`...retrieval.search`): `VectorStoreDocumentRetriever` wraps a `VectorStore` with its own topK, similarityThreshold, and (optionally per-request) filter expression. - `DocumentJoiner` (`...retrieval.join`): `ConcatenationDocumentJoiner` merges/dedups results when multiple expanded queries each return documents. 3. **Post-Retrieval — refine the candidate set.** - `DocumentPostProcessor` (`...postretrieval.document`): rank, compress, or select documents (e.g., drop redundant chunks, re-order by relevance) before augmentation. 4. **Generation/Augmentation — inject context into the prompt.** - `QueryAugmenter` (`...generation.augmentation`): `ContextualQueryAugmenter` formats the retrieved Documents into the prompt. Its key options: - `allowEmptyContext` — when **false** (default) and nothing was retrieved, it augments with an *empty-context* prompt so the model returns a safe 'I don't have information' style answer instead of hallucinating; when **true**, it lets the model answer from its own knowledge. - Customizable prompt templates for both the normal and empty-context cases. **Wiring:** ```java var advisor = RetrievalAugmentationAdvisor.builder() .queryTransformers(RewriteQueryTransformer.builder().chatClientBuilder(builder).build()) .documentRetriever(VectorStoreDocumentRetriever.builder() .vectorStore(vectorStore).topK(6).similarityThreshold(0.7).build()) .queryAugmenter(ContextualQueryAugmenter.builder().allowEmptyContext(false).build()) .build(); chatClient.prompt().advisors(advisor).user(q).call().content(); ``` **QAA vs RAA — choosing:** - Use **QuestionAnswerAdvisor** for the common case: single query, single store, simple injection. Less config, fewer LLM calls. - Use **RetrievalAugmentationAdvisor** when you need any of: query rewriting/compression for chat, translation for multilingual data, multi-query expansion for recall, post-retrieval reranking/compression, or explicit empty-context handling. **Gotchas:** - Query transformers and expanders **cost extra LLM calls** (latency + tokens) per request — measure before enabling. - `MultiQueryExpander` multiplies retrieval calls; the `DocumentJoiner` must dedup or you inject redundant context. - Some components in the RAG module were incubating in early 1.0 releases; verify exact builder names against your Spring AI version. - Empty-context behavior is opt-in via `ContextualQueryAugmenter` — the default guards against hallucination, which is a behavior change to be aware of if you expected the model to always answer.

  • What does ContextualQueryAugmenter's allowEmptyContext flag control?
    When false (default), if retrieval returns no documents it augments with an empty-context prompt so the model responds with a safe 'no information' answer instead of hallucinating; when true, it permits the model to answer from its own parametric knowledge.
  • Why would you add a MultiQueryExpander, and what's its cost?
    It generates several query variants to improve recall for ambiguous questions. Cost: an extra LLM call to produce variants plus one retrieval per variant, so latency, tokens, and the need for a DocumentJoiner to dedup results all increase.

saying these in an interview costs you the question

  • Treating RetrievalAugmentationAdvisor as just a renamed QuestionAnswerAdvisor with no added stages
  • Thinking query rewriting/expansion is free (each is an extra LLM call)
  • Assuming empty retrieval always yields a refusal — it depends on allowEmptyContext
  • Forgetting a DocumentJoiner is needed when multiple expanded queries each return documents

context