How does QuestionAnswerAdvisor perform context injection, and how do you control what it retrieves?
answer
- SearchRequest: topK, similarityThreshold, filterExpression
- {question_answer_context} placeholder in PromptTemplate
- similaritySearch before the model call
- Override template but keep the placeholder
- Empty context -> possible hallucination (QAA doesn't guard)
basics
~20 sIt runs a similarity search using a SearchRequest (topK, similarityThreshold, filter), formats the returned Documents' text, and substitutes them into a prompt template at the {question_answer_context} placeholder before the model call. You tune retrieval via the SearchRequest and override the template.
solid answer
~40 sQuestionAnswerAdvisor is a call/stream Advisor in the ChatClient chain. Before the request reaches the model, its `before` logic builds a `SearchRequest` from the user's query and calls `vectorStore.similaritySearch(...)`. You configure retrieval by passing a `SearchRequest` to the advisor's builder: `topK` (how many chunks), `similarityThreshold` (min relevance), and a `FilterExpression` for metadata filtering (e.g., only a given tenant or category). The retrieved Documents' content is joined and injected into a `PromptTemplate` that carries a `{question_answer_context}` placeholder (and `{query}`); the default template instructs the model to answer from the context. You can override that template with `.promptTemplate(...)`. You can also pass per-request filter expressions via advisor context parameters. Order matters if you stack advisors (memory, logging), controlled by `getOrder()`.
code
java · 20 linesvar advisor = QuestionAnswerAdvisor.builder(vectorStore)
.searchRequest(SearchRequest.builder()
.topK(6)
.similarityThreshold(0.75)
.build())
.promptTemplate(new PromptTemplate("""
Answer ONLY from the context. If it is not there, say you don't know.
Context:
{question_answer_context}
Question: {query}
"""))
.build();
String answer = chatClient.prompt()
.user("What is our refund window?")
// per-request metadata filter via advisor context
.advisors(a -> a.param(QuestionAnswerAdvisor.FILTER_EXPRESSION, "category == 'policy'"))
.advisors(advisor)
.call()
.content();go deeper
Know it retrieves and injects context via a placeholder; details optional.
Explain SearchRequest knobs (topK/threshold/filter) and the {question_answer_context} template.
Cover template override rules, per-request filter params, empty-context hallucination risk, and advisor ordering.
Reason about token-budget vs recall trade-offs, filter-DSL portability across stores, and when to graduate to modular RAG.
**Where it runs:** `QuestionAnswerAdvisor` implements Spring AI's advisor contract (`BaseAdvisor` / `CallAdvisor` + `StreamAdvisor`). Advisors form an interceptor chain around the model call; QAA does its work *before* the request is sent and simply reads the response after. **The retrieval step — `SearchRequest`:** QAA turns the user's question into a query and issues `vectorStore.similaritySearch(SearchRequest)`. You configure the request via the advisor builder: ```java QuestionAnswerAdvisor.builder(vectorStore) .searchRequest(SearchRequest.builder() .topK(6) // number of chunks to return .similarityThreshold(0.75) // drop weak matches (0..1) .filterExpression("category == 'faq'") // metadata filter .build()) .build(); ``` - **`topK`** bounds how many Documents are injected — bigger = more recall but more tokens and noise. - **`similarityThreshold`** filters out low-relevance hits; too high can return nothing. - **`filterExpression`** applies structured metadata filtering (portable filter DSL translated per store), essential for multi-tenant or category-scoped retrieval. It can be set statically or supplied per-request through the advisor's context (`FILTER_EXPRESSION` parameter passed via `.advisors(a -> a.param(...))`). **The injection step — the prompt template:** QAA formats the retrieved Documents into a `PromptTemplate`. The default template contains a `{question_answer_context}` placeholder where the joined document text goes, plus the user `{query}`, and text like "answer the query using the context; if unknown, say you don't know." You override it with `.promptTemplate(new PromptTemplate(customText))` — the custom template **must** keep the `{question_answer_context}` placeholder or injection breaks. **Response side:** by default QAA does not attach source citations; if you need the retrieved Documents surfaced, you capture them yourself (e.g., a custom advisor or reading `ChatResponse` metadata where supported). **Gotchas:** - **Empty retrieval** still calls the model with an empty context block — the model may then hallucinate unless the template instructs otherwise. (The newer `RetrievalAugmentationAdvisor` handles empty context explicitly.) - **Token budget:** `topK` × chunk size can exceed the context window; balance with chunking and threshold. - **Filter DSL portability:** expression is translated to each store's native filter; unsupported operators fail per store. - **Advisor ordering:** if combined with `MessageChatMemoryAdvisor`, order affects whether retrieval uses the rewritten/expanded query — QAA uses the raw user text, it does not rewrite the query (that's RetrievalAugmentationAdvisor's job). **When to use:** QAA is ideal for straightforward single-query RAG. Reach for `RetrievalAugmentationAdvisor` when you need query transformation, expansion, or explicit empty-context handling.
- What happens if the similarity search returns no documents?QuestionAnswerAdvisor still calls the model with an empty {question_answer_context}, so the model may answer from its own knowledge or hallucinate. Guard via the prompt template's instructions, or use RetrievalAugmentationAdvisor's ContextualQueryAugmenter which handles empty context explicitly.
- How would you scope retrieval to a single tenant?Store a tenant id in each Document's metadata at ingest, then set a filterExpression (statically or per-request via the FILTER_EXPRESSION advisor param) so the SearchRequest only matches that tenant's chunks.
saying these in an interview costs you the question
- Saying QuestionAnswerAdvisor rewrites or expands the query (it uses the raw user text; rewriting is RetrievalAugmentationAdvisor)
- Thinking you can freely replace the template without keeping the {question_answer_context} placeholder
- Assuming it automatically refuses to answer when nothing is retrieved
- Believing filterExpression is a raw SQL WHERE clause rather than Spring AI's portable filter DSL