skip to content

How does QuestionAnswerAdvisor perform context injection, and how do you control what it retrieves?

level: seniorimportance: should knowfreq 55%

answer

  1. SearchRequest: topK, similarityThreshold, filterExpression
  2. {question_answer_context} placeholder in PromptTemplate
  3. similaritySearch before the model call
  4. Override template but keep the placeholder
  5. Empty context -> possible hallucination (QAA doesn't guard)

basics

~20 s

It runs a similarity search using a SearchRequest (topK, similarityThreshold, filter), formats the returned Documents' text, and substitutes them into a prompt template at the {question_answer_context} placeholder before the model call. You tune retrieval via the SearchRequest and override the template.

solid answer

~40 s

QuestionAnswerAdvisor is a call/stream Advisor in the ChatClient chain. Before the request reaches the model, its `before` logic builds a `SearchRequest` from the user's query and calls `vectorStore.similaritySearch(...)`. You configure retrieval by passing a `SearchRequest` to the advisor's builder: `topK` (how many chunks), `similarityThreshold` (min relevance), and a `FilterExpression` for metadata filtering (e.g., only a given tenant or category). The retrieved Documents' content is joined and injected into a `PromptTemplate` that carries a `{question_answer_context}` placeholder (and `{query}`); the default template instructs the model to answer from the context. You can override that template with `.promptTemplate(...)`. You can also pass per-request filter expressions via advisor context parameters. Order matters if you stack advisors (memory, logging), controlled by `getOrder()`.

code

java · 20 lines
java
var advisor = QuestionAnswerAdvisor.builder(vectorStore)
    .searchRequest(SearchRequest.builder()
        .topK(6)
        .similarityThreshold(0.75)
        .build())
    .promptTemplate(new PromptTemplate("""
        Answer ONLY from the context. If it is not there, say you don't know.
        Context:
        {question_answer_context}
        Question: {query}
        """))
    .build();

String answer = chatClient.prompt()
    .user("What is our refund window?")
    // per-request metadata filter via advisor context
    .advisors(a -> a.param(QuestionAnswerAdvisor.FILTER_EXPRESSION, "category == 'policy'"))
    .advisors(advisor)
    .call()
    .content();

go deeper

for a junior

Know it retrieves and injects context via a placeholder; details optional.

for a middle

Explain SearchRequest knobs (topK/threshold/filter) and the {question_answer_context} template.

for a senior

Cover template override rules, per-request filter params, empty-context hallucination risk, and advisor ordering.

for a principal

Reason about token-budget vs recall trade-offs, filter-DSL portability across stores, and when to graduate to modular RAG.

**Where it runs:** `QuestionAnswerAdvisor` implements Spring AI's advisor contract (`BaseAdvisor` / `CallAdvisor` + `StreamAdvisor`). Advisors form an interceptor chain around the model call; QAA does its work *before* the request is sent and simply reads the response after. **The retrieval step — `SearchRequest`:** QAA turns the user's question into a query and issues `vectorStore.similaritySearch(SearchRequest)`. You configure the request via the advisor builder: ```java QuestionAnswerAdvisor.builder(vectorStore) .searchRequest(SearchRequest.builder() .topK(6) // number of chunks to return .similarityThreshold(0.75) // drop weak matches (0..1) .filterExpression("category == 'faq'") // metadata filter .build()) .build(); ``` - **`topK`** bounds how many Documents are injected — bigger = more recall but more tokens and noise. - **`similarityThreshold`** filters out low-relevance hits; too high can return nothing. - **`filterExpression`** applies structured metadata filtering (portable filter DSL translated per store), essential for multi-tenant or category-scoped retrieval. It can be set statically or supplied per-request through the advisor's context (`FILTER_EXPRESSION` parameter passed via `.advisors(a -> a.param(...))`). **The injection step — the prompt template:** QAA formats the retrieved Documents into a `PromptTemplate`. The default template contains a `{question_answer_context}` placeholder where the joined document text goes, plus the user `{query}`, and text like "answer the query using the context; if unknown, say you don't know." You override it with `.promptTemplate(new PromptTemplate(customText))` — the custom template **must** keep the `{question_answer_context}` placeholder or injection breaks. **Response side:** by default QAA does not attach source citations; if you need the retrieved Documents surfaced, you capture them yourself (e.g., a custom advisor or reading `ChatResponse` metadata where supported). **Gotchas:** - **Empty retrieval** still calls the model with an empty context block — the model may then hallucinate unless the template instructs otherwise. (The newer `RetrievalAugmentationAdvisor` handles empty context explicitly.) - **Token budget:** `topK` × chunk size can exceed the context window; balance with chunking and threshold. - **Filter DSL portability:** expression is translated to each store's native filter; unsupported operators fail per store. - **Advisor ordering:** if combined with `MessageChatMemoryAdvisor`, order affects whether retrieval uses the rewritten/expanded query — QAA uses the raw user text, it does not rewrite the query (that's RetrievalAugmentationAdvisor's job). **When to use:** QAA is ideal for straightforward single-query RAG. Reach for `RetrievalAugmentationAdvisor` when you need query transformation, expansion, or explicit empty-context handling.

  • What happens if the similarity search returns no documents?
    QuestionAnswerAdvisor still calls the model with an empty {question_answer_context}, so the model may answer from its own knowledge or hallucinate. Guard via the prompt template's instructions, or use RetrievalAugmentationAdvisor's ContextualQueryAugmenter which handles empty context explicitly.
  • How would you scope retrieval to a single tenant?
    Store a tenant id in each Document's metadata at ingest, then set a filterExpression (statically or per-request via the FILTER_EXPRESSION advisor param) so the SearchRequest only matches that tenant's chunks.

saying these in an interview costs you the question

  • Saying QuestionAnswerAdvisor rewrites or expands the query (it uses the raw user text; rewriting is RetrievalAugmentationAdvisor)
  • Thinking you can freely replace the template without keeping the {question_answer_context} placeholder
  • Assuming it automatically refuses to answer when nothing is retrieved
  • Believing filterExpression is a raw SQL WHERE clause rather than Spring AI's portable filter DSL

context