skip to content

In LangChain v1, how do you wire a retriever into a question-answering chain?

level: middleimportance: must knowfreq 62%

answer

  1. Three calls: retrieve, format, prompt
  2. The helper is pre-1.0 packaging
  3. Retrieval here is unconditional
  4. Keep the Documents, not just the string
  5. k times chunk size is a token bill

basics

~20 s

Retrieve, format, prompt: call retriever.invoke(question), join the returned Documents' page_content into a context string, and render it into a prompt template beside the question. The pre-1.0 create_retrieval_chain helper packaged that shape and belongs to the legacy compatibility surface, not to v1's slim core.

solid answer

~40 s

The pattern is three explicit steps. Call `retriever.invoke(question)` to get `list[Document]`; format those Documents into a single context string, usually by joining `page_content` with a separator and a source marker; then render both into a prompt template — a `ChatPromptTemplate` with `{context}` and `{question}` — and call the model. You return the answer **and** the Documents, because citations and evaluation both need the retrieved set. The pre-1.0 `create_retrieval_chain(retriever, combine_docs_chain)` helper wrapped exactly this, exposing `input`, `context` and `answer` keys, and it always retrieved on every invocation. LangChain v1 slimmed the `langchain` package and moved that legacy chain namespace into a separate classic compatibility package, so new code composes the steps directly instead. Composing it yourself is not more work; it is the same three calls with the seams visible.

code

python · 17 lines
python
from langchain_core.prompts import ChatPromptTemplate

prompt = ChatPromptTemplate.from_template(
    "Answer using only the context. If it does not cover the question, say you do not know.\n\n"
    "Context:\n{context}\n\nQuestion: {question}"
)

def answer(question: str):
    docs = retriever.invoke(question)
    if not docs:
        return "No relevant documents found.", []
    context = "\n\n".join(
        f"[{i}] {d.metadata.get('source', '?')}: {d.page_content}"
        for i, d in enumerate(docs, start=1)
    )
    msg = model.invoke(prompt.format_messages(context=context, question=question))
    return msg.content, docs

go deeper

for a junior

Be able to name the three steps — retrieve, format into the prompt, call the model — and say that the retrieved Documents carry the context.

for a middle

Explain what the legacy helper packaged, its input/context/answer keys, and why v1 favours composing the steps explicitly.

for a senior

Show the operational instincts: token budgeting before the call, abstaining on empty retrieval, citation-shaped formatting, and returning evidence for debugging.

for a principal

Decide where retrieval belongs in the product's control flow at all — always-on chain versus a decision the system makes — and set the grounding and citation contract the whole org builds against.

## The shape of the thing Stripped of framework vocabulary, retrieval-augmented question answering is: 1. `docs = retriever.invoke(question)` 2. `context = format(docs)` 3. `answer = model.invoke(prompt.format_messages(context=context, question=question))` Everything else is packaging. Understanding that the packaging is thin is most of what an interviewer is checking, because engineers who learned the helper first often cannot say what happens between input and output. ## What the legacy helper did `create_retrieval_chain(retriever, combine_docs_chain)` was the pre-1.0 constructor. Given a question under the `input` key it ran the retriever, put the Documents under `context`, ran a documents chain (usually built by `create_stuff_documents_chain(llm, prompt)`, which stuffs every document into one prompt) and returned a dict carrying `input`, `context` and `answer`. Its predecessor, `RetrievalQA`, returned sources under a `source_documents` key when asked to — which is why that key still shows up in old code and old answers. Two behaviours of that helper are worth stating explicitly because they are what people forget: - **It always retrieves.** There is no conditional. A greeting, a follow-up like "and the second one?", a question about the conversation itself — all of them hit the vectorstore and inject four chunks of unrelated text. Deciding *whether* to retrieve is a different control-flow problem and is not something this chain does. - **It stuffs.** The whole retrieved set goes into one prompt. That is the correct default and the reason `k` is a token-budget decision rather than a quality dial; it also means one oversized chunk can overflow the context window at request time rather than at build time. ## Why v1 pushes you to compose it yourself LangChain 1.x slimmed the top-level `langchain` package considerably, and the legacy chain constructors moved out into a separate classic compatibility package. In interviews, saying "`create_retrieval_chain` is the pre-1.0 shape, and in v1 I compose the steps" reads as current; presenting it as the recommended modern API reads as training-data-shaped. The practical argument for composing is the seams. When you write the three steps, you can: - **Format documents deliberately.** Numbering the chunks and prefixing each with its source lets the model produce inline citations you can verify: `[1] handbook.pdf p.12`. A generic join gives the model no handle to cite with. - **Budget tokens before the call.** Count the formatted context and truncate or drop the lowest-ranked documents rather than discovering the overflow as an API error. - **Short-circuit on empty retrieval.** If the retriever returns nothing — very possible with a score threshold — you can answer "I don't have anything on that" without paying for a generation that will hallucinate. - **Return the evidence.** Keeping the Documents alongside the answer is what makes citations, offline evaluation and incident debugging possible. A chain that returns only a string throws away the only artefact that explains why the answer was what it was. ## The prompt matters more than the wiring The grounding instruction is doing the real work: tell the model to answer only from the supplied context, and to say it does not know when the context does not cover the question. Without that, the model blends its parametric knowledge with the retrieved text and you lose the property that made retrieval worth building. ## Conversational queries One wrinkle appears the moment there are two turns. "How much is it?" is unretrievable on its own — the referent lives in the previous turn. The standard fix is a rewrite step that condenses the history plus the new message into a standalone search query before it reaches the retriever, so the vectorstore sees "how much does the premium support plan cost" instead of a pronoun. That is an extra model call before retrieval, and it is why conversational RAG latency is structurally higher than single-shot. ## Failure modes - Retrieved context far under the intended size because `chunk_size` was in characters, not tokens. - Confident answers to off-topic questions, because retrieval is unconditional and the prompt never told the model it may abstain. - No citations available at answer time, because the Documents were discarded after formatting. - Context-window overflow in production but not in testing, because `k` chunks of *worst-case* size were never summed.

  • Why does a follow-up question like 'and how much does it cost?' retrieve badly?
    Because the retriever only sees that string, and the referent of 'it' lives in the previous turn. The embedding of a pronoun-laden fragment lands nowhere useful. The standard fix is a query-rewriting step that condenses the conversation history plus the new message into a standalone search query before retrieval, at the price of one extra model call per turn.
  • How do you get verifiable citations out of this pattern?
    Format the context with explicit markers — number each chunk and prefix it with its source and page from the Document metadata — and instruct the model to cite those markers inline. Then return the Document list alongside the answer so the application can resolve each marker back to a real source and, if you want, verify that every cited id was actually in the retrieved set.
  • What breaks first when you raise k from 4 to 20?
    Cost and latency immediately, on every request, since all twenty chunks are prompt tokens. Then answer quality, because relevant material gets buried mid-context where models attend least, and because more chunks means more chances for a plausible-but-wrong passage to be quoted. If you want a wide net, retrieve wide and then compress or rerank down before the prompt.

saying these in an interview costs you the question

  • Presents create_retrieval_chain as the current v1 API
  • Thinks the chain decides whether retrieval is needed
  • Returns only the answer string and cannot cite
  • Never instructs the model to abstain when context is thin
  • Forgets that k chunks all land in one prompt

context