skip to content

How does RunnableWithMessageHistory feed prior turns into a LangChain LCEL chain?

level: middleimportance: must knowfreq 68%

answer

  1. wraps a runnable, adds a store
  2. session id rides in config
  3. factory returns the history object
  4. three keys: input, history, output
  5. loads before, appends after, never trims

basics

~20 s

RunnableWithMessageHistory wraps a runnable. On each call it uses the session id from config['configurable'] to fetch that conversation's history, prepends the stored messages to the input, invokes the inner chain, then appends the new input and output messages back to the store.

solid answer

~40 s

`RunnableWithMessageHistory` (in `langchain_core.runnables.history`) is the LCEL-era memory wrapper. You give it an inner runnable and a `get_session_history` callable; at invoke time it reads the session id from `config={"configurable": {"session_id": ...}}`, calls your factory, loads the stored messages and injects them, runs the inner chain, then writes the new human input and the AI output back to the same history object. When the inner chain takes a dict you tell it which keys to use with `input_messages_key`, `history_messages_key` and `output_messages_key`; the history key must match the placeholder variable in your prompt. Invoking without the configurable key raises rather than silently running stateless. Two things to say out loud: it never trims, so the prompt grows every turn, and LangChain 1.x steers durable multi-step state toward graph-level persistence instead.

code

python · 16 lines
python
from langchain_core.chat_history import InMemoryChatMessageHistory
from langchain_core.messages import AIMessage, HumanMessage
from langchain_core.runnables import RunnableLambda
from langchain_core.runnables.history import RunnableWithMessageHistory

store: dict[str, InMemoryChatMessageHistory] = {}

def get_session_history(session_id: str) -> InMemoryChatMessageHistory:
    return store.setdefault(session_id, InMemoryChatMessageHistory())

fake_model = RunnableLambda(lambda msgs: AIMessage(f"saw {len(msgs)} messages"))
with_history = RunnableWithMessageHistory(fake_model, get_session_history)

cfg = {"configurable": {"session_id": "u1"}}
print(with_history.invoke([HumanMessage("hi")], config=cfg).content)
print(with_history.invoke([HumanMessage("again")], config=cfg).content)

go deeper

for a junior

Be able to say that the wrapper takes a chain plus a function returning that session's history, and that the session id is passed in config under configurable.

for a middle

Explain the load-invoke-append cycle and the roles of input_messages_key, history_messages_key and output_messages_key, including what a key mismatch looks like.

for a senior

Talk about what the wrapper costs at turn fifty, where you insert trimming, and the write-path edge cases: failed turns, retries and streams that never finish.

for a principal

Frame the boundary: a flat message log fits chat transcripts, while branching, resumable workflow state belongs in a persistence layer built for it, and say which you would adopt.

## What the wrapper does `RunnableWithMessageHistory` turns a stateless chain into a conversational one without changing the chain. It is a runnable itself, so it keeps `invoke` / `stream` / `batch` and composes like any other. Per call it performs four steps: resolve the session, load history, run the inner chain with history injected, and persist the turn. ## The session comes from config, not from the input The conversation key travels in the runnable config, not the payload: `chain.invoke(payload, config={"configurable": {"session_id": "abc"}})`. The wrapper hands that value to your `get_session_history` callable, which returns a `BaseChatMessageHistory` for that conversation. Because the factory is yours, the backend is your choice — an in-process dict for tests, Redis or SQL in production — and the chain above never knows the difference. Forgetting the configurable value is an error, not a silent stateless run. One session id is rarely enough in a real product; you usually want user plus conversation. `history_factory_config` accepts a list of `ConfigurableFieldSpec` objects, which lets the factory take several parameters (for example `user_id` and `conversation_id`) that callers pass under `configurable`. ## Wiring the keys If the inner runnable takes a bare list of messages, no keys are needed — history is simply prepended. If it takes a dict (the common case with a prompt template), you must say which key is which. - `input_messages_key` — the dict key holding this turn's user input. - `history_messages_key` — the key the loaded history is injected under. It must match the placeholder variable your prompt reserves for prior messages; a mismatch means the model sees an empty conversation while the store quietly fills up. - `output_messages_key` — needed when the chain returns a dict rather than a message or string, so the wrapper knows which field to persist. ## What gets persisted After the inner chain returns, the wrapper appends both sides of the turn. Consequences worth stating in an interview: a failed turn may leave the input recorded without an answer; retries can double-write; and streaming persists once the stream completes, so a client that disconnects mid-stream can produce a half-recorded turn. If exactly-once semantics matter, write to the store yourself instead of relying on the wrapper. ## The cost it hides Nothing here bounds size. Every turn reloads the full history and re-sends it, so prompt tokens grow roughly linearly with conversation length, latency drifts up, and eventually the provider rejects the request for exceeding the context window. Memory in LangChain is a context-window budget problem, and this wrapper deliberately does not solve it — you compose a trimming or summarizing step into the inner chain, ahead of the prompt, so it runs on every invocation. ## Where it sits in LangChain 1.x This is the successor to the pre-1.0 `ConversationChain` plus `BaseMemory` pairing, which coupled memory to a specific chain class. The wrapper is composable instead: any runnable, any store. For richer state — branching, resuming a partially completed run, human approval mid-flight — LangChain now points at graph-level persistence rather than a flat message log, and the honest interview answer names that boundary: this wrapper is right for chat transcripts and thin for workflow state.

  • How do you key history by user and conversation instead of a single session id?
    Pass history_factory_config a list of ConfigurableFieldSpec entries naming the extra fields, for example user_id and conversation_id, and write get_session_history to accept those parameters. Callers then supply both under config['configurable'], and the factory scopes storage by user, which also stops a client-supplied conversation id from reaching another user's transcript.
  • The store fills up but the model behaves as if every turn is the first. What is wrong?
    Almost always a key mismatch: history_messages_key does not match the variable your prompt reserves for prior messages, so history is injected under a name the template never renders. The wrapper still persists the turn, which is why the store looks healthy. Check that the two names agree.
  • Where does trimming belong when you use this wrapper?
    Inside the wrapped runnable, ahead of the prompt, so it runs on every invocation against the freshly loaded history. The wrapper itself never bounds anything: it loads all stored messages and appends more. Trimming after the prompt, or once at construction time, does nothing for turn fifty.

saying these in an interview costs you the question

  • Passes session id in the input payload instead of config
  • Expects the wrapper to trim history automatically
  • Thinks one global history object serves all users
  • Confuses this with the pre-1.0 ConversationChain plus BaseMemory pairing
  • Assumes the prompt placeholder name does not have to match

context