In Google ADK, how does MemoryService recall differ from the session's event history?
answer
- one thread versus everything before it
- exhaustive replay versus selective search
- someone has to ingest the session
- recall arrives as a tool call
- the in-memory implementation is a stub
basics
~20 sSession history is this thread's events, reloaded in full every turn. A MemoryService is a separate searchable store of past sessions: you ingest finished ones with add_session_to_memory, and the agent queries them on demand through the load_memory tool.
solid answer
~40 sThey are different services on the `Runner` with different lifetimes. The `SessionService` owns one conversation: its `Session` holds the ordered `Event` list that is loaded and given to the model every turn, so it grows and eventually competes for context. The `MemoryService` is the cross-session store — you call `add_session_to_memory(session)` when a conversation ends to ingest it, and later `search_memory(app_name=..., user_id=..., query=...)` returns matching fragments. Agents reach it by adding the `load_memory` tool, which makes recall a model decision: it calls the tool when it thinks something was said before. `InMemoryMemoryService` does basic keyword matching and is a prototyping stub; `VertexAiMemoryBankService` backs it with the managed service for real semantic recall. Nothing is ingested automatically — if you never add sessions, the tool finds nothing.
code
python · 29 linesfrom google.adk.agents import LlmAgent
from google.adk.memory import InMemoryMemoryService
from google.adk.runners import Runner
from google.adk.sessions import InMemorySessionService
from google.adk.tools import load_memory
memory_service = InMemoryMemoryService()
agent = LlmAgent(
name="assistant",
model="gemini-2.0-flash",
instruction=(
"Call load_memory only when the user refers to something "
"from an earlier conversation."
),
tools=[load_memory],
)
runner = Runner(
app_name="support_bot",
agent=agent,
session_service=InMemorySessionService(),
memory_service=memory_service,
)
async def close_conversation(session) -> None:
# Nothing is remembered until this happens.
await memory_service.add_session_to_memory(session)go deeper
Know that session history is this conversation and memory is a separate store for past ones, and that the agent reaches memory through the load_memory tool.
Explain the explicit ingestion step with add_session_to_memory, the search_memory signature's scoping arguments, and why the in-memory implementation only does keyword matching.
Show operational judgment: when to ingest, how to instruct the model so recall is neither constant nor never, and how stale recalled facts can beat a correct value held in state.
Own memory as a data-governance surface — retention and deletion across session, memory and artifact stores, tenant and user scoping as an authorization concern, and the cost curve of recall on every turn.
## Two services, two questions ADK separates "what happened in this conversation" from "what do we know from previous conversations", and gives each its own service on the `Runner`. The `SessionService` answers the first. A `Session` is a thread: an id, an event list, a state dict. Every turn, that history is loaded and shapes the request sent to the model. It is exhaustive and ordered, and it is *not* selective — a 300-turn session is a 300-turn session, which is why long threads become a context and cost problem rather than a memory solution. The `MemoryService` answers the second. It is a searchable corpus over material from other, usually completed, sessions for the same user or app. It is selective by construction: a query returns fragments, not a transcript. ## The ingestion step people forget Memory is not a side effect of running an agent. The service exposes `add_session_to_memory(session)`, and something in your application has to call it — typically when a conversation is closed, or on a schedule that sweeps finished sessions. Until then, `search_memory` over that user returns nothing, and the usual bug report is "the load_memory tool always says it found nothing". The corollary is that ingestion is a policy decision you own: which sessions are worth remembering, and what should never be ingested at all. ## How the agent reaches memory The idiomatic path is the built-in `load_memory` tool. You import it and put it in the agent's `tools` list, and the model decides when to call it, passing a query. Recall then behaves like any other tool call: it appears in the event stream as a function call and a function response, it costs a model round trip, and it is visible in a trace. There is also a preload-style tool that injects memories ahead of the model call instead of waiting for the model to ask; that trades the round trip for unconditional context spend. Because it is a tool, its usefulness depends on the instruction. An agent with `load_memory` and no guidance will either never call it or call it constantly. A line such as "call load_memory when the user refers to something from a previous conversation" is doing real work. ## Implementations and what they honestly do `InMemoryMemoryService` stores ingested sessions in the process and matches queries by basic keyword overlap. It is a stub for tests and demos: no embeddings, no ranking to speak of, and it dies with the process. Reading a demo transcript where it worked well and concluding that ADK gives you semantic memory out of the box is a mistake. `VertexAiMemoryBankService` delegates to Google's managed memory service, which extracts and consolidates durable facts and serves genuine semantic retrieval. That is the production shape, and it comes with the corresponding coupling to Vertex AI and its data-handling terms. ## Memory versus user: state A fair follow-up is "why not just use `user:` state?" Because they solve different shapes of problem. `user:` state is a handful of known keys you decide to carry — preferred language, tier, current order. Memory is open-ended: you do not know in advance which sentence from three weeks ago will matter, so you keep the material and search it. In practice a good system uses both: state for the small set of facts the agent should always have, memory for everything else, retrieved on demand. ## Operating concerns **Scoping and leakage.** `search_memory` is called with an `app_name` and a `user_id`; getting that wiring wrong is how one user's history reaches another. Treat memory scope with the same seriousness as an authorization check. **Staleness.** Ingested sessions record what was true then. A user who changed address a year ago has both facts in memory, and retrieval has no inherent sense of which one is current. Where correctness matters, keep the authoritative value in state or your own database and use memory for colour and context. **Deletion.** A user's right to be forgotten now spans two stores plus artifacts. Any retention design has to cover the session store, the memory backend and anything you ingested from them. **Cost.** Recall is a tool call and its result enters the context; unbounded recall on every turn is a quiet way to double token spend without improving answers.
- An agent has the load_memory tool but always reports finding nothing. What do you check first?Whether anything was ever ingested. `MemoryService` is not populated by running the agent; your application must call `add_session_to_memory(session)` for completed sessions. After that, check that the search is scoped to the same `app_name` and `user_id` the sessions were stored under, and — if you are on `InMemoryMemoryService` — that the process has not restarted, since that implementation keeps everything in memory.
- When would you prefer user: session state over a MemoryService for something the agent must remember?When the fact is small, known in advance and must be correct every turn — language, subscription tier, the id of the order under discussion. State is a deterministic lookup with no retrieval risk and no extra round trip. Memory is for open-ended material you cannot enumerate ahead of time, where an occasional miss is acceptable and a search is worth its cost.
- What are the risks of an agent calling load_memory on every turn?Latency and tokens: each call is a model round trip plus retrieved text entering the context, on turns that did not need it. It also raises the chance of pulling in stale or irrelevant material that misleads the answer. Constrain it with instruction — recall only when the user references the past — or use a preload approach with a tight budget when recall is genuinely always relevant.
saying these in an interview costs you the question
- Assuming sessions are ingested into memory automatically
- Treating InMemoryMemoryService as semantic retrieval
- Confusing loading session history with searching memory
- Ignoring app_name and user_id scoping on memory search
- Trusting recalled facts as current without a freshness check