skip to content

Your ADK agent runs on several replicas — which session, memory and artifact services do you pick?

level: principalimportance: should knowfreq 34%

answer

  1. per-process stores fail behind a load balancer
  2. three services, three decisions
  3. operate it yourself or accept the coupling
  4. append-only means a growth plan
  5. deletion has to reach every store

basics

~20 s

Every in-memory service must go: they are per-process, so a request landing on another replica sees no session. Pick a shared backend for each — a database or Vertex session service, a durable memory service, and GcsArtifactService for blobs — and decide retention per store.

solid answer

~40 s

The `Runner` takes three independent stores, and each is a separate decision. Sessions: `InMemorySessionService` breaks the moment two replicas exist, because a user's next turn may land elsewhere and find no session; use `DatabaseSessionService(db_url=...)` against a shared Postgres you operate, or `VertexAiSessionService` if you are already committed to Vertex AI Agent Engine. Memory: `InMemoryMemoryService` is a keyword stub that dies with the process, so production recall means a managed backing such as `VertexAiMemoryBankService`. Artifacts: `GcsArtifactService` puts bytes in Cloud Storage where every replica and a later session can read them, while `InMemoryArtifactService` loses them. Beyond availability, weigh the operational bill — you own schema, backups, retention and deletion on a self-run database, and you inherit a vendor's data terms on the managed path.

go deeper

for a junior

Know that the in-memory services are for local development only and that a deployed agent needs a shared backend for sessions.

for a middle

Name the durable alternatives for each store — a database or Vertex session service, a managed memory service, GcsArtifactService — and explain why a second replica breaks the in-memory ones.

for a senior

Reason about the failure modes you inherit: store outages, partial turns, concurrent writers on one session, and unbounded session growth without a retention policy.

for a principal

Own the commitment itself — self-operated versus managed per store, tenancy and scoping as an authorization boundary, deletion spanning sessions, memory and artifacts, and the cost curve of every persisted key being reloaded each turn.

## Three stores, three decisions ADK's `Runner` is constructed with a session service, and optionally a memory service and an artifact service. Nothing forces them to be equally durable, and treating them as one "persistence setting" is the first mistake. Sessions hold conversation threads and small state; memory holds searchable material from past sessions; artifacts hold bytes. They have different sizes, different access patterns and different deletion obligations. ## What multi-replica actually breaks The in-memory implementations are per-process dictionaries. With one process they behave like a real backend, which is why they survive so far into a project. Put two replicas behind a load balancer and the failure is intermittent rather than obvious: a conversation works while sticky routing holds and forgets everything when a request lands elsewhere or a pod is rescheduled. The same applies to artifacts (a file uploaded on replica A is missing on B) and to memory (ingested sessions are visible only to the process that ingested them). So the baseline is: **any horizontally scaled deployment needs external storage for every store it actually uses.** "Actually uses" matters — an agent with no artifacts does not need an artifact service at all, and adding one you do not need is a store to secure and delete from for no benefit. ## Self-operated versus managed `DatabaseSessionService` with a SQLAlchemy URL gives you a store you can inspect, back up, and keep inside your own compliance boundary. In exchange you own migrations of ADK's schema across framework upgrades, connection pooling under concurrency, and growth: every turn appends events, so a chatty agent's session table grows monotonically until you define a retention policy. Nobody does that for you. `VertexAiSessionService` and `VertexAiMemoryBankService` remove that operational load and, for memory, give real semantic recall rather than the keyword stub. The cost is coupling: your conversational data lives in a specific vendor's service under its terms and regions, and portability away from it becomes a project rather than a config change. That is a genuine architectural commitment, not an implementation detail, and it should be made explicitly. A common middle position is a self-operated Postgres for sessions — the store most likely to hold regulated content and most likely to be queried in incident review — with managed memory, where the semantic capability is the point and the data is derived rather than authoritative. ## Sizing and hygiene **Keep state small.** Everything persisted in session state is reloaded with the session on every subsequent turn. Ids, flags and short preferences belong there; documents and payloads do not. Use `temp:` for intermediates that must never be written and artifacts for bytes, storing only the filename in state. **Bound history.** Sessions are append-only. Decide when a thread ends and a new one begins, and what happens to old ones: archived, ingested into memory, deleted. "Forever" is a decision too, and usually the wrong one. **Scope carefully.** Session and memory access are keyed by `app_name` and `user_id`. Those arguments are effectively an authorization boundary; deriving them from anything a client can influence is how one user's history reaches another. Multi-tenant deployments should also remember that `app:`-prefixed state is shared across everyone under that app name. **Plan deletion once, across three stores.** A user-deletion request has to reach sessions, memory and artifacts, plus any of your own tables holding ids. Designing that after launch is much harder than designing it before. ## Failure behaviour Externalising storage introduces a new failure mode: the store can be down. Decide up front what the agent does when the session service is unavailable — fail the request loudly, or serve a degraded stateless turn — and make sure a partial turn leaves coherent history, since events are appended as they are produced rather than in one commit. Also consider concurrency: two workers driving one session both append, and last-delta-wins on a contested state key, so either serialise per session or keep a session to a single writer. ## The answer an interviewer is listening for Not a service name, but the reasoning: which stores this agent genuinely uses, what each must survive, who operates it, how big it gets, how it is scoped, and how it is deleted. Naming `DatabaseSessionService` is the easy half; owning retention, tenancy and the vendor commitment is the half that distinguishes a lead.

  • Sticky sessions would keep the in-memory service working. Why is that not a good answer?
    It converts a correctness property into a routing coincidence. Any deploy, pod eviction, autoscaling event or node failure loses conversations, and the loss is silent and user-visible. It also blocks rolling upgrades and makes replicas non-interchangeable. Stickiness is a legitimate optimisation on top of shared storage — never a substitute for it.
  • Your session table grows without bound on a chatty agent. What do you do?
    Treat retention as a design decision rather than a cleanup task: define when a thread ends, ingest anything worth remembering into the memory service, then archive or delete old sessions on a schedule. Keep durable state small so each session row is light, and push bulky content into artifacts. Trimming history mid-thread is not the fix, since the log is what state is derived from.
  • How do you decide between a self-operated database and the managed Vertex services?
    By where the data must live and who should carry the operations. Self-operated keeps conversational data inside your boundary and lets you inspect and back it up, at the cost of owning schema upgrades, scaling and retention. Managed removes that load and gives real semantic memory, but couples your conversational data to one vendor's service, regions and terms. Split the decision per store rather than making it once.

saying these in an interview costs you the question

  • Relying on sticky routing to keep in-memory sessions alive
  • Treating the three services as a single persistence switch
  • Assuming a managed service removes retention obligations
  • Letting session tables grow with no end-of-thread policy
  • Deriving user_id for session lookup from client-supplied input

context