What does Google ADK's ParallelAgent isolate between branches, and what do they share?
answer
- one thing per branch, one thing shared
- branches cannot read each other's messages
- the shared store is where they collide
- distinct keys per branch, always
- the merge is an agent you add yourself
basics
~20 sParallelAgent gives each sub-agent its own branch of the conversation history, so branches do not see each other's messages. They still run against one shared session state, so two branches writing the same key race and the loser's work vanishes silently.
solid answer
~50 s`ParallelAgent` runs every agent in its `sub_agents` list concurrently over a single invocation. What it isolates is **context**: each branch gets its own branch of the event history, so branch A's tool calls and replies are not in branch B's prompt — that is the point, since fan-out is only useful when the branches are independent. What it does **not** isolate is **session state**: all branches read and write the same store, so distinct `output_key` values per branch are mandatory, and any shared key is a genuine race whose winner is whichever branch finished last. Two more operational consequences: the branches hit your model provider simultaneously, so a fan-out of eight is eight concurrent requests against one rate limit; and there is no built-in join step, so the merge is a separate agent you place after the parallel block, usually by nesting the `ParallelAgent` inside a `SequentialAgent`.
code
python · 19 linesfrom google.adk.agents import LlmAgent, ParallelAgent, SequentialAgent
market = LlmAgent(name="market", model="gemini-2.0-flash",
instruction="Summarise market risk for the company named by the user.",
output_key="summary_market")
legal = LlmAgent(name="legal", model="gemini-2.0-flash",
instruction="Summarise legal risk for the company named by the user.",
output_key="summary_legal")
fan_out = ParallelAgent(name="analysts", sub_agents=[market, legal])
merger = LlmAgent(
name="merger", model="gemini-2.0-flash",
instruction=("Combine the analyses below into one brief. Note any section that is missing.\n\n"
"MARKET:\n{summary_market?}\n\nLEGAL:\n{summary_legal?}"),
output_key="risk_brief",
)
risk_review = SequentialAgent(name="risk_review", sub_agents=[fan_out, merger])go deeper
Know that ParallelAgent runs its sub_agents at the same time rather than one after another, and that each branch should write its own distinct output_key.
Explain the split precisely — event history is per branch so branches do not see each other's messages, while session state is shared — and describe the merger agent placed after the parallel block.
Raise the operational consequences unprompted: silent key collisions, races on shared state, a burst of concurrent model calls against one rate limit, worst-case tail latency, and building partial-failure tolerance into each branch.
Own fan-out width as a capacity decision with measured concurrency limits, and set the team convention for branch key naming and partial-result contracts so a degraded branch produces a documented marker rather than an ambiguous absence.
## What fan-out buys you `ParallelAgent` exists for one reason: independent work that would otherwise be serialised. Three market analyses that each take six seconds cost eighteen seconds in a `SequentialAgent` and roughly six in a `ParallelAgent`. Since neither composition type calls a model itself, the whole saving is real wall-clock latency, not tokens — you still pay for every branch's inference. ## The isolation: separate branches of history Each sub-agent runs in its own branch of the invocation's event history. Practically, this means branch A's messages, tool calls and tool results are not injected into branch B's prompt. That is deliberate and it is what makes concurrency safe at the prompt level: if branches shared one linear history, they would be interleaving messages into each other's context in a non-deterministic order, and each branch's prompt would balloon with work it does not care about. The corollary is that a branch cannot see, and must not depend on, what a sibling branch is doing. Any design where branch B reads what branch A "just produced" is not parallel — it is a sequence you have accidentally made concurrent, and it will pass in testing whenever A happens to finish first. ## The sharing: one session state Session state is not branched. All branches read and write the same store, and that is the single most important operational fact about `ParallelAgent`: - **Distinct output keys are mandatory.** If two branches both declare `output_key="summary"`, both write it, and the value you keep is whichever branch completed last. Nothing errors. You lose an entire branch's work and the run looks successful. Name keys per branch — `summary_market`, `summary_legal`, `summary_technical`. - **Reads are unsynchronised.** A branch that reads a key another branch is writing gets whatever happens to be there at that instant. If a placeholder in a branch's instruction references a key that only a sibling produces, you have built a race. - **State is also how the merge works.** Because state survives the parallel block, a downstream agent can read all branch keys at once. That is the intended join. ## There is no join step `ParallelAgent` finishes when its branches finish; it does not synthesise anything. The merge is your job and it is an agent. The canonical structure is a `SequentialAgent` whose members are the `ParallelAgent` and then a synthesiser `LlmAgent` whose instruction templates in each branch's key: `SequentialAgent(sub_agents=[ParallelAgent(sub_agents=[a, b, c]), synthesiser])` Write the synthesiser's placeholders in the optional form so a branch that produced nothing degrades to a partial answer rather than breaking prompt assembly. ## The operational hazards a senior is expected to raise **Concurrency against shared limits.** Eight branches are eight simultaneous model requests plus whatever their tools call. You are now generating a burst against a per-minute quota, a downstream API, or a connection pool that was sized for serial traffic. Fan-out width is a capacity decision, not a stylistic one — and rate-limit rejections under a burst look like flaky agents, not like a quota problem, in the trace. **Failure semantics.** Treat an exception in one branch as something that can take down the parallel step rather than quietly yielding a partial result. If best-effort is what you want, make each branch tolerant on its own — catch inside the branch's tools and have the branch write an explicit "unavailable" value to its key — so the merger sees a missing-data marker rather than the whole invocation failing. Designing for partial results is a decision you make in the branches, not something you get for free. **Tail latency, not average.** A parallel block costs as much as its slowest branch. Adding a fourth branch that is usually fast but occasionally times out converts a reliable six-second step into an occasionally-thirty-second one. Fan-out improves the mean and worsens the tail; if you have an end-to-end SLO, bound each branch's own work rather than assuming concurrency has made latency a non-issue. **Debuggability.** Branched history is a gift when reading a trace — each branch reads as its own coherent conversation — but it means "what did the agent see" now has three answers. When you reconstruct an incident, reconstruct per branch, and remember that the shared state a branch read may have been written by a sibling mid-flight. ## The summary an interviewer is listening for "Context is per branch, state is shared." Everything else — distinct output keys, an explicit merger agent, rate-limit bursts, worst-case tail latency, deliberate partial-failure handling — falls out of that one sentence, and being able to derive it rather than recite it is what separates the senior answer from the documentation answer.
- How do you combine the branches' results into one answer?Add a merger agent after the parallel block — typically `SequentialAgent(sub_agents=[ParallelAgent(...), synthesiser])`. Each branch writes a distinct `output_key`, and the synthesiser's instruction templates all of those keys in, using the optional `{key?}` form so a branch that produced nothing degrades to a partial answer instead of breaking prompt assembly.
- Is a fan-out of twenty branches a good idea?Rarely. Twenty branches are twenty simultaneous model calls plus their tool traffic, all against one rate limit and one connection pool, and the step still costs as much as the slowest branch. Beyond a handful, batch the work inside fewer branches, or bound concurrency at the tool layer, and treat fan-out width as a capacity decision you have measured.
- Why not just let branches read each other's state to coordinate?Because there is no ordering between branches, so any such read is a race that passes in testing whenever the writer happens to finish first. If branch B needs branch A's output, they are not independent and belong in a `SequentialAgent`. Shared state under `ParallelAgent` is for collecting results, not for communicating between concurrent branches.
- How do you get a best-effort result when one branch fails?Build the tolerance into the branch. Have its tools catch their own errors and have the branch write an explicit "unavailable" value to its own key, so the merger sees a missing-data marker. Do not assume the parallel step will hand you a partial result automatically — an unhandled exception in a branch can take the step down with it.
saying these in an interview costs you the question
- Believing each branch gets its own copy of session state
- Letting two branches share one output_key
- Expecting ParallelAgent to merge the branch outputs for you
- Having one branch read a key another branch is writing
- Assuming fan-out is free because the framework handles concurrency