What does Google ADK's Runner.run_async yield, and how do you spot the final reply?
answer
- not a return value, a stream
- async for over an async generator
- tool calls are visible items too
- one boolean separates draft from answer
- partial fragments must be accumulated
basics
~20 sRunner.run_async is an async generator of Event objects — one per model chunk, tool call, tool response and agent message. Check event.is_final_response() to find the turn's answer; event.partial marks streaming fragments you should not treat as final.
solid answer
~40 s`Runner.run_async(user_id=..., session_id=..., new_message=...)` returns an async generator, so you consume it with `async for event in ...`. Each `Event` is an atomic thing that happened during the turn: a model text chunk, a function call the model requested, the function response a tool produced, a transfer to another agent, or a state change carried in `event.actions`. The Runner appends each one to the session through the session service as it goes, which is why history exists even if the client disconnects mid-turn. To render a chat reply you filter: `event.partial` is true for streaming fragments, `event.get_function_calls()` and `event.get_function_responses()` identify tool traffic, and `event.is_final_response()` marks the event whose `content` is the agent's completed answer for this turn. `Runner.run` is a synchronous wrapper over the same stream.
code
python · 28 linesimport asyncio
from google.adk.agents import LlmAgent
from google.adk.runners import InMemoryRunner
from google.genai import types
async def main() -> None:
agent = LlmAgent(name="assistant", model="gemini-2.0-flash")
runner = InMemoryRunner(agent=agent, app_name="demo")
session = await runner.session_service.create_session(
app_name="demo", user_id="u1"
)
message = types.Content(role="user", parts=[types.Part(text="Hello")])
async for event in runner.run_async(
user_id="u1", session_id=session.id, new_message=message
):
if event.partial:
continue
for call in event.get_function_calls():
print("tool:", call.name)
if event.is_final_response() and event.content:
print("answer:", event.content.parts[0].text)
asyncio.run(main())go deeper
Know that you loop over run_async with async for and that each item is an Event. Remember to create a session first and check is_final_response() before printing.
Explain the event taxonomy — model chunks, function calls, function responses, control events — and how partial differs from final. Mention that events are appended to the session as they are yielded.
Show how you would build a UI and a trace log from one stream, attribute events by author in multi-agent runs, and reason about what a mid-turn disconnect leaves in the session.
Own the operating cost of the design: per-event writes against a shared session store, backpressure and cancellation semantics, and what the event stream must expose for auditing an agent's actions in production.
## The runtime loop in one paragraph ADK's `Runner` is the thing that actually drives a turn. You give it an `app_name`, the root agent, and the services it should use (session, and optionally memory and artifact). Calling `run_async` with a `user_id`, a `session_id` and a `new_message` (a `google.genai.types.Content`) starts the loop: the Runner loads the session, hands the agent its context, and then *yields back to you* every time something happens. The agent does not return a string at the end — it emits a stream of `Event` objects, and the final answer is just one of them. ## What an Event carries An `Event` has an `author` (the literal string `"user"` or the emitting agent's name), an optional `content` holding `parts` of text or function payloads, an `actions` object, and flags such as `partial`. The `content` is the same `types.Content` shape the model API uses, which is why a function call and a paragraph of prose can both live in one event type. The helpers matter more than the fields. `event.get_function_calls()` returns the tool invocations the model asked for in this event; `event.get_function_responses()` returns the results being fed back. `event.is_final_response()` is the one you branch on for UI: it is true for the event that concludes the agent's turn with a user-facing answer, and false for intermediate model chunks, tool traffic and control events. Streaming fragments additionally set `event.partial` — accumulating those gives you token-by-token rendering, and treating one as final gives you a truncated answer, which is the classic first bug. ## Why a stream and not a return value Three things fall out of the streaming design. First, **observability is free**. Because tool calls and tool results are events, you can log or display the agent's reasoning trail without instrumenting anything. The dev UI is built on exactly this stream. Second, **history is written as you go**. The Runner appends events to the session through the session service during the turn, not after it. If the process dies halfway through a long tool sequence, the session still contains everything that had already happened. Third, **side effects are transported, not hidden**. State changes and artifact saves ride on `event.actions` (`state_delta`, `artifact_delta`), so "the agent changed something" is a visible item in the same stream as "the agent said something". ## Consuming it correctly The canonical loop looks like this: ``` async for event in runner.run_async(user_id=uid, session_id=sid, new_message=msg): if event.partial: stream_to_ui(event) elif event.is_final_response() and event.content: show(event.content.parts[0].text) ``` A few practical points. You must create the session first — `await session_service.create_session(app_name=..., user_id=...)` — and pass its `id`; the Runner will not invent one for you. The `new_message` is a `types.Content` with `role="user"`, not a bare string. And you must drain the generator: abandoning it part-way leaves the turn incomplete, which is very different from cancelling a blocking call that had already finished server-side. `Runner.run` exists as a synchronous convenience for scripts and notebooks; it wraps the same async machinery, so anything you learn about events applies to both. Behaviour of the loop can be tuned with a `RunConfig` passed to the call, which is where streaming mode and per-run limits live. ## Failure modes worth naming Agents that call tools produce many events before the answer; UIs that print every event with text in it show the user the model's intermediate musings. Conversely, code that only looks at the last event yielded can miss the final response when control events follow it — branch on `is_final_response()` rather than on position. And in a multi-agent setup, events from sub-agents carry that sub-agent's name in `author`, so filtering by author is how you attribute output correctly rather than assuming every event came from the root agent.
- Why does the Runner append events to the session as it yields them rather than at the end of the turn?So history survives a crash or a client disconnect mid-turn, and so state changes carried in `event.actions.state_delta` are committed at the point they happen rather than in one risky batch. It also means a second reader of the session sees the conversation progressing. The cost is more writes to the session store, which is a real consideration on a database-backed service.
- In a multi-agent ADK app, how do you tell which agent produced a given event?`event.author` holds the emitting agent's `name` (or `"user"` for the incoming message), so events from a sub-agent are attributed to it rather than to the root agent. Control events such as a transfer also surface in `event.actions`. Filtering the stream by author is how a UI shows per-agent traces instead of one undifferentiated log.
- What happens if you break out of the async for loop early?You stop consuming the generator, so the turn does not run to completion — remaining model output and tool calls are not produced. Anything already yielded was already appended to the session, so you are left with a truthful but partial history. If you need an early stop, prefer a termination condition the agent understands over abandoning the stream.
saying these in an interview costs you the question
- Expecting run_async to return a string answer
- Rendering partial events as if they were the final reply
- Assuming the last yielded event is always the answer
- Thinking tool calls are hidden from the event stream
- Passing a plain string instead of a types.Content message