In Google ADK, how does an .evalset.json file differ from a .test.json file?
answer
- one session versus a dataset of cases
- same invocation atom underneath
- session_input seeds initial state
- the dev UI Eval tab writes one of them
- criteria live in a separate config file
basics
~20 sA test file captures one simple session and is meant to run cheaply and often, like a unit test. An eval set file holds many named eval cases, each a full multi-turn session with its own starting state, and is the integration-tier artifact you record from the dev UI.
solid answer
~40 sBoth describe recorded agent behaviour with the same building block — an invocation carrying `user_content`, the expected `final_response`, and `intermediate_data.tool_uses` listing the tool calls that should happen. The difference is granularity. A `.test.json` file is a single session, usually short, and you keep many of them next to the agent so a change to one tool breaks one small file. An `.evalset.json` is a dataset: an `eval_set_id` plus a list of `eval_cases`, each with its own `eval_id`, a multi-turn `conversation`, and optionally `session_input` seeding app name, user id and initial state. Eval sets are what you record interactively in `adk web`'s Eval tab, and they are heavier to run. Both are executed by `adk eval <agent_module> <file>` or by `AgentEvaluator.evaluate(...)` from pytest.
code
json · 19 lines{
"eval_set_id": "weather_cases",
"eval_cases": [
{
"eval_id": "asks_for_london",
"conversation": [
{
"invocation_id": "inv-1",
"user_content": {"role": "user", "parts": [{"text": "Weather in London?"}]},
"final_response": {"role": "model", "parts": [{"text": "It is 15C and cloudy in London."}]},
"intermediate_data": {
"tool_uses": [{"name": "get_weather", "args": {"city": "London"}}]
}
}
],
"session_input": {"app_name": "weather", "user_id": "u1", "state": {}}
}
]
}go deeper
Know that ADK stores recorded agent behaviour in JSON: small test files for one session, eval set files for several multi-turn cases, and that you can record a case from the dev UI.
Explain the shared invocation schema — user content, expected final response, and the tool calls under intermediate_data — and that session_input lets an eval case start from non-empty state.
Talk about how these artifacts rot: recorded phrasings break on prompt changes, recorded tool arguments break on signature changes, and live-data recordings age. Say when you re-record versus adjust criteria.
Own the split: which behaviours deserve small always-run test files, which deserve expensive multi-turn eval sets, and how folder-scoped criteria let you hold different bars for deterministic tool paths and open-ended conversation.
## Two artifacts, one schema ADK's evaluation story rests on recorded expectations. Two file shapes carry them, and interviewers ask about the difference because it reveals whether you have actually run the tooling. ## The shared unit: the invocation Whatever the file, the atom is an **invocation**: one user turn and everything the agent did in response. An invocation records `user_content` (the user's message, in the standard content/parts shape), `final_response` (the text the agent should end up producing) and `intermediate_data`, whose `tool_uses` list names each tool call with its arguments. That triple is what the scorers consume: the tool calls feed the trajectory metric, the final response feeds the text-match metric. ## Test files — small, many, cheap A test file (`*.test.json`) represents a single, simple agent–model interaction: one session, often one or two turns. The idea is unit-test ergonomics. You keep a folder of them beside the agent, each pinning one behaviour — "asks for a city and calls get_weather with it", "refuses when the tool errors" — so a failure points at one thing. They are cheap enough to run on every change during development, and because each is small, updating one after an intentional behaviour change is a small diff rather than a re-record of a long conversation. ## Eval sets — few, multi-turn, recorded An eval set file (`*.evalset.json`) is a dataset. Its top level carries an `eval_set_id` and a list of `eval_cases`; each case has an `eval_id`, a `conversation` (the list of invocations, in order) and optionally a `session_input` block giving `app_name`, `user_id` and an initial `state` dictionary. That `session_input` matters: it lets a case start from a non-empty world — a user already identified, a preference already set — which is exactly what a multi-turn conversation needs and what a bare test file does not express. Eval sets are the integration tier. A case can be ten turns long and exercise a whole session's worth of state accumulation, sub-agent transfer and tool sequencing. They are also the artifact the dev UI produces: run a conversation by hand in `adk web`, open the Eval tab, create or pick an eval set, and save the session as a case. That recording path is the practical reason eval sets exist — capturing a real session beats hand-writing a ten-turn JSON blob. ## Running either one The CLI takes both: `adk eval <path_to_agent_module> <path_to_eval_file>`. You can restrict a run to specific cases by appending case ids to the eval set path, and `--config_file_path` points at the criteria file. From Python, `AgentEvaluator.evaluate(agent_module=..., eval_dataset_file_path_or_dir=...)` runs a file or a whole directory, which is how you fold ADK evals into a pytest suite and let your existing test runner report them. ## Thresholds live beside the files Neither file carries its own pass criteria. A `test_config.json` placed in the same folder supplies a `criteria` block that applies to the files in that folder. That folder-scoped design is deliberate: you can hold strict criteria over a folder of tool-trajectory tests and looser criteria over a folder of open-ended conversational cases without editing any case. ## Practical failure modes Recorded expectations rot. A recorded `final_response` is one phrasing the model happened to produce; a prompt tweak that improves answers can fail the case on wording alone. Recorded `tool_uses` assume the tool surface is stable — rename an argument and every case that touched it fails at once, which is informative but noisy. And a case recorded from a session whose tools hit live third-party services encodes that day's data into the expected response. The discipline is to record small cases for anything you want stable, keep multi-turn eval sets few and deliberately chosen, and re-record rather than patch when behaviour intentionally changes.
- What does the session_input block on an eval case buy you?It seeds the session before the case runs — app name, user id and an initial `state` dictionary. That lets a case start from a world that already has a user preference or an identifier set, so you can test the branch that depends on prior state without replaying the turns that produced it. Test files, being single simple sessions, typically start empty.
- How would you run ADK eval cases inside an existing pytest suite?Use `AgentEvaluator.evaluate(agent_module=..., eval_dataset_file_path_or_dir=...)` from `google.adk.evaluation.agent_evaluator` in an async test. It accepts a single file or a directory, so a folder of test files becomes one pytest case, and failures surface through your normal test reporter instead of a separate CLI run.
- A prompt change made every recorded case fail on wording. What do you do?First check whether the tool trajectories still pass — if they do, the agent's behaviour is intact and only phrasing moved. Then either relax the response threshold in the folder's `test_config.json` criteria, or re-record the expected responses. Patching individual expected strings by hand is the option that quietly desyncs the suite from the agent.
saying these in an interview costs you the question
- Thinking an eval set is just a bigger test file with no state seeding
- Believing pass thresholds are stored inside the eval case
- Hand-writing multi-turn eval sets instead of recording them
- Assuming recorded expected responses stay valid across prompt changes
- Not knowing evals can be driven from pytest as well as the CLI