In DeepEval, how do you build a multi-turn ConversationalTestCase?
answer
- failures that live between turns
- role and content, in order
- do not flatten the transcript
- the assistant's replies are the output
- what a conversational golden can safely store
basics
~20 sConstruct ConversationalTestCase with a turns list of Turn objects, each carrying a role of user or assistant and its content. Optional scenario and expected_outcome describe what the conversation was supposed to achieve, and conversational metrics score the exchange as a whole.
solid answer
~50 s`ConversationalTestCase` from `deepeval.test_case` takes `turns=[Turn(role=..., content=...)]`, where `role` is `"user"` or `"assistant"`, plus optional `scenario`, `expected_outcome` and `user_description` describing what the exchange was meant to accomplish. It exists because a single-turn `LLMTestCase` cannot express the failures that only appear over several turns — forgetting something stated three turns back, contradicting an earlier commitment, or drifting out of role. Flattening the transcript into one `input` string destroys the turn boundaries the conversational metrics rely on. On the dataset side, `ConversationalGolden` stores the *setup* rather than a fixed transcript: a `scenario`, an `expected_outcome`, and a description of the user, because the assistant's replies differ on every run so a frozen transcript is not a reusable golden. Version note: in deepeval 4.1.9 the turn list is `turns` of `Turn` objects; SDKs before 3.0 used a list of `LLMTestCase` messages instead.
code
python · 12 linesfrom deepeval.test_case import ConversationalTestCase, Turn
test_case = ConversationalTestCase(
scenario="User asks to change a delivery address after the order shipped",
expected_outcome="Assistant explains it cannot be changed and offers a carrier redirect",
turns=[
Turn(role="user", content="i need to change my delivery address, order 88213"),
Turn(role="assistant", content="Order 88213 has already shipped, so the address is locked."),
Turn(role="user", content="so what can i do"),
Turn(role="assistant", content="I can request a redirect with the carrier for order 88213."),
],
)go deeper
Know that a multi-turn conversation uses ConversationalTestCase with a turns list, and that each Turn has a role of user or assistant plus its content.
Explain why the type exists — memory, consistency and role failures span turns — and name scenario and expected_outcome. Be ready to say why flattening a transcript into one input breaks the metrics.
Talk about the golden side: what can safely be stored when the assistant's replies are the output, and the determinism-versus-realism tradeoff between a scripted and a simulated user, including the per-turn cost.
Decide how much multi-turn coverage is worth its cost, which conversational behaviours are worth gating on, and how you keep multi-turn results comparable when the user side is itself model-driven.
## Why a separate test case type exists An `LLMTestCase` describes one exchange: an input went in, an output came out, score it. That is the right shape for a question-answering endpoint and the wrong shape for a chatbot, because the interesting chatbot failures are *relational* — they exist between turns, not inside one: - The user gives their order number in turn two; the assistant asks for it again in turn six. - The assistant promises a refund in turn three and denies eligibility in turn seven. - The assistant is configured as a formal banking agent and by turn nine is telling jokes. - The user's goal is never reached, though every individual reply was locally reasonable. None of these is visible in any single output. You have to look at the sequence. ## The object ``` ConversationalTestCase( turns=[Turn(role="user", content="..."), Turn(role="assistant", content="..."), ...], scenario="...", expected_outcome="...", ) ``` - `turns` — the ordered exchange. Each `Turn` carries a `role` (`"user"` or `"assistant"`) and its `content`. - `scenario` — plain-English description of the situation the conversation is exercising. - `expected_outcome` — what a successful conversation should end with. - `user_description` — who the user is, when that shapes what counts as a good exchange. Conversational metrics consume the whole test case rather than one output field, which is exactly why the structure must be preserved. ## Why not just flatten the transcript? A tempting shortcut is to join the transcript into one string, put it in `LLMTestCase.input`, and score the last reply as `actual_output`. It runs, and it measures the wrong thing. The metric can no longer tell which text the assistant produced and which the user typed, so it cannot attribute a contradiction to the assistant. It cannot tell you *when* the failure started. And the single-turn metrics you would then be applying were designed to judge one response against one input, so their scores on a mashed-up transcript are not meaningful readings of anything. ## The golden side: ConversationalGolden The dataset question is subtler than the test-case question. What is the reusable, stored artifact for a conversation? It cannot be a frozen transcript with the assistant's replies included, because those replies are the output under test — they change every run and with every model. A golden containing them would be scoring a recording, not a system. So `ConversationalGolden` stores the *setup*: the `scenario` the user is in, the `expected_outcome` that defines success, and a `user_description`. The user turns are the fixed part; the assistant turns are produced at evaluation time by driving your application through the scenario, and the resulting exchange becomes a `ConversationalTestCase`. This mirrors the single-turn split exactly — the golden holds what you maintain, the test case holds what this run produced. Driving the conversation is where multi-turn evaluation gets genuinely harder than single-turn. A scripted user replays fixed user turns regardless of what the assistant said, which is deterministic but unrealistic once the assistant asks a question the script does not answer. A simulated user is driven by a model following the scenario, which is realistic and adaptive but adds non-determinism and cost — every turn is another model call, on both sides. ## Practical notes - Keep conversations short and targeted. A twelve-turn golden that fails tells you very little about *where*; three focused goldens of four turns each localize the failure. - Write the `expected_outcome` before the transcript exists. If you cannot state what success looks like in one sentence, the scenario is too vague to score. - Budget for cost: multi-turn evaluation multiplies calls by turn count on the application side, again on the simulated-user side if you use one, and again in whatever judge the metric runs. ## Version context In deepeval 4.1.9, `ConversationalTestCase` takes `turns` of `Turn` objects. SDK versions before 3.0 modelled a conversation as a list of `LLMTestCase` messages instead, so older blog posts and snippets will not run unmodified — if you see `messages=[LLMTestCase(...)]`, you are reading pre-3.0 material and it needs porting.
- Why can't a ConversationalGolden just store the full transcript including the assistant's replies?Because those replies are the thing under test. They change with every model, prompt and temperature setting, so a stored transcript is a recording of one past run rather than a reusable standard. The golden therefore keeps the stable part — the scenario, the user's situation and the expected outcome — and the assistant turns are produced at evaluation time by driving your application through that scenario.
- What is the tradeoff between a scripted user and a simulated one when driving multi-turn evals?A scripted user replays fixed turns, so runs are deterministic and cheap, but the script breaks down the moment the assistant asks something it does not answer, and it never probes an unexpected path. A model-driven simulated user adapts and reaches realistic conversations, at the price of non-determinism between runs and roughly doubling the model calls, since both sides of every turn now cost money.
- What goes wrong if you flatten a transcript into a single LLMTestCase input?The metric loses the turn boundaries, so it cannot tell which text the assistant produced and cannot attribute a contradiction or a forgotten detail to it. You also lose the location of the failure — you learn the conversation was bad, not that it went wrong at turn six. And single-turn metrics were designed to judge one response against one input, so their scores on a concatenated transcript are not readings of anything well-defined.
- How long should a conversational golden be?Short and targeted. A twelve-turn conversation that fails tells you the exchange was bad without telling you where, and it costs twelve turns of model calls on every run. Three four-turn goldens, each exercising one behaviour — memory across turns, consistency of a commitment, staying in role — localize the failure and are far cheaper to run and to read when they go red.
saying these in an interview costs you the question
- Joins the transcript into one string and scores it as a single output
- Stores the assistant's replies inside the golden as the expected transcript
- Thinks multi-turn evaluation is just single-turn run repeatedly
- Writes twelve-turn goldens and cannot say where they failed
- Uses the pre-3.0 messages-of-LLMTestCase shape on a current SDK