skip to content

In a CrewAI Task, what is expected_output for and why is it required?

level: middleimportance: must knowfreq 72%

answer

  1. the second mandatory string on a task
  2. what done looks like, not what to do
  3. it goes into the prompt, not into validation
  4. it is the loop's convergence criterion
  5. numbers beat adjectives

basics

~20 s

expected_output is the natural-language contract for a finished task: it is injected into the agent's prompt, tells the agent when to stop iterating and what shape the final answer takes, and becomes the text downstream tasks and reviewers judge the result against. CrewAI requires it on every Task.

solid answer

~50 s

A `Task` carries two mandatory strings: `description` (what to do) and `expected_output` (what a finished answer looks like). CrewAI concatenates both into the prompt, so `expected_output` is the agent's stopping condition — the loop keeps going until the model believes it has produced something matching that description, then emits it as the final answer. It is required precisely because an agent loop with no acceptance criterion does not converge: it either stops on the first plausible sentence or burns iterations. Practically, treat it as a spec rather than a wish: state the artifact, the size, the structure and the must-have fields ("a JSON object with keys title, summary and three source URLs"), not "a good summary". A vague `expected_output` is the single most common cause of unusable crew output, and it is also what the next task in the chain and any guardrail are implicitly checked against.

go deeper

for a junior

Know that every task needs both description and expected_output, and be able to rewrite a vague criterion like 'a good summary' into something countable such as 'five bullets, one sentence each'.

for a middle

Explain that expected_output is prompt text, that it acts as the agent's stopping condition, and that it costs tokens on every loop iteration — then show how tightening it changes agent behaviour.

for a senior

Demonstrate the contract view: the upstream expected_output is the downstream task's input format, and a guardrail is its executable form. Talk about keeping prose criteria and schemas consistent so you do not get schema-valid but empty output.

for a principal

Own the specification discipline across a crew — who writes these contracts, how they are reviewed, and how you keep them from drifting away from the code that parses the results as the pipeline grows.

## The two strings Every CrewAI `Task` is built from `description` and `expected_output`. Constructing a task without `expected_output` fails validation — this is deliberate, not an oversight. `description` is the instruction; `expected_output` is the acceptance criterion. Splitting them is a design decision that separates *the work* from *what done looks like*, and it is why CrewAI tasks read like tickets rather than prompts. ## How it reaches the model CrewAI assembles a prompt for the owning agent from the agent's role/goal/backstory, the task description, any context from upstream tasks, the available tool descriptions, and the expected output. The expected output lands near the end, in the section that tells the model how to present its final answer. So it is not metadata that CrewAI inspects — it is prompt text the model reads, which has two consequences: it costs tokens on every iteration of the agent's loop, and it is only as effective as the model's instruction-following. ## Why it functions as a stopping condition An agent runs a loop: think, optionally call a tool, observe, repeat, and eventually produce a final answer. Something has to decide *eventually*. In CrewAI that decision is the model's own judgment of whether it has produced the expected output. If the criterion is "a report", almost any paragraph qualifies and the agent stops early. If it is "exactly five findings, each with a one-sentence rationale and the source type", the model has a checklist to run against its own draft, and — importantly — a reason to go back to a tool when an item is missing. Tightening `expected_output` is usually a cheaper fix for a lazy agent than raising the iteration cap. ## What good expected_output looks like Specify, in roughly this order: 1. **Artifact type** — a markdown report, a JSON object, a bulleted list, a single sentence. 2. **Size** — number of items, word count, section count. Numbers are the highest-leverage part. 3. **Structure** — the sections or keys, named. 4. **Content constraints** — what must be present (dates, citations, units) and what must not (speculation, preamble). 5. **Anti-preamble clause** — "output only the JSON, no commentary" if the value is machine-consumed. Bad: "A thorough analysis of the topic." Good: "A markdown document with an H2 per competitor (3 total), each containing a 2-sentence positioning summary and a bulleted list of 2 pricing facts with dates." ## Its second life: downstream and review `expected_output` does not stop mattering when the task ends. The finished `TaskOutput` object carries the `expected_output` string alongside the produced text, so callbacks, logs and evaluation harnesses can see the contract next to the result. Downstream tasks that consume this task's output through context are, in effect, relying on that contract holding — a chain is only as stable as the tightest description in it. And when you attach a guardrail, the guardrail is the *executable* version of what `expected_output` states in prose. ## The relationship to structured output If you also set `output_pydantic` or `output_json`, CrewAI will convert the final answer into that schema after the agent finishes. That does not make `expected_output` redundant: the schema constrains *shape*, `expected_output` steers *content and effort*. In practice you write both, and you keep them consistent — a schema demanding five items next to an expected_output asking for "a few" produces retries and blank fields. ## Common failure modes - **Duplicating the description.** If `expected_output` restates the instruction, it adds tokens and no acceptance criterion. - **Describing the process, not the artifact.** "Search the web thoroughly" belongs in `description`. - **Over-specifying for a small model.** A twelve-clause contract given to a weak model produces partial compliance; either simplify or enforce with a schema plus a guardrail. - **Letting it drift from the consumer.** When downstream code parses the output, the parser and the expected_output must change together. ## What interviewers are checking That you understand `expected_output` as prompt-level control with a concrete job (convergence and shape), not as documentation; that you can turn a vague criterion into a specific one on the spot; and that you know where it stops being enough — at which point structured output and guardrails take over.

  • An agent keeps returning a two-line answer when you wanted a full report. Do you raise max_iter or rewrite expected_output?
    Rewrite `expected_output` first. A short answer usually means the agent believed it had satisfied the criterion, so more iterations just re-derive the same conclusion at extra cost. Give it a countable target — section count, item count, required fields — so the model can detect its own draft is incomplete. Raise the iteration cap only when logs show the agent running out of steps mid-work.
  • If you already set output_pydantic, is expected_output still doing anything?
    Yes. The schema fixes the shape of the final object; `expected_output` still drives how much work the agent does and what goes *into* those fields. A schema with a `findings: list[str]` field will happily accept one shallow finding. Keep the prose criterion aligned with the model — same counts, same required content — or you get schema-valid but empty results.
  • How does expected_output interact with the next task in a chain?
    The downstream task receives this task's produced text as context, so the upstream contract effectively becomes the downstream task's input format. If task B's description assumes a bulleted list, task A's `expected_output` has to promise one. When a chain breaks, compare the two strings before touching the agents — mismatched contracts between adjacent tasks are the usual cause.

saying these in an interview costs you the question

  • Treating expected_output as documentation the framework only stores
  • Restating the description instead of describing the artifact
  • Describing the process rather than the finished output
  • Assuming an output schema makes the prose criterion unnecessary
  • Fixing shallow answers by raising the iteration cap first

context