skip to content

What does CrewAI's Crew.kickoff() return, and what does that object carry?

level: juniorimportance: should knowfreq 58%

answer

  1. prints like a string, isn't one
  2. final task text lives in raw
  3. one entry per executed task
  4. structured fields only when asked
  5. the run's cost rollup

basics

~20 s

Crew.kickoff() returns a CrewOutput, not a string. It carries raw (the final task's text), structured pydantic and json_dict forms when a task requested them, tasks_output with one TaskOutput per task, and token_usage for the whole run.

solid answer

~50 s

`crew.kickoff()` returns a `CrewOutput` object. Its `raw` attribute is the plain-text result of the final task, and printing the object gives you that same text, which is why beginners think they got a string back. The rest is what makes it useful in production. `pydantic` and `json_dict` are populated only when the final task asked for structured output; otherwise they are empty, so code that assumes a model instance will break on a plainly-configured crew. `tasks_output` is a list of `TaskOutput` objects, one per executed task, which is how you inspect intermediate work instead of guessing what happened between the first and last step. `token_usage` aggregates the run's prompt, completion and total tokens plus the number of successful requests, which is your per-run cost signal. In a service, log `token_usage`, persist `tasks_output` for debugging, and treat `raw` as the payload only after checking whether you actually wanted the structured form.

code

python · 11 lines
python
result = crew.kickoff(inputs={"topic": "vector databases"})

print(result.raw)                      # final task's text
print(len(result.tasks_output))        # one TaskOutput per executed task
print(result.tasks_output[0].raw)      # first task's own output

usage = result.token_usage
print(usage.total_tokens, usage.prompt_tokens, usage.completion_tokens)

if result.pydantic is not None:        # only when the task asked for it
    print(result.pydantic)

go deeper

for a junior

Know that kickoff returns a CrewOutput and that result.raw is the text you usually want. Being able to say 'it is an object, not a string' already answers most of this question.

for a middle

Explain the field map: raw for text, pydantic and json_dict only when the final task requested structure, tasks_output for per-task results, token_usage for the run's aggregate cost.

for a senior

Show what you do with it in production: log the token rollup as a cost metric, persist tasks_output as the trace you will debug from, and never treat completion as correctness.

for a principal

Own it as the observability contract of a crew run — the only per-run cost and trace signal the framework gives you — and decide what your platform stores, tags and alerts on before crews multiply across teams.

## The return type Every successful `crew.kickoff()`, and every element of the list returned by the batch form, is a `CrewOutput`. This matters immediately because the object's string representation is the final task's raw text — so `print(result)` looks exactly like a string, and code written on that assumption (`result.strip()`, `json.loads(result)`) fails the first time someone touches the object rather than its text. ## The fields **`raw`** — the final task's output as text. This is the headline result of the run and the field you use when the crew's job is to produce prose. **`pydantic`** — a model instance, present only when the final task was configured to produce a typed object. If you did not ask for structured output, this is empty. The correct pattern is to check before using it, not to assume. **`json_dict`** — a parsed dictionary, present under the same condition when the task was configured for JSON output. There is also a JSON-string accessor, which raises if the task never requested JSON — a deliberate loud failure rather than a silent empty string. **`tasks_output`** — a list of `TaskOutput`, one entry per task the crew executed, in execution order. Each entry carries that task's own raw text (and its structured forms if it requested them), plus identifying detail about the task. This is the single most useful debugging field in the object: when a four-task crew produces a bad final answer, `tasks_output` tells you which step went wrong, without re-running anything. **`token_usage`** — aggregated usage metrics for the whole run: prompt tokens, completion tokens, total tokens, cached prompt tokens where the provider reports them, and the number of successful requests. It is a per-run rollup across all agents, not a per-agent breakdown. ## Using it in a service A production wrapper around a crew typically does four things with the output object: 1. **Decides the payload.** Structured field if the task requested one, `raw` otherwise. Never both, and never guessing. 2. **Emits cost telemetry.** `token_usage.total_tokens` and the request count, tagged with whatever job identifier you have, so cost regressions are visible as a time series rather than a surprise invoice. 3. **Persists `tasks_output`.** Intermediate outputs are the trace you will want when a user reports a wrong answer next week. Re-running the crew reproduces nothing reliably, because model output is not deterministic. 4. **Validates before shipping.** A `CrewOutput` is only evidence that the run completed, not that the answer is correct. Nothing in the object asserts quality. ## Common misreadings - *'kickoff returns a string.'* It returns an object that prints like one. - *'pydantic is always there.'* Only when the final task requested typed output. - *'raw concatenates every task.'* It does not — `raw` is the final task's text; earlier steps live in `tasks_output`. - *'token_usage is per agent.'* It is the run-level rollup. - *'a returned CrewOutput means success.'* It means execution finished; correctness is your evaluation problem. ## Why interviewers ask Because it separates people who have run a crew in a notebook from people who have shipped one. The notebook user prints `result` and is done. The engineer who has operated a crew knows the token rollup is the only cost handle the framework gives them per run, and that `tasks_output` is the difference between debugging a pipeline and re-rolling the dice. ## Version note Describes the CrewAI 0.x `CrewOutput` shape.

  • Why is tasks_output more useful than re-running the crew when an answer looks wrong?
    Because a re-run is a different sample. LLM output is not deterministic, and in a hierarchical crew even the routing can differ, so the second run may not reproduce the fault at all. `tasks_output` is the actual trace of the run that failed: it shows which step's output was already wrong, so you fix that task's description, tools or agent instead of guessing.
  • What does token_usage let you do that provider-side billing does not?
    Attribute cost to a run while you still know its inputs. The rollup comes back attached to the result, so you can tag it with the job, tenant or crew version and watch cost per run as a metric. Provider invoices aggregate across everything your key does, which tells you the bill went up but not which crew change caused it.
  • If the final task requested structured output, is raw still populated?
    Yes — `raw` carries the model's text and the structured field carries the parsed form, so both are available. Prefer the structured field as your payload when you asked for one: it has already been parsed and validated, whereas re-parsing `raw` yourself duplicates work the framework did and reintroduces the failure mode you asked for structure to avoid.

saying these in an interview costs you the question

  • Says kickoff returns a plain string
  • Assumes result.pydantic is populated without configuring structured output
  • Thinks raw concatenates the output of every task
  • Treats token_usage as a per-agent breakdown
  • Treats a returned CrewOutput as proof the answer is correct

context