skip to content

How does a CrewAI task in tasks.yaml get wired to a @task method?

level: juniorimportance: should knowfreq 55%

answer

  1. two files under a config directory
  2. the attribute becomes a parsed dict
  3. the decorator also registers the object
  4. declaration order is execution order

basics

~10 s

In a @CrewBase project, tasks.yaml stores each task's description and expected_output under a named key. A @task-decorated method returns Task(config=self.tasks_config['that_key']), and the decorator registers the returned Task into self.tasks for the crew.

solid answer

~40 s

A `@CrewBase` class points at `config/tasks.yaml` through its `tasks_config` attribute. Each top-level key in that file is one task, holding at least `description` and `expected_output`, and usually an `agent:` entry naming an `@agent` method in the same class. In Python you write a method per task decorated with `@task` that returns `Task(config=self.tasks_config["research_task"])`; the decorator both builds the object and appends it to `self.tasks`, in declaration order, so `Crew(agents=self.agents, tasks=self.tasks, ...)` picks them up without you listing them twice. Placeholders like `{topic}` inside the YAML strings are interpolated from the inputs dict you pass when the crew runs. Anything YAML cannot express — a guardrail callable, `output_pydantic`, a `callback` — you pass as an extra keyword argument alongside `config=`.

code

yaml · 14 lines
yaml
research_task:
  description: >
    Research recent developments in {topic} and collect the key findings.
  expected_output: >
    A list of exactly 5 findings, each one sentence, each naming a source type.
  agent: researcher

report_task:
  description: >
    Turn the research findings into a short briefing on {topic}.
  expected_output: >
    A markdown briefing with an intro, three sections and a conclusion.
  agent: reporter
  output_file: briefing.md

go deeper

for a junior

Be able to point at the two YAML files, name the two required strings in a task entry, and say that a @task method returns Task(config=...) which the decorator adds to self.tasks.

for a middle

Explain that tasks_config is a parsed dict by the time you index it, that declaration order is execution order in a sequential crew, and that {placeholders} come from the run inputs.

for a senior

Show where the split breaks down: guardrails, structured-output models and explicit context lists have to be Python, so be clear about which half of a task's definition each teammate owns and how prompt edits get reviewed.

for a principal

Own the argument for the layout itself — YAML buys reviewable prompt diffs and non-engineer edits, and costs an indirection that hides wiring bugs. Decide when a team is better served by plain Python task construction.

## What the YAML layout is for CrewAI lets you keep the *prose* of a crew — role, goal, backstory, task description, expected output — in YAML, and keep the *code* — Python objects, tools, callbacks, structured-output models — in a class. The motivation is that the prose is what you iterate on daily (and what a non-engineer can review), while the wiring is stable. The convention is two files, `config/agents.yaml` and `config/tasks.yaml`, next to a crew module. ## The class The class is decorated with `@CrewBase` and declares where the YAML lives: `agents_config = "config/agents.yaml"` and `tasks_config = "config/tasks.yaml"`. At runtime those attributes are replaced with the *parsed dictionaries*, so `self.tasks_config["research_task"]` is the dict for that YAML key, not a path. This is the single most common point of confusion: you write a path, you read a dict. ## One YAML key = one task A task entry carries: - `description` — what the agent should do; the instruction block that goes into the prompt. - `expected_output` — the contract for what a finished answer looks like. - `agent` — optional; the *method name* of an `@agent` in the same class, which is how the task gets an owner without importing anything. - optional extras such as `output_file`. Both long strings are normally written with YAML block scalars (`>` or `|`) so they can span lines without escaping. ## The @task decorator ``` @task def research_task(self) -> Task: return Task(config=self.tasks_config["research_task"]) ``` The decorator does two things. First, it memoizes/constructs the `Task` when the crew is assembled. Second — the part people miss — it **registers** the task into `self.tasks`, in the order the methods are declared in the class body. That is why the `@crew` method can say `tasks=self.tasks` and never enumerate them. The same mechanic applies to `@agent` and `self.agents`. Order matters: under a sequential process the declaration order of the `@task` methods is the execution order. Reordering methods reorders the crew. ## Interpolation Any `{placeholder}` inside `description`, `expected_output`, or `output_file` is filled from the inputs dictionary supplied when the crew is started (`inputs={"topic": "vector databases"}`). This is plain string interpolation over the task's fields — it is not a template engine, and a placeholder with no matching input is an error rather than an empty string. Keep the placeholder names identical across `agents.yaml` and `tasks.yaml` so one input feeds both. ## Mixing YAML and Python `config=` is a starting point, not a straitjacket. You can pass additional keyword arguments in the same constructor call, and they apply on top of the YAML: ``` @task def report_task(self) -> Task: return Task( config=self.tasks_config["report_task"], output_pydantic=Report, context=[self.research_task()], ) ``` Things that *must* live in Python because YAML cannot hold a Python object: `output_pydantic` / `output_json` model classes, `guardrail` callables, `callback` functions, tool instances, and any explicitly built `context` list of `Task` objects. ## When to skip YAML entirely For a small script or a demo, constructing `Task(description=..., expected_output=..., agent=...)` directly in Python is simpler and fully supported; `@CrewBase` is a project convention, not a requirement. The YAML layout earns its keep when (a) more than a couple of tasks exist, (b) the descriptions are long and reviewed by someone who does not read Python, or (c) you want to diff prompt changes cleanly in code review without noise from surrounding code. ## What interviewers are checking That you know the YAML is *configuration for the same `Task` object*, not a separate DSL; that `@task` is what puts the task into the crew's list; and that interpolation comes from run inputs. A candidate who thinks the YAML is parsed by the LLM, or who cannot say where `self.tasks` comes from, has only copied a template.

  • If you need a guardrail or an output_pydantic model on a YAML-defined task, where does it go?
    In the Python method, as extra keyword arguments beside `config=`. YAML can only hold data, so a callable or a Pydantic class cannot be expressed there. `Task(config=self.tasks_config["report_task"], output_pydantic=Report, guardrail=check_length)` merges the YAML fields with the Python-only ones on the same object.
  • How does a task in tasks.yaml end up owned by a specific agent?
    The YAML entry's `agent:` value is the *method name* of an `@agent` in the same `@CrewBase` class, and CrewAI resolves it when the crew is assembled. You can also omit it in YAML and pass `agent=self.researcher()` in the Python method. Under a hierarchical process a manager may assign work instead, so an explicit owner is what makes routing deterministic.
  • What happens if tasks.yaml uses {topic} but the run supplies no such input?
    Interpolation fails rather than silently substituting an empty string — the placeholder has no value to bind. Treat inputs as part of the crew's public signature: validate them at the call site, and keep placeholder names consistent between agents.yaml and tasks.yaml so a single input feeds both the role text and the task description.

saying these in an interview costs you the question

  • Thinking tasks_config stays a file path at runtime
  • Believing the LLM reads the YAML file directly
  • Listing tasks twice — in the class and in Crew(tasks=[...])
  • Assuming YAML can hold a guardrail or Pydantic class
  • Not knowing method declaration order sets execution order

context