skip to content

What does wrapping each crew in a CrewAI Flow step give you over one big crew?

level: seniorimportance: should knowfreq 55%

answer

  1. a step is just Python calling kickoff
  2. the seam between crews is code
  3. skipped work costs nothing
  4. retry one crew, not the run
  5. deterministic outside, agentic inside

basics

~20 s

A flow step is ordinary Python: it calls Crew().kickoff(inputs=...) itself, stores the result on self.state, and a router branches on it. Sequencing, retries and skipping become code you control instead of decisions a manager model makes each run.

solid answer

~50 s

Inside a `@start` or `@listen` method you just build a crew and call `crew.kickoff(inputs={...})`; the returned `CrewOutput` — `.raw` for text, `.pydantic` when the task declared a schema — gets written to `self.state`, and the next step or a `@router` consumes it. What that buys is deterministic seams *between* crews: you can run cheap non-LLM code in between (validate, dedupe, hit an API, write to a database), skip an expensive crew entirely when a router says the cached answer is good, retry one crew without re-running the others, and attribute token cost per crew instead of per run. The cost is that you wrote the sequence yourself, so the system can no longer reorder work it did not anticipate. That is the right trade exactly when the sequence is known — which is most production workflows.

code

python · 20 lines
python
from crewai import Agent, Crew, Task
from crewai.flow.flow import Flow, listen, router, start


class ResearchFlow(Flow):
    @start()
    def research(self):
        analyst = Agent(role="Analyst", goal="Summarize a topic", backstory="Careful reader")
        task = Task(description="Summarize {topic}", expected_output="Two paragraphs", agent=analyst)
        result = Crew(agents=[analyst], tasks=[task]).kickoff(inputs={"topic": self.state["topic"]})
        self.state["summary"] = result.raw
        return result.raw

    @router(research)
    def check(self, summary):
        return "expand" if len(summary) < 200 else "done"

    @listen("done")
    def finish(self):
        return self.state["summary"]

go deeper

for a junior

Know that a flow step is plain Python and can simply call kickoff on a crew, put the result on self.state, and let the next step read it. There is no special wiring to learn.

for a middle

Explain what the seam between crews is for: non-LLM work, structured output on state, and a router branching on the result. Be able to name what CrewOutput carries and why you store the distilled field rather than the raw text.

for a senior

Argue the trade concretely — skipped crews cost nothing, retries become per crew, token usage is attributable, and prompts stop growing because you choose what each crew sees. Acknowledge that you gave up runtime adaptability to get it.

for a principal

Own the split as a design position: freeze the decisions that measurement shows never vary, keep the model where the ordering genuinely depends on findings, and be able to defend the boundary with cost, variance and on-call debuggability rather than taste.

## The two ways to compose Given four units of work, CrewAI offers two shapes. One crew with many tasks and, in a hierarchical process, a manager model deciding who does what next. Or several small crews, each invoked from its own flow step, with the ordering written in Python. Both run the same agents against the same models. The difference is *who decides the order* — and therefore what varies run to run, what you can test, and what you pay. ## The mechanics There is no special API for this. A flow step is a normal method, so it constructs (or looks up) a crew and calls `crew.kickoff(inputs={...})`. Inputs typically come off `self.state`, so the flow decides what context this crew sees. The call returns a `CrewOutput`: `.raw` for the text, `.pydantic` or `.json_dict` when the final task declared structured output, `.tasks_output` for per-task results, and `.token_usage` for what it cost. The step writes what matters onto `self.state` and returns whatever the next step needs. That is the whole integration surface, and its plainness is the point: between two crew calls you are in Python, with no model in the loop. ## What the seams buy **Deterministic work between crews.** Validation, deduplication, a database write, an idempotency check, a call to an internal service. Doing these as flow steps costs zero tokens. Doing them as agent tasks costs tokens *and* can be skipped or fumbled by the model. **Skipping.** A router can decide that the cached answer is fresh, that the request failed policy checks, or that the first crew already answered well enough — and the expensive crew is simply never invoked. Inside one hierarchical crew, an agent that decides not to act has still read its task and reasoned about it, on the clock. **Retry granularity.** If the extraction crew produced malformed JSON, you can re-run *that* crew — possibly with a stricter prompt or a different model wired into the retry path — without re-running the research crew that cost ten times as much. Retry logic lives in the step, where you can bound it with a counter on state. **Cost attribution.** Each `CrewOutput` carries its own token usage, so per-crew cost lands on state and in your logs. A single large crew gives you one number and a guess. **Context control.** Each crew gets exactly the inputs the step passes it. Inside a hierarchical crew, context accumulates across tasks whether or not the next task needs it, and prompts grow. Flows let you deliberately *not* forward the previous crew's raw output. **Testability.** A step is a method. You can construct the flow, set state, call the step with the crew stubbed, and assert what it wrote. Manager-driven delegation has no comparable seam. ## What it costs You wrote the sequence, so the system cannot adapt to a situation you did not anticipate. If the work genuinely requires discovering the plan at runtime — an open-ended research request where the next move depends on what was just found — hardcoding steps produces a rigid pipeline that fails on the interesting inputs. There is also more code: models, routers, state fields, tests. And you own the failure semantics, because a crew that raises inside a step propagates out of `kickoff()` with no retry unless you wrote one. ## The practical middle The shape that holds up is **deterministic outside, non-deterministic inside**. The flow encodes the parts of the workflow that are genuinely fixed — intake, three phases, persistence, notification — and each phase is a crew whose agents have real latitude within their phase. Judgment sits where judgment is needed, in bounded units whose cost you can measure, and the skeleton is code. A good migration heuristic: run the hierarchical crew, log the manager's delegation decisions across a few dozen runs, and look at which ones never varied. Those are edges you should be writing as flow steps; you are paying a model to re-derive them every run. The ones that did vary are the places to keep the model in charge. ## Operational notes Put a run correlation id — the flow state's generated id — into whatever you log around each `kickoff` call, so a support ticket resolves to one execution and its per-crew outputs. Keep crew construction out of module import time if the crew reads state-dependent configuration, so each step builds the crew it actually needs. And be deliberate about what a step forwards: the most common regression is a step passing an entire previous `CrewOutput.raw` into the next crew's inputs, which quietly doubles the prompt of every downstream call.

  • How do you retry just one crew inside a flow?
    Write the retry in the step: keep an attempt counter on self.state, call kickoff again on failure or on a failed validation of the output, and route to a fallback branch once the counter passes your limit. Nothing in the flow layer retries a step for you, and that explicitness is why the granularity is per crew instead of per run.
  • What is the risk of passing a crew's full raw output into the next crew's inputs?
    Prompt growth that compounds. Each crew's output becomes the next one's prompt tokens, so a four-crew flow that forwards everything can pay for the first crew's output three more times. Forward the distilled field the next crew actually needs — ideally structured output rather than raw text — and keep bulky artifacts behind a reference.
  • When is one hierarchical crew still the better choice?
    When the plan genuinely has to be discovered at runtime — open-ended research or triage where the next move depends on what was just found, and the useful orderings are too many to enumerate. Hardcoding steps there produces a pipeline that handles the demo and fails on the interesting inputs; let the model sequence and spend your effort on bounding cost and iterations.
  • How do you decide which delegation decisions to freeze into flow steps?
    Measure. Run the crew, log what the manager chose across a few dozen realistic runs, and separate decisions that never varied from those that did. The invariant ones are edges you are paying a model to rediscover every run — move them into steps and routers. Leave the genuinely variable ones with the model.

saying these in an interview costs you the question

  • Saying a flow needs a special API to invoke a crew
  • Assuming a failed crew inside a step retries automatically
  • Believing a skipped branch still costs tokens
  • Passing every previous output forward to be safe
  • Claiming flows replace crews rather than sequence them

context