skip to content

What does setting planning=True on a CrewAI Crew do, and what does it cost?

level: middleimportance: nice to knowfreq 33%

answer

  1. an LLM call before any task runs
  2. the plan is appended to descriptions
  3. it does not reorder anything
  4. paid again on every kickoff
  5. planned without seeing any data

basics

~20 s

planning=True makes CrewAI call a planning LLM before execution, produce a step-by-step plan for the crew's tasks, and append that plan to each task's description. It costs an extra LLM call per kickoff and can bake in wrong assumptions.

solid answer

~50 s

With `planning=True`, CrewAI runs a planning phase before any task executes. It sends the crew's tasks and agents to a planning model — `planning_llm` if you set one, otherwise the default — and asks for a concrete step-by-step plan. That plan is then appended to each task's description, so every agent starts with explicit instructions about how its step should be carried out. What it buys is consistency on crews where agents otherwise improvise their approach: vague task descriptions get concretised once, up front, in a way that references the other tasks. What it costs is an extra LLM call on every kickoff, longer prompts for every subsequent task (the plan text rides along in each description), and a plan written with zero execution feedback — no tool results, no retrieved documents. If the planner guesses wrong about the data, every agent inherits that wrong assumption. It is a prompt-engineering aid, not a control-flow feature: it does not reorder tasks or change who runs them.

code

python · 12 lines
python
from crewai import Crew, Process

crew = Crew(
    agents=[researcher, writer],
    tasks=[research, write],
    process=Process.sequential,
    planning=True,
    planning_llm="gpt-4o",   # plan once on a strong model
)

result = crew.kickoff(inputs={"topic": "vector databases"})
print(result.token_usage.total_tokens)   # compare against planning=False

go deeper

for a junior

Know that planning=True makes CrewAI generate a step-by-step plan before the run and add it to the tasks, and that it costs an extra model call.

for a middle

Explain the mechanism precisely: a planning LLM (planning_llm if set) is called once per kickoff and its plan is appended to each task's description, so control flow is unchanged and every task's prompt grows.

for a senior

Argue the cost case — a call per kickoff plus inflated prompts on every task — and describe the A/B you would run against a fixed input set comparing quality and token_usage before keeping the flag on.

for a principal

Position it as dynamic prompt engineering: decide organisationally when generated plans are worth regenerating per run versus writing the plan into task descriptions once, and require evidence rather than defaults.

## What the flag actually does `Crew(planning=True)` adds a pre-execution phase. Before the first task runs, CrewAI collects the crew's task descriptions and its agents' roles and goals, sends them to a planning model, and asks for a step-by-step plan describing how the work should be done. The resulting plan is appended to each task's description, so when an agent finally receives its prompt, that prompt now contains both the original task text and the planner's instructions for that step. You choose the planning model with `planning_llm`. Leaving it unset uses the default model. Setting it lets you plan on a stronger model than the one your workers use — often the best value, since planning happens once per run while worker calls happen many times. ## What it does not do This is where candidates go wrong. Planning does **not**: - reorder the tasks list; - reassign tasks to different agents; - turn a sequential crew into a dynamic one; - re-plan when a task fails or returns something unexpected. It is text injected into prompts before the run. Control flow still comes from `process` — sequential order, or a manager's delegation under hierarchical. Confusing `planning=True` with `Process.hierarchical` is the classic error: one writes a better prompt, the other changes who decides. ## The costs **A call per kickoff.** Planning is not cached across runs, so every kickoff pays for it. On a crew invoked thousands of times in a batch, that is thousands of planning calls. **Longer prompts everywhere.** Because the plan is appended to each task description, every task's prompt grows by the plan's length. On a six-task crew you pay for that text six times, not once, and it eats context you might need for retrieved content. **Blind planning.** The planner sees configuration, not reality. It has not called a tool, not read a document, not seen how large the input actually is. A plan that says 'summarise the three most relevant results' when retrieval returns none is now an instruction every agent tries to follow. **Harder debugging.** The prompt an agent receives is no longer the text you wrote; it is your text plus generated text that differs between runs. When output changes and your code did not, the plan is a new suspect. ## When it earns its place Use it when task descriptions are genuinely underspecified and you cannot easily improve them by hand — exploratory crews, crews assembled from user-supplied goals, or crews where the sequence of steps really does depend on the input. Skip it when your tasks already have precise descriptions and expected outputs, because then you are paying a model to restate what you already wrote, on every single run. The senior framing: planning is dynamic prompt engineering. If you run the same crew shape repeatedly, a plan you write once by hand is cheaper, more stable and easier to review than a plan regenerated on every kickoff. If the shape genuinely varies per input, generated planning is doing work you cannot pre-write. ## Measuring the decision Run the crew both ways on a fixed input set. Compare output quality with whatever evaluation you have, and compare `CrewOutput.token_usage` across the two configurations. If quality is flat and tokens are up, the flag is pure cost. That measurement takes an afternoon and it is exactly the evidence an interviewer wants to hear you would gather. ## Version note Describes the CrewAI 0.x `planning` / `planning_llm` crew options.

  • Is planning=True an alternative to Process.hierarchical?
    No — they solve different problems. Planning injects generated instructions into task descriptions before the run and leaves control flow untouched. Hierarchical changes who decides which agent runs which task, at run time, with a manager agent. You can enable both, and then you are paying for a plan and a manager, which is worth doing only if you can show each is earning its tokens.
  • Why can a generated plan make results worse rather than better?
    Because it is written blind. The planner sees task descriptions and agent roles, not tool results or retrieved documents, so it can assert steps that do not match reality — assuming data exists, assuming a tool returns a particular shape. Agents then follow those instructions faithfully, and a confident wrong plan is harder to recover from than a vague description would have been.
  • How would you decide whether to keep planning enabled for a crew that runs thousands of times a day?
    Measure it. Run a fixed evaluation set with the flag on and off, compare quality scores and compare `CrewOutput.token_usage` totals. At that volume the planning call and the inflated per-task prompts are a real line item, so keep it only for a measurable quality win. If the crew shape is stable, hand-writing the plan into the task descriptions once usually wins outright.

saying these in an interview costs you the question

  • Thinks planning=True reorders tasks or reassigns agents
  • Believes the plan is generated once and cached across runs
  • Confuses planning with hierarchical delegation
  • Assumes the planner can see tool output or retrieved documents
  • Ignores that the plan inflates every task's prompt

context