A CrewAI crew's token spend tripled after a refactor — how do you find where it went?
answer
- compare requests versus tokens per request
- which task inflated
- read the config diff first
- the manager talks, and it bills
- rate cap while you investigate
basics
~20 sStart from CrewOutput.token_usage to confirm the regression per run, then attribute it: capture a run log with output_log_file and verbose, inspect tasks_output, and check whether process, planning or memory changed. Cap the rate with max_rpm while investigating.
solid answer
~50 sFirst, quantify: `CrewOutput.token_usage` gives prompt, completion and total tokens plus successful request count for one run, so run the old and new configurations on the same fixed input and compare. A jump in *request count* points at more loops or delegation turns; a jump in *prompt tokens* at the same request count points at bigger prompts. Then attribute. Turn on `verbose` and write a full trace with `output_log_file` so you can see every agent turn, and read `CrewOutput.tasks_output` to find which task inflated. Use `step_callback` or `task_callback` to record per-step timing and content programmatically instead of eyeballing console output. Then check the usual suspects in the diff: a switch to `Process.hierarchical` adds manager delegation and review turns around every task; `planning=True` adds a call per kickoff and lengthens every task prompt; enabling memory adds retrieved context to prompts; a task now carrying more tools means longer tool schemas in every call. Set `max_rpm` while you investigate so a runaway is slow rather than expensive.
code
python · 18 linesfrom crewai import Crew, Process
def on_step(step_output):
print("step:", step_output)
crew = Crew(
agents=[researcher, writer],
tasks=[research, write],
process=Process.sequential,
verbose=True,
output_log_file="crew-run.log",
step_callback=on_step,
max_rpm=20,
)
result = crew.kickoff(inputs={"topic": "vector databases"})
u = result.token_usage
print(u.total_tokens, u.prompt_tokens, u.completion_tokens, u.successful_requests)go deeper
Know that CrewOutput carries token_usage and that verbose logging exists, so a cost question starts with measuring one run rather than guessing.
Explain what inflates a crew's prompts: hierarchical manager turns, planning text appended to every task, memory context, and tool schemas carried on every call. Point at tasks_output to localise the growth.
Show the method — fixed input set, compare request count against tokens per request, attribute to a task with a run log and callbacks, then decide whether the extra spend bought measurable quality. Cap with max_rpm while investigating.
Own the practice: per-run token usage logged with the crew version, alerts on the moving mean, and a standing evaluation set that reports quality and cost together so configuration changes are traded deliberately rather than discovered on an invoice.
## Step 1 — make the regression measurable You cannot debug a cost problem you have not pinned to a number. Every `crew.kickoff()` returns a `CrewOutput` whose `token_usage` reports prompt tokens, completion tokens, total tokens, cached prompt tokens where the provider reports them, and the count of successful requests. Run the previous configuration and the new one on the same fixed input set and diff those figures. The shape of the diff already narrows the search: - **More requests, similar tokens per request** → more turns. Agents are looping, or a manager is delegating and reviewing. - **Same requests, more prompt tokens** → prompts got bigger. Injected plan text, retrieved memory, larger tool schemas, more context threaded between tasks. - **More completion tokens** → agents are being asked for, or are volunteering, longer output; expected-output wording or a model change is the usual cause. ## Step 2 — attribute it to a task `CrewOutput.tasks_output` gives one `TaskOutput` per executed task in order. Read it: a task whose output ballooned, or a task that clearly re-did work an earlier task already produced, is your candidate. For the turn-by-turn view, set `verbose` on the crew and send the trace to a file with `output_log_file` so you can grep it rather than scroll a terminal. For anything repeated, wire `step_callback` and `task_callback` to record structured records per step and per task — that turns a one-off investigation into a standing metric. ## Step 3 — check the configuration diff In practice a tripling almost always comes from one of a short list of crew-level changes: **Process.** Moving from `Process.sequential` to `Process.hierarchical` puts a manager in front of and behind every task: a delegation decision, often clarifying questions to coworkers, then a review. That is easily a doubling on its own, and it is the single most common cause of the symptom. **Planning.** `planning=True` adds a planning call on every kickoff *and* appends the generated plan to each task's description, so the plan's tokens are paid once per task, not once per run. **Memory.** Enabling crew memory means retrieved context is added to prompts, and embedding calls are made. It is a real feature with a real per-run price. **Tools.** Every tool available to an agent contributes its schema to that agent's prompt on every call. Adding a rich toolset to a chatty agent multiplies across all its turns. **Agent loop bounds.** A raised per-agent iteration ceiling lets a struggling agent keep trying; combined with a hierarchical manager that re-delegates on a poor result, this is how a run goes from expensive to pathological. ## Step 4 — bound the blast radius While you investigate, set `max_rpm` on the crew. It caps requests per minute across the crew, which does not reduce total tokens for a completed run but does convert a runaway from a fast expensive failure into a slow visible one — and it keeps you inside provider rate limits during repeated test runs. `cache=True` helps when agents repeat identical tool calls, which is common in retry-heavy crews, though it does not deduplicate LLM calls themselves. ## Step 5 — decide what to keep The outcome should be a deliberate trade, not a rollback reflex. If hierarchical routing genuinely improved output quality on your evaluation set, the extra tokens are the price of that quality and belong in your cost model. If the manager consistently routes the same task to the same agent, the routing was static and the tokens bought nothing — encode the order in the task list and go back to sequential. If planning did not move your quality metric, turn it off; if it did, consider writing its output into the task descriptions by hand so you stop paying for regeneration. ## Step 6 — make it not recur The engineering answer, and what an interviewer is listening for, is that per-run token usage should be a logged metric with the crew version attached, not something you discover from an invoice. Log `token_usage` on every run, alert on the per-run mean moving, and pin a small evaluation set that runs both quality and cost on every configuration change. Cost regressions then arrive as a graph with a commit next to them. ## Version note Describes the CrewAI 0.x crew observability surface: `token_usage`, `tasks_output`, `verbose`, `output_log_file`, `step_callback`, `task_callback`, `max_rpm`, `cache`.
- Does max_rpm reduce the token cost of a run?No. It caps requests per minute across the crew, so a run takes longer rather than costing less — the same calls still happen. Its value is containment and rate-limit safety: a runaway loop becomes slow and visible instead of burning budget at full speed, and repeated test runs stay under the provider's ceiling while you investigate.
- Why does adding tools to an agent raise cost even when those tools are never called?Because tool schemas are part of the prompt. Every call that agent makes carries the definitions of all its available tools, so an unused toolset is a fixed tax on every turn of that agent's loop. On a chatty agent in a multi-task crew that multiplies quickly, which is why tool assignment should be per-agent and minimal rather than a shared catalogue.
- If hierarchical mode doubled cost but improved quality, what would you do?Verify the quality gain on a fixed evaluation set rather than on impressions, then decide with numbers. If the win is real, put the cost in the model and look for cheaper routing — a smaller manager model, fewer clarifying turns, pinning the steps whose assignment never varies. If the manager always routes identically, the routing was static and belongs in the task list.
saying these in an interview costs you the question
- Assumes token cost scales only with the number of tasks
- Thinks max_rpm reduces total tokens rather than request rate
- Ignores that a hierarchical manager adds LLM turns per task
- Believes unused tools on an agent are free
- Debugs by re-running and eyeballing output instead of reading token_usage