skip to content

How does CrewAI cache tool results, and what does cache_function control?

level: middleimportance: should knowfreq 44%

answer

  1. on by default at the crew level
  2. keyed by tool plus arguments
  3. a predicate over (args, result)
  4. errors and volatile data must not stick
  5. side-effecting tools want it off

basics

~20 s

CrewAI caches tool results during a crew run, keyed by tool and arguments, so an identical repeat call returns the stored value instead of re-executing. Setting cache_function on a tool lets you decide per result whether it is stored; Crew(cache=False) disables caching entirely.

solid answer

~50 s

Caching is on by default: `Crew(cache=True)` is the default, and CrewAI's cache handler stores each tool result keyed by the tool and the arguments it was called with. When an agent — often a *different* agent in the same crew — calls that tool with identical arguments again, it gets the cached observation back without your `_run` ever executing. That is a real saving, because agent loops repeat calls constantly. The hook you get is `cache_function`, a callable taking `(args, result)` and returning a boolean; assign it on the tool object and CrewAI consults it to decide whether *this* result is worth storing. Use it to refuse caching failures, empty results, or anything time-sensitive: caching an error string means every later attempt in the run replays the error. If the tool is non-idempotent — it sends mail, files a ticket, mutates state — you generally want `Crew(cache=False)` or a `cache_function` that always returns False, so the model's second call actually performs the action.

code

python · 14 lines
python
from crewai.tools import tool


@tool("Quote price")
def quote_price(sku: str) -> str:
    """Return the current quoted price for a SKU, or SKU=unavailable."""
    return f"{sku}=41.90"


def cache_only_real_quotes(args, result):
    return not result.endswith("=unavailable")


quote_price.cache_function = cache_only_real_quotes

go deeper

for a junior

Know that CrewAI caches tool results within a crew run so repeated identical calls do not re-execute, and that Crew(cache=False) turns that off.

for a middle

Explain the key (tool plus arguments), the default (cache=True), and the cache_function(args, result) -> bool hook, with a concrete predicate such as refusing to cache empty or error results.

for a senior

Show what a default-on cache does to real systems: swallowed retries after a cached failure, stale volatile lookups, and side-effecting tools whose second call silently never happens. Say how you would debug a suspect observation.

for a principal

Own the policy rather than the flag: which tool classes may be cached at all, idempotency keys as the contract for write tools, and how caching interacts with your call metrics and cost model when instrumentation sits below the cache.

## Why a tool cache exists at all An agent loop is repetitive by construction. The LLM decides, calls a tool, reads the observation, decides again — and it very often re-issues a call it already made, either because the earlier observation scrolled out of usefulness or because the prompt nudged it back to the same idea. In a multi-agent crew the repetition is worse: two agents with the same tool will independently look up the same fact. CrewAI's cache exists to make that repetition cheap. ## The default and the switch `Crew` takes a `cache` parameter and it defaults to `True`. With caching on, CrewAI's cache handler records each tool invocation's result keyed by the tool and the arguments that produced it. A later invocation with the same tool and the same arguments is answered from that store; your `_run` body is not entered. The cache lives with the crew execution — it is an in-process store, not a durable shared cache across deployments, so do not reason about it as if it were Redis. ## cache_function: the per-result decision Because "should this be cached" is a property of the *result*, not just of the tool, CrewAI exposes `cache_function`. It is a callable with the signature `(args, result) -> bool`. Assign it on the tool object (`my_tool.cache_function = my_predicate`) or declare it on a `BaseTool` subclass. When it returns False, the result is not stored, and the next identical call executes for real. Useful predicates in practice: - **Refuse to cache failures.** If the tool returns an error string or an empty payload, caching it poisons the rest of the run: the agent retries, gets the same stale error, and concludes the world is broken. This is the single most common caching bug in CrewAI crews. - **Refuse to cache volatile data.** A price quote, a queue depth, a "current on-call" lookup — the second call in a long run is asking precisely because time has passed. - **Cache only the expensive successes.** A search tool that costs an API credit per call and returns stable results for the same query is the ideal cache candidate. ## Non-idempotent tools are the sharp edge The cache is keyed on arguments, and it does not know whether your tool reads or writes. A tool that posts a Slack message, creates a Jira ticket, or charges a card will, on a second identical call, return the cached observation and *not* perform the side effect — or, depending on how the agent phrases the arguments, perform it twice when you assumed once. Neither is a caching bug so much as a design mismatch: side-effecting tools should either disable caching (`cache_function` returning False) or be made genuinely idempotent by carrying an idempotency key in their arguments, so a repeat with the same key is safe whether it hits the cache or your backend. ## Observability A cached call produces no outbound request and no log line inside `_run`, which is exactly why caching confuses people debugging a crew: the tool "didn't run" and the metrics do not move. When you are diagnosing why an agent keeps seeing an outdated observation, turning `cache` off is a cheap experiment that isolates the cache from prompt or tool bugs. Conversely, when you are measuring how many external calls a crew really makes, remember the cache sits between the agent's decisions and your instrumentation. ## Interaction with the rest of the loop Caching does not reduce token cost. The cached observation is still appended to the agent's context and still consumed on the next model call — you save the API call and the latency of your tool, not the prompt. If context growth from repeated large observations is the problem, the fix is a smaller, summarised tool return, not caching. ## What to say in an interview State the default (`cache=True`), the key (tool plus arguments), the hook (`cache_function(args, result) -> bool`), and then move immediately to judgment: what you refuse to cache and why. The interesting part is not that the cache exists, it is that a default-on cache silently changes the behaviour of failing and side-effecting tools, and that a two-line predicate is how you take that back.

  • Why is caching a tool's error string particularly damaging inside an agent loop?
    Because the agent's natural recovery move is to retry the same call. With the failure cached, the retry never reaches your code — it replays the stored error, the agent concludes the capability is unavailable, and it either gives up or burns its remaining iterations. A `cache_function` that returns False for error and empty results keeps the retry path real.
  • Does CrewAI's tool cache reduce token spend?
    No. It skips the tool execution and its latency, but the cached observation is still inserted into the agent's context and still billed on the next model call. If repeated large observations are inflating your prompt, the fix is to return less from the tool — trim, summarise, or paginate — rather than to rely on caching.
  • How would you make a ticket-creating tool safe under caching?
    Either turn caching off for it with a `cache_function` that always returns False, or make it genuinely idempotent: take an explicit idempotency key as an argument and have the backend treat a repeat with the same key as a no-op returning the original result. The second option is more robust, because it also survives retries that never touch the cache.

saying these in an interview costs you the question

  • Assumes tool caching is opt-in and off by default
  • Thinks the cache persists across separate crew runs or processes
  • Believes caching lowers token cost as well as call count
  • Caches error and empty results, then wonders why retries fail
  • Gives a side-effecting write tool the default cache behaviour

context