In Prefect, what does setting cache_key_fn on a task do?
answer
- Prefect asks a question before it works
- a string decides whether the body runs
- same inputs, same key, same answer
- the run lands in Cached, not Running
basics
~20 scache_key_fn computes a string key from the run context and the task's inputs. If a completed task run already recorded that key and it has not expired, Prefect skips the function body and returns the stored result in a Cached state.
solid answer
~50 s`cache_key_fn` is a function Prefect calls **before** running the task, with `(context, parameters)`; the string it returns becomes that task run's cache key. Prefect looks the key up in its backend: if a task run already completed with the same key and `cache_expiration` has not elapsed, the new run transitions straight to `Cached` — a state whose type is Completed — and the previously stored return value is handed to downstream tasks without executing your code. Returning `None` disables caching for that run. The shipped helper `task_input_hash` builds the key from the task's inputs. Two caveats matter: the key lives in the Prefect backend and is not automatically namespaced per task, so a careless key can collide across flows, and a hit can only return a value if that result was persisted somewhere the new run can read. Prefect 3 expresses the same idea declaratively with `cache_policy`.
code
python · 14 linesfrom datetime import timedelta
from prefect import flow, task
from prefect.tasks import task_input_hash
@task(cache_key_fn=task_input_hash, cache_expiration=timedelta(hours=1))
def load_day(day: str) -> int:
print(f"expensive work for {day}")
return len(day)
@flow
def main():
load_day("2026-08-20")
# same argument, same key -> Cached, body does not run
load_day("2026-08-20")go deeper
Be able to state that cache_key_fn returns a string, that an identical string within the expiry means the body is skipped, and that task_input_hash is the ready-made version keyed on inputs.
Explain the mechanics: the key is looked up in the Prefect backend before the run starts, a hit produces a Cached state whose type is Completed, and the returned value comes from persisted result storage rather than from memory.
Show the operational judgment — expiry choice for still-growing partitions, key collisions across flows, and the danger of caching side-effecting tasks so a hit silently skips the effect while reporting success.
Own the policy: where caching belongs in an expensive pipeline, what a cache entry costs in storage and confusion, and whether a team standard should be an explicit key function or Prefect 3's composable cache policies.
## The problem caching solves A Prefect task is a Python function decorated with `@task`. By default the body runs every time the task is called inside a flow. When the body is expensive — a wide warehouse query, a model fit, a metered API call — you want the second invocation with the same inputs to reuse the first one's answer instead of paying again. `cache_key_fn` is how you tell Prefect when two runs count as "the same work". ## The signature `cache_key_fn` is an ordinary function that Prefect calls *before* the task body, passing the run context and a dict of the task's resolved parameters: ```python def key_from_day(context, parameters): return f"load-orders-{parameters['day']}" ``` Whatever string it returns is the task run's **cache key**. Returning `None` means "do not cache this particular run" — the task simply executes. ## What Prefect does with the key Before transitioning the run to `Running`, Prefect asks its backend (a local server, a self-hosted server, or Prefect Cloud) whether a completed task run already carries that key. If one exists and the configured expiration has not passed, the new run is set to `Cached` instead. `Cached` is a *state name* whose *state type* is Completed, so downstream tasks treat it as success and receive the previous run's return value. Your function body never executes — which is the whole point, and also the whole hazard when the body has side effects. ## cache_expiration `cache_expiration` takes a `datetime.timedelta` and bounds how long an entry is considered valid: ```python @task(cache_key_fn=task_input_hash, cache_expiration=timedelta(hours=1)) def fetch(day: str): ... ``` Omit it and the entry has no expiry — the key is good for as long as the record survives in the backend. Set it when the upstream data can change underneath a key that would otherwise stay identical (for example a "today" partition that is still being appended to). ## A cache hit needs a reachable result A Prefect state does not carry your data; it carries a *reference* to a result. So a cache hit can only hand back a value if that value was written to result storage that the reading run can actually open. Locally this works by default because results land under the Prefect home directory on the same machine. Across containers or workers on different hosts it does not, unless you configure shared `result_storage`. In Prefect 3, declaring a cache policy turns result persistence on for you, precisely because caching without persistence is meaningless. ## Cache keys are global, not per-task The key is a string in a shared table. Nothing stops two different tasks from producing the same string, and if they do they will share results — a genuinely confusing bug where a task "returns" another task's output. Defend against it by folding something task-identifying into the key (the task name, a version constant), or by using the built-in policies that already include task identity. ## Prefect 3: cache policies Prefect 3 keeps `cache_key_fn` but adds composable `cache_policy` values — `INPUTS`, `TASK_SOURCE`, `RUN_ID`, `FLOW_PARAMETERS`, `NO_CACHE` — which you combine with `+`: ```python @task(cache_policy=INPUTS + TASK_SOURCE) def transform(rows): ... ``` The default policy is scoped to the flow run, so repeat calls inside one flow run are reused while a later flow run recomputes. `TASK_SOURCE` folds the function's source into the key, so editing the body invalidates old entries — a common expectation that a pure inputs hash does not satisfy. An explicit `cache_key_fn` takes precedence over the policy. ## Forcing a recompute Set `refresh_cache=True` on the task (or via `.with_options(refresh_cache=True)`) to make a run ignore any existing entry, execute, and overwrite the stored result under the same key. That is the tool for "the upstream was wrong, recompute this one". ## When not to cache Do not cache tasks whose value is the side effect — sending mail, publishing a message, writing a file with a run-specific name — because a hit silently skips the effect while reporting success. Do not cache non-deterministic reads unless the key encodes the thing that varies. And be wary of caching very large return values: every hit and miss is an object round-tripped through your result storage, which can cost more than recomputing something cheap.
- How do you force a single run to ignore an existing cache entry?Set `refresh_cache=True` on the task, or call it as `my_task.with_options(refresh_cache=True)`. That run executes the body regardless of an existing entry and overwrites the stored result under the same key, so later runs pick up the corrected value. Changing an input that feeds the key, or editing the body under a source-aware policy, has the same effect indirectly.
- Why can two unrelated tasks accidentally share a cached result?Cache keys are strings in a shared backend table with no automatic per-task namespace. If two key functions can emit the same string — say both return the date partition — the second task hits the first task's entry and returns its value. Fold task identity into the key, or use a policy that includes task source, to keep the namespaces apart.
- What does returning None from cache_key_fn do?It disables caching for that specific run: Prefect skips the lookup, executes the task body, and stores no entry. It is the clean way to make caching conditional — for example never caching dry runs, backfill runs, or calls whose inputs include a value you know is non-deterministic.
It is a coat check: the key function writes the ticket, and if a matching ticket is already on the rack you get the coat back instead of buying a new one.
saying these in an interview costs you the question
- Thinks the cache lives in worker memory rather than the backend
- Assumes a cache hit works without any persisted result
- Believes cache keys are automatically unique per task
- Confuses a cache hit with a retry, which re-executes the body
- Caches a task whose only purpose is a side effect