skip to content

When should you use a Prefect subflow instead of a task for a step in a pipeline?

level: middleimportance: should knowfreq 48%

answer

  1. Nesting the same decorator, not a new one
  2. One is a step, one is a stage
  3. The inner call gets its own run record
  4. Retries and runner are configured per level
  5. Grouping, not a way to go parallel

basics

~20 s

Calling a Prefect flow from inside another flow creates a subflow run with its own flow run record, parameter validation, retries and task runner. Use a task for a single unit of work, a subflow to group and reuse a whole multi-task pipeline.

solid answer

~50 s

In Prefect a flow can call another flow. That inner call becomes a **subflow run**: a first-class flow run, linked to its parent, with its own parameters validated against type hints, its own `retries` and `timeout_seconds`, and its own `task_runner`. A task is the wrong unit for that — a task is one step, and it cannot contain other orchestrated steps with independent retry behaviour. So the rule of thumb is granularity: use `@task` for a single unit of work you want retried and cached, and a subflow when a step is itself a small pipeline you want to run standalone, reuse across parents, retry as a whole, or give a different concurrency model. The cost is overhead and an extra layer in the UI, so do not wrap every three tasks in a subflow. By default the call blocks until the subflow run finishes.

code

python · 29 lines
python
from prefect import flow, task
from prefect.task_runners import ThreadPoolTaskRunner


@task(retries=2)
def extract(region: str) -> list[dict]:
    return [{"region": region}]


@task
def load(rows: list[dict]) -> int:
    return len(rows)


@flow(retries=1, timeout_seconds=900,
      task_runner=ThreadPoolTaskRunner(max_workers=4))
def load_region(region: str) -> int:
    return load(extract(region))


@flow
def nightly(regions: list[str]) -> int:
    total = 0
    for region in regions:
        try:
            total += load_region(region)   # blocking subflow run
        except Exception:
            continue                       # one region may fail
    return total

go deeper

for a junior

Know that calling a flow-decorated function from inside another flow simply works and produces a nested run, and that a task is for one unit of work.

for a middle

Explain what the subflow run actually gets — its own record, parameter validation, retries, timeout and task runner — and why a task cannot provide any of that composition.

for a senior

Show restraint and diagnosis: subflows cost overhead and nesting, they block by default, and reaching for them to get parallelism is a modelling error you should be able to name.

for a principal

Own the decomposition strategy — which stages become independently deployable flows with stable parameter contracts owned by other teams, and where the boundary between one pipeline and several belongs.

## What a subflow is There is no special decorator. Any flow called from inside a running flow becomes a **subflow run**: ```python @flow def load_region(region: str) -> int: raw = extract(region) return load(transform(raw)) @flow def nightly(regions: list[str]) -> int: return sum(load_region(r) for r in regions) ``` Each `load_region(...)` call produces its own flow run, recorded separately, linked to the parent run so the UI shows the nesting. By default the call blocks: the parent waits for the subflow run to reach a final state before continuing, and the returned value is the subflow's return value. ## What a subflow gives you that a task does not **Its own composition.** A subflow contains tasks — and further subflows. A task cannot orchestrate steps below it; it is a leaf. If the step you are modelling is really "extract, transform and load one region", a subflow expresses that and a task flattens it into an opaque blob. **Its own run record.** The subflow appears as a flow run with its own duration, state and logs. When a nightly job covers twelve regions, twelve subflow runs make it obvious which region was slow or failed, and the parent stays readable. **Its own configuration.** Retries, timeout and task runner are set per flow. A subflow can retry as a unit while its parent does not, or use a different task runner from the parent — for example a distributed runner for one heavy stage and threads for the rest. Retrying a subflow re-executes its tasks, which is exactly why the subflow's steps must be idempotent. **Its own parameters.** Subflow arguments go through the same type-hint coercion and validation as any flow, so a bad value fails the subflow run before its tasks execute. **Standalone runnability.** The same function can be called directly on its own for a single region, or deployed independently, without any of the parent's machinery. ## What it costs A subflow is heavier than a task: an extra run record, extra state transitions, more API traffic, another layer of nesting in the UI. Wrapping every couple of tasks in a subflow makes a run graph that is technically correct and miserable to read. Prefer flat: tasks in a flow, subflows only where the grouping means something operationally. A second, subtler cost: because a blocking subflow call is a Python call, the parent's flow process is occupied while it runs. Composing very many subflows sequentially in a single parent makes one long-lived process holding all the state. ## Choosing, in practice Use a **task** when the step is one unit of work — one query, one API call, one file written. Tasks are the retry, cache and concurrency boundary, and they are cheap. Use a **subflow** when at least one of these is true: - The step contains several tasks that only make sense together, and you want them retried, timed out or observed as a unit. - You want the same pipeline reused from multiple parents or run standalone. - The step needs a different task runner or a different concurrency profile from the rest. - The step is owned by a different team and evolves separately, so a stable flow signature is a useful contract. Use **neither** — a plain function — when the code is a pure helper with nothing to observe or retry. ## A common mis-modelling People reach for subflows to get parallelism, having noticed that mapping tasks is easy but calling a flow per element is not. Fanning out work over elements is a task-mapping job; a blocking subflow per element runs them one after another unless you deliberately arrange otherwise. If the driver is concurrency, the right first move is mapping tasks or a wider task runner, not nesting flows. ## Failure propagation A failed subflow run raises into the parent at the call site, just as a failed directly-called task does, so the parent fails unless you catch it. If you want the parent to survive a failed region, wrap the call in `try`/`except` and decide explicitly — the same choice you would make with any Python call. This is one of the pleasant consequences of Prefect's model: control flow around failures is just Python. ## Answering the question Lead with the mechanic (calling a flow from a flow creates a linked subflow run with its own parameters, retries and task runner), then the granularity rule (task = one unit of work; subflow = a reusable multi-task stage), then the caution (overhead and nesting; do not use it to get parallelism). That progression shows you know both what it is and when not to reach for it.

  • If a subflow run fails, what happens to the parent flow run?
    The failure raises at the call site in the parent, so the parent run fails unless you catch it. Because the call is ordinary Python, wrapping it in try/except lets you continue with the remaining work — useful when one region out of twelve is allowed to fail without sinking the nightly job.
  • You want twelve regions processed concurrently. Is a subflow per region the right tool?
    Usually not. A subflow call blocks by default, so twelve of them run one after another. Concurrency in Prefect comes from mapping or submitting tasks onto the task runner, so fan out at the task level; reach for subflows when the grouping and separate configuration are what you actually want.
  • Can a subflow use a different task runner from its parent?
    Yes. The task runner is configured per flow, so a subflow handling a heavy stage can use a distributed runner while the parent keeps the default threads. That is one of the strongest reasons to make a stage a subflow rather than a cluster of tasks in the parent.

A task is a paragraph and a subflow is a chapter: the chapter can be read on its own, has its own structure inside, and can be rewritten as a unit — but a book made entirely of one-paragraph chapters is harder to read, not easier.

saying these in an interview costs you the question

  • Thinking a subflow needs its own special decorator
  • Using subflows to obtain parallelism instead of mapping tasks
  • Wrapping every two or three tasks in a subflow
  • Assuming a failed subflow is silently skipped by the parent
  • Believing a task can contain other orchestrated tasks

context