skip to content

In Prefect, how does calling a task with .submit() differ from calling it directly?

level: middleimportance: must knowfreq 68%

answer

  1. One returns a value, one a handle
  2. Who blocks, and when
  3. The handle is redeemed later
  4. Passing the handle onward creates an edge
  5. Concurrency comes from the runner, not the call

basics

~20 s

Calling a Prefect task directly runs it inline in the flow and returns its value, so steps are sequential. Calling .submit() hands it to the flow's task runner and returns a PrefectFuture immediately, letting independent task runs overlap; .result() waits for the value.

solid answer

~40 s

A direct call — `rows = fetch(region)` — executes the task run right there in the flow's own thread and returns the real value, so the flow is sequential. `future = fetch.submit(region)` schedules the task run on the flow's **task runner** and returns a `PrefectFuture` straight away, so the flow keeps going and independent work overlaps. You resolve a future with `future.result()`, which blocks until the run finishes and re-raises its exception on failure, or `future.wait()` if you only need completion. Dependencies are inferred from data: passing a future into another `.submit()` call makes Prefect wait for it and pass the resolved value, and `wait_for=[other_future]` expresses an ordering edge with no data flowing. Prefect waits for outstanding submitted runs before the flow run finishes, so you cannot accidentally abandon them.

code

python · 25 lines
python
from prefect import flow, task


@task
def fetch(region: str) -> int:
    return len(region)


@task
def total(counts: list[int]) -> int:
    return sum(counts)


@flow
def sequential(regions: list[str]) -> int:
    # each call blocks until it returns a real value
    counts = [fetch(r) for r in regions]
    return total(counts)


@flow
def concurrent(regions: list[str]) -> int:
    # every fetch starts before the first result is needed
    futures = [fetch.submit(r) for r in regions]
    return total([f.result() for f in futures])

go deeper

for a junior

Know that a direct call gives you the value and blocks, while .submit() gives you a future you later resolve with .result(). Recognising both call styles in code is enough at this level.

for a middle

Explain the mechanics: what a PrefectFuture is, that submission delegates to the task runner, that exceptions surface at resolution, and how passing a future creates a data dependency.

for a senior

Demonstrate the operational reading — spot the submit-then-immediately-resolve anti-pattern, use wait_for for ordering without data, and reason about which failures a late .result() hides.

for a principal

Own the concurrency model as a design choice: how much fan-out a flow process should hold, when threads stop paying, and when the work belongs on a distributed runner or split into separate flows entirely.

## Two calling conventions, one function A task-decorated function can be invoked two ways inside a flow, and the difference is where the work runs and what you get back. **Direct call.** `rows = fetch("eu")` creates a task run and executes it inline, in the flow's own thread, returning the actual return value. The flow blocks until it finishes. This is the simplest form and is exactly right for a linear pipeline where each step needs the previous step's output anyway. **Submitted call.** `fut = fetch.submit("eu")` creates the task run and hands it to the flow's **task runner**, returning a `PrefectFuture` immediately. The flow continues to the next statement while the work proceeds. This is how you get concurrency in Prefect. ```python futures = [fetch.submit(r) for r in regions] rows = [f.result() for f in futures] ``` All the fetches start before the first `.result()` blocks, so the wall-clock cost is roughly the slowest fetch rather than the sum. ## What a future is and how you resolve it A `PrefectFuture` is a handle to a task run that may still be in flight. The operations you need: - `future.result()` — blocks until the run finishes and returns its value. If the run failed, the exception is re-raised here, which is what makes failure propagate to the flow. Pass `raise_on_failure=False` to get the failed state back instead of an exception when you want to handle it yourself. - `future.wait()` — blocks until the run reaches a final state without returning the value. - `future.state` — the run's state object, useful for inspecting rather than raising. The timing detail that catches people out: exceptions surface when you resolve, not when you submit. A loop that submits fifty tasks and only later collects results will not raise on the fifth failure until you reach its `.result()`. ## How dependencies are expressed There is no explicit edge syntax. Two mechanisms cover everything: **Data dependencies.** Pass a future straight into another task: ```python raw = extract.submit(region) clean = transform.submit(raw) # waits for raw, receives its value ``` Prefect resolves the future before running `transform`, so the second task run cannot start until the first completes and it receives the plain value, not the future. **Ordering-only dependencies.** When the second task does not consume the first's output — it reads a table the first wrote, say — use the reserved keyword: ```python refresh.submit(wait_for=[load_future]) ``` That is the equivalent of an edge that carries no payload, and it is the correct fix for the classic bug where two submitted tasks touching the same table race because nothing linked them. ## Where the concurrency actually comes from `.submit()` on its own does not create threads — it delegates to the flow's task runner. The default runner in Prefect 3 executes submitted runs as threads inside the flow process, which is genuinely concurrent for I/O-bound work (HTTP calls, warehouse queries, object-storage reads) and largely pointless for CPU-bound pure-Python work because of the interpreter lock. Swapping in a distributed runner changes where submitted runs execute without changing any of the `.submit()` calls. Directly-called tasks bypass the runner entirely, no matter which runner is configured. ## Flow completion and outstanding futures Prefect waits for submitted task runs to reach a final state before the flow run itself is finished, so a fire-and-forget submit is not silently dropped. But "not dropped" is not the same as "handled": if you never call `.result()`, the exception never propagates through your code, and your flow logic proceeds as though the work succeeded. When correctness depends on the result, resolve the future. ## Async flows If the flow and tasks are `async def`, awaiting a task call runs it concurrently under asyncio in the usual way, and the same `.submit()` mechanics remain available. Mixing sync and async carelessly is a common source of confusion — keep a flow consistently one or the other. ## Choosing between them Use a direct call when the next line needs the value anyway; the code is simpler and the traceback is more direct. Use `.submit()` when several units of work are independent and slow — the fan-out over regions, files or partitions. A pipeline where every step feeds the next gains nothing from submitting, and sprinkling `.submit()` plus an immediate `.result()` on the next line is a common anti-pattern: it reintroduces the blocking you were trying to remove while adding a layer of indirection. ## The interview shape Say the three things: direct call returns a value and blocks; `.submit()` returns a `PrefectFuture` and delegates to the task runner; futures passed as arguments become dependencies. Then add the operational nuance — failures surface at resolution time, and `wait_for` covers ordering without data.

  • Two submitted tasks write to the same table and must not overlap, but neither consumes the other's output. How do you order them?
    Pass `wait_for=[first_future]` to the second `.submit()` call. That creates an ordering edge with no data flowing, so the second task run starts only after the first reaches a final state. Without it, the two runs are independent and the task runner is free to execute them at the same time.
  • A flow submits fifty tasks and one fails, but the flow log shows it continuing for another minute. Why?
    Submission does not raise. The failure is recorded on that task run, but the exception only reaches your code when you resolve the future with `.result()`, so a loop that submits everything first and collects afterwards keeps going until it reaches the failed one. Resolve earlier, or inspect states, if you need to fail fast.
  • Does calling .submit() guarantee the work runs on another machine?
    No. `.submit()` only hands the run to the flow's configured task runner. The default executes submitted runs as threads inside the same process; distributing across machines requires a distributed task runner and its cluster. The call site is identical either way, which is the point of the abstraction.

A direct call is ordering at the counter and standing there until the coffee is handed over. .submit() is taking a numbered ticket — you are free to order three more drinks, and the ticket is redeemed for the actual cup when you call .result().

saying these in an interview costs you the question

  • Believing .submit() alone spawns processes or remote workers
  • Calling .result() on the next line and expecting parallelism
  • Expecting a failed submitted task to raise at submission time
  • Assuming two submitted tasks touching one table are ordered automatically
  • Confusing a future with the value it will eventually hold

context