skip to content

Prefect

A Python-native orchestrator where flows and tasks are decorated functions and the graph emerges from actual execution. Interviewers use it to contrast dynamic orchestration with Airflow's static, scheduling-first DAG model.

on this pageshow

explore

questions

18

In Prefect, what is a deployment and how does it differ from running a flow locally?

level: juniorimportance: must knowfreq 78%

answer

  1. the flow itself is just Python
  2. something has to make it remotely triggerable
  3. a record on the server, not the code
  4. entrypoint, schedule, work pool, parameters
  5. created by prefect deploy or flow.deploy()

basics

~20 s

A Prefect deployment is a server-side record of a flow — entrypoint, code location, default parameters, schedules, work pool and infrastructure overrides — so the API can schedule and trigger runs on remote infrastructure instead of you calling the function yourself.

solid answer

~40 s

Calling a `@flow`-decorated function runs it right there in your Python process: Prefect tracks the run and its states, but nothing schedules it and nothing runs it when your laptop is closed. A **deployment** is the object that makes a flow remotely invokable. It stores the entrypoint (`flows/etl.py:etl_flow`), where the code lives, default parameters, tags, a version, zero or more schedules, and the **work pool** whose worker should execute it. You create one with `prefect deploy` (driven by `prefect.yaml`) or programmatically with `flow.from_source(...).deploy(...)`. After that the API can create flow runs — on a cron/interval/rrule schedule, from the UI, from `prefect deployment run`, from an automation, or from `run_deployment()` inside another flow — and a worker picks each run up and executes it on the configured infrastructure.

code

yaml · 19 lines
yaml
# prefect.yaml
name: pipelines
prefect-version: 3.0.0

pull:
  - prefect.deployments.steps.git_clone:
      repository: https://github.com/acme/pipelines.git
      branch: main

deployments:
  - name: nightly
    entrypoint: flows/etl.py:etl_flow
    parameters:
      day: "yesterday"
    schedules:
      - cron: "0 4 * * *"
        timezone: "UTC"
    work_pool:
      name: k8s-prod

go deeper

for a junior

Be ready to state plainly that a deployment is a stored, schedulable definition of a flow and to name what it holds: entrypoint, parameters, schedule and work pool. Know that prefect deploy creates it.

for a middle

Explain the two creation paths (prefect.yaml plus prefect deploy, versus flow.from_source(...).deploy()), what flow.serve() does differently, and how a run travels from a schedule to a worker.

for a senior

Show how deployments fit CI: version pinning, promoting the same flow to dev and prod deployments with different job variables, and why deployment creation is safe to run on every merge.

for a principal

Own the convention across teams — naming, ownership, how many deployments per flow, whether code ships by git pull or baked image, and how deployment definitions stay reviewable in version control.

## The problem a deployment solves In Prefect you write ordinary Python and decorate it: ```python from prefect import flow, task @task def extract(day: str) -> list[dict]: ... @flow def etl_flow(day: str = "today"): return extract(day) if __name__ == "__main__": etl_flow("2026-08-20") ``` Running `python flows/etl.py` executes the flow immediately in that process. Prefect still observes it — a flow run appears in the UI with states, task runs and logs — but the *trigger* was you, and the *infrastructure* was whatever machine you were sitting at. Nothing here is scheduled, nothing is reproducible on a server, and there is no way for a teammate or an event to start it. A **deployment** closes that gap. It is a record stored by the Prefect API that says: this flow, found at this entrypoint, fetched from this location, with these default parameters, on these schedules, should run in this work pool with these infrastructure settings. ## What a deployment actually stores - **Entrypoint** — `path/to/file.py:flow_function_name`. Prefect imports the module and calls that flow object. - **Code source** — how the executing process obtains the code: pull steps in `prefect.yaml` (for example `prefect.deployments.steps.git_clone`), a `flow.from_source(...)` reference to a git URL or storage bucket, or code baked into a Docker image. - **Parameters** — defaults for the flow's arguments, overridable per run. Prefect validates them against the function's type annotations (pydantic under the hood), so a bad parameter fails fast at run creation. - **Schedules** — a list; each may be `cron`, `interval` (seconds, with an optional anchor date), or `rrule`, each with a timezone and an active flag. - **Work pool (and optional work queue)** — the queue of pending runs a worker polls. - **Job variables** — per-deployment overrides of the work pool's infrastructure defaults (image, env vars, CPU/memory, namespace). - **Metadata** — name, version, description, tags, concurrency-related settings. A deployment is identified as `flow-name/deployment-name`, which is what you pass to `prefect deployment run`. ## How you create one The file-driven route is `prefect.yaml` at the project root plus `prefect deploy`: ```yaml deployments: - name: nightly entrypoint: flows/etl.py:etl_flow parameters: day: "yesterday" schedules: - cron: "0 4 * * *" timezone: "UTC" work_pool: name: k8s-prod ``` The Python route is equivalent and often nicer in CI: ```python flow.from_source( source="https://github.com/acme/pipelines.git", entrypoint="flows/etl.py:etl_flow", ).deploy(name="nightly", work_pool_name="k8s-prod") ``` There is also `flow.serve(name="nightly", cron="0 4 * * *")`, which creates a deployment *and* runs a long-lived process that executes its runs as subprocesses. It is the fastest way to get something scheduled, but that process is your only execution capacity — no work pool, no remote infrastructure, no horizontal scaling. Teams graduate from `serve` to a work pool when they need runs on containers or Kubernetes. ## The mental model to say out loud Deploying does **not** upload your flow code to Prefect and does not "run" anything. It registers a pointer plus a configuration. Execution stays in your infrastructure: the API creates a flow run in `Scheduled` state, a worker polling the deployment's work pool claims it, materializes the configured infrastructure, fetches the code, and runs the flow there. If no worker is polling that pool, the run simply sits and eventually shows as `Late` — a deployment by itself is not execution capacity. This is also the piece with no direct Airflow equivalent. In Airflow, a DAG file placed in the DAGs folder is simultaneously the definition, the schedule and the thing the scheduler executes. In Prefect the flow is plain Python that runs anywhere, and the deployment is a separate, explicitly created object that adds remote triggering, scheduling and infrastructure to it. ## Version note This describes Prefect 3.x, where `prefect deploy` / `flow.deploy()` and workers with work pools are the model. Prefect 2 originally used `prefect deployment build`, agents, and infrastructure/storage blocks attached to the deployment; agents were removed in 3.x in favour of workers.

  • Does creating a deployment upload your flow code to Prefect?
    No. The deployment records *where* the code lives — a git repository via a pull step or `from_source`, an object-storage path, or a Docker image — plus the entrypoint inside it. At run time the worker executes those pull steps and imports the flow. Prefect stores orchestration metadata, states and logs, not your source or your data.
  • When would you use flow.serve() instead of a work pool and worker?
    For a single always-on machine and a handful of flows: `serve` creates the deployment and runs its runs as subprocesses of that one long-lived process. It is ideal for prototypes, internal tools and small teams. Move to a work pool when you need containerized or Kubernetes execution, per-deployment infrastructure overrides, or horizontal scaling across many workers.
  • How do you trigger one deployment from inside another Prefect flow?
    Call `run_deployment(name="flow-name/deployment-name", parameters={...})`. It creates a flow run for that deployment through the API, so the child executes on *its* work pool and infrastructure rather than in the caller's process. By default the caller waits for it to finish and it appears as a subflow; you can opt out of waiting.

The flow is a recipe you can cook yourself any time; the deployment is the standing order posted in the kitchen saying which recipe, with which ingredients, at what hour, on which stove.

saying these in an interview costs you the question

  • Says deploying uploads or copies the flow code to Prefect
  • Thinks a deployment executes the flow by itself, no worker needed
  • Confuses the deployment with the @flow decorator or the flow run
  • Believes you must deploy before you can run a flow at all
  • Calls the deployment 'Prefect's DAG file' dropped in a folder

context

open as a page

In Prefect, what do the @flow and @task decorators do to a Python function?

level: juniorimportance: must knowfreq 80%

basics

~20 s

@flow turns a function into an orchestrated flow run with tracked state, validated parameters and logging. @task marks a unit of work called inside it, giving each call its own tracked run with retries, caching and concurrency options.

open as a page

In Prefect, what does setting cache_key_fn on a task do?

level: juniorimportance: must knowfreq 55%

basics

~20 s

cache_key_fn computes a string key from the run context and the task's inputs. If a completed task run already recorded that key and it has not expired, Prefect skips the function body and returns the stored result in a Cached state.

open as a page

In Prefect, what is a work pool and what does a worker do with it?

level: middleimportance: must knowfreq 74%

basics

~20 s

A Prefect work pool is a typed queue of scheduled flow runs plus the default infrastructure template for running them. A worker process polls one pool, claims runs, and launches each on that infrastructure — a subprocess, container, or Kubernetes job.

open as a page

In Prefect, how does calling a task with .submit() differ from calling it directly?

level: middleimportance: must knowfreq 68%

basics

~20 s

Calling a Prefect task directly runs it inline in the flow and returns its value, so steps are sequential. Calling .submit() hands it to the flow's task runner and returns a PrefectFuture immediately, letting independent task runs overlap; .result() waits for the value.

open as a page

In Prefect, what distinguishes a Crashed state from a Failed state?

level: middleimportance: must knowfreq 60%

basics

~20 s

Failed means the run executed and its code raised or reported an error, so Prefect recorded the exception. Crashed means the run was cut off by its environment — process killed, container evicted, out of memory — so it never reported its own outcome.

open as a page

In Prefect, how does a worker get your flow code when a deployment run starts?

level: middleimportance: should knowfreq 52%

basics

~20 s

Prefect stores a pointer, not your code. At run time the worker executes the deployment's pull steps — typically a git clone or a storage download — or uses code already baked into the container image, then imports the flow named by the entrypoint.

open as a page

In Prefect, what does calling a task's .map() method do to the arguments you pass?

level: middleimportance: should knowfreq 55%

basics

~20 s

Prefect's .map() creates one task run per element of each iterable argument, zipped element-wise, and runs them through the task runner. Non-iterable values are broadcast to every child run; wrap an iterable you want passed whole in unmapped().

open as a page

When should you use a Prefect subflow instead of a task for a step in a pipeline?

level: middleimportance: should knowfreq 48%

basics

~20 s

Calling a Prefect flow from inside another flow creates a subflow run with its own flow run record, parameter validation, retries and task runner. Use a task for a single unit of work, a subflow to group and reuse a whole multi-task pipeline.

open as a page

A Prefect deployment's scheduled runs are piling up as Late — how do you diagnose it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Late means Prefect created the runs on schedule but nothing claimed them. Check that a healthy worker of the right type is polling that exact work pool and queue, that neither is paused, and that a pool, queue or global concurrency limit is not saturated.

open as a page

A Prefect task calling a flaky API fails intermittently. How do you configure retries and timeouts?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Set retries and retry_delay_seconds on the @task. retry_delay_seconds accepts a list of per-attempt delays or exponential_backoff(), with retry_jitter_factor to spread a thundering herd, and timeout_seconds bounds a hung call so it fails instead of blocking the flow forever.

open as a page

In Prefect, why might an on_failure hook on a flow never fire in production?

level: seniorimportance: should knowfreq 40%

basics

~20 s

State hooks are client-side code that runs in the process owning the run. If that process is killed the hook never executes, and on_failure only covers Failed states — a Crashed or Cancelled run needs on_crashed or on_cancellation instead.

open as a page

In Prefect, why can a persisted task result be unreadable to a worker on another machine?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Prefect stores a reference in the state and the payload in result storage, which defaults to the local filesystem of the machine that ran the task. A different worker, container or pod has no such path, so the reference resolves to nothing.

open as a page

How would you choose between Prefect Cloud and a self-hosted Prefect server for a team?

level: principalimportance: should knowfreq 34%

basics

~20 s

Both run the same orchestration API and the same hybrid execution model, so the decision is about who operates the control plane. Self-hosting means you run the API, database and UI; Cloud adds hosted availability, workspaces, access control and serverless push and managed work pools.

open as a page

In Prefect, how would you cap concurrent calls that many flows make to one shared API?

level: principalimportance: should knowfreq 30%

basics

~20 s

Put the cap in Prefect's concurrency system rather than in each flow: a task tag concurrency limit throttles every task run carrying that tag across the workspace, and a global concurrency limit with the concurrency or rate_limit context manager guards arbitrary code blocks.

open as a page

In Prefect, how do a deployment's job variables relate to the work pool's base job template?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

A Prefect work pool's base job template declares the infrastructure settings available to its runs and their defaults; a deployment's job_variables override those defaults for that deployment only, and an individual run can override them again at trigger time.

open as a page

In Prefect, what does create_markdown_artifact attach to a flow run?

level: middleimportance: nice to knowfreq 25%

basics

~20 s

It publishes a block of rendered markdown to the run in the Prefect UI — a row-count summary, a data-quality report, a diff. Give it a key and Prefect versions it, so you can page through the same artifact across every run.

open as a page

In Prefect, when would you replace the default ThreadPoolTaskRunner with a Dask or Ray runner?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

Prefect's default ThreadPoolTaskRunner executes submitted task runs as threads inside the flow process, which suits I/O-bound work. Swap in DaskTaskRunner or RayTaskRunner when task runs are CPU-heavy or memory-heavy and need to spread across a cluster.

open as a page