skip to content

In Prefect, what is a work pool and what does a worker do with it?

level: middleimportance: must knowfreq 74%

answer

  1. one side is config, the other is a process
  2. the queue also carries infrastructure defaults
  3. who calls whom, and in which direction
  4. type must match on both sides
  5. kubernetes, docker, process — and push

basics

~20 s

A Prefect work pool is a typed queue of scheduled flow runs plus the default infrastructure template for running them. A worker process polls one pool, claims runs, and launches each on that infrastructure — a subprocess, container, or Kubernetes job.

solid answer

~50 s

A **work pool** is the bridge between the Prefect API and your compute. It has a **type** — `process`, `docker`, `kubernetes`, or a cloud/serverless variant — and a **base job template**: the JSON schema of infrastructure settings (image, env, CPU, namespace) with defaults that every run in the pool inherits. Deployments are assigned to a pool, and their scheduled runs land in it. A **worker** is a long-lived process you run in your own infrastructure (`prefect worker start --pool k8s-prod`). Its type must match the pool's type. It polls the API *outbound* for runs in its pool, and for each one it merges the base job template with the deployment's job variables and submits the resulting job — a subprocess, a Docker container, a Kubernetes Job — then reports state back. Pools and their work queues carry concurrency limits and priorities. **Push** pools are the exception: Prefect Cloud submits runs directly to serverless infrastructure with no worker at all.

code

bash · 8 lines
bash
# create a pool that runs each flow run as a Kubernetes Job
prefect work-pool create k8s-prod --type kubernetes

# start a worker inside the cluster that drains that pool
prefect worker start --pool k8s-prod

# cap how many flow runs from the pool execute at once
prefect work-pool set-concurrency-limit k8s-prod 20

go deeper

for a junior

Know that a work pool is where a deployment's scheduled runs wait and that a worker process you start elsewhere picks them up. Remember the command shape prefect worker start --pool <name>.

for a middle

Explain the base job template, that worker type must match pool type, and the direction of the connection: the worker polls outbound, the API never dials in.

for a senior

Demonstrate operating them — sizing pools per environment and resource class, concurrency limits and queue priorities, worker high availability, and what a worker outage actually does to the backlog.

for a principal

Own the topology and its blast radius: how many pools, who may create them, whether serverless push pools are worth handing cloud credentials to Prefect, and how pool boundaries enforce isolation between teams.

## Two objects, two jobs Prefect splits remote execution into a piece the API owns and a piece you run. The **work pool** is server-side configuration. It has a name, a **type**, and a **base job template**. The type declares what kind of infrastructure runs live on: `process` (a subprocess on the worker's host), `docker` (a container), `kubernetes` (a Job in a cluster), plus cloud variants such as ECS, Cloud Run and Azure Container Instances. The base job template is a JSON-schema document exposing variables — image, environment, CPU/memory, service account, namespace, and so on — with defaults for the whole pool. The pool is also a queue: every scheduled run of every deployment assigned to it waits there. The **worker** is a process you start in the environment where work should execute: ```bash prefect work-pool create k8s-prod --type kubernetes prefect worker start --pool k8s-prod ``` Its loop is simple: poll the API for runs in its pool that are ready, claim one, build the infrastructure configuration, submit it, monitor it, and report state and logs back. The worker itself does not execute your flow code in-process (a `process` worker forks a subprocess; a Kubernetes worker creates a Job and watches it). Worker type must match pool type, and non-process worker types ship in integration libraries (`prefect-kubernetes`, `prefect-docker`, `prefect-aws`). ## Why polling matters The worker opens an **outbound** connection to the Prefect API; the API never connects into your network. That is the hybrid model: you open no inbound firewall hole, your flow code and the data it touches never leave your infrastructure, and Prefect sees orchestration metadata — run states, logs, parameters, timings. It also means workers are cattle: start more of them against the same pool for more throughput, and a worker crash strands nothing permanently, because unclaimed runs stay in the pool. It also means the pool is **not** execution capacity. A pool with no worker is a queue nobody is draining; its runs go `Scheduled` then `Late`. ## Work queues, concurrency and priority Every pool has at least a default **work queue**, and you can add more. Queues within a pool have a **priority** (lower number is served first) and their own concurrency limits, which is how you keep a flood of low-value backfill runs from starving a latency-sensitive deployment sharing the same compute. Limits stack: ```bash prefect work-pool set-concurrency-limit k8s-prod 20 ``` Caps concurrent flow runs from that pool; a queue limit caps a slice of it. A worker can be told to serve only specific queues. When a limit is saturated, further runs simply wait — and show as `Late` if they wait past their scheduled time, which is a benign cause of lateness people often misdiagnose. ## Choosing a pool type - **process** — simplest; runs on the worker host. Great for a single VM, dev, or light workloads. All runs share that host's dependencies and resources. - **docker** — each run gets a container from an image, so dependencies are isolated per deployment. Needs a Docker daemon on the worker host. - **kubernetes** — each run is a Job; you get cluster scheduling, per-run CPU/memory requests, node selectors and real isolation. The heaviest to operate. - **push (serverless)** — Prefect Cloud submits runs straight to Cloud Run, ECS, ACI or Modal using credentials you store. **No worker to run or babysit**, at the cost of giving Cloud credentials into your account. - **managed** — Prefect Cloud runs the flow on Prefect-operated infrastructure; least setup, least control. A common production layout is several pools, not one: separate pools per environment (dev/staging/prod), and sometimes per resource class (a `kubernetes` pool with big memory limits for heavy jobs, a `process` pool for cheap ones). Pools are also a natural blast-radius boundary — pausing a pool stops all its deployments at once. ## Do not map this onto Airflow An Airflow **pool** is a slot-based concurrency limit on tasks within one Airflow installation; an Airflow **executor** is a global setting for how the scheduler runs tasks. A Prefect **work pool** is neither: it is a per-deployment routing target that carries both the queue *and* the infrastructure template, chosen per deployment rather than once per installation. That per-deployment choice — this flow on Kubernetes, that one in a subprocess, from the same Prefect instance — is the practical difference worth stating in an interview. ## Version note Work pools and workers are the Prefect 3.x model (introduced during 2.x). Prefect 2 originally used **agents** polling work queues with infrastructure blocks attached to deployments; agents were removed in 3.x.

  • Why does the worker poll the Prefect API rather than the API pushing runs to the worker?
    So execution can live behind your firewall. The worker makes outbound calls only, so no inbound ports or ingress are needed, and your flow code and data never leave your infrastructure — Prefect receives states, logs and metadata. It also makes workers trivially horizontally scalable and disposable: unclaimed runs just stay in the pool.
  • You need one deployment on Kubernetes with 16 GB and another as a cheap subprocess. How do you arrange the pools?
    Two pools, because the pool type fixes the infrastructure kind: a `kubernetes` pool whose base job template sets generous memory defaults, and a `process` pool on a small VM. Assign each deployment to the right pool, and use job variables to tune per-deployment settings within the Kubernetes pool rather than creating a pool per flow.
  • What happens to runs in a Prefect work pool when every worker for it is stopped?
    They stay `Scheduled` and are flagged `Late` once their start time passes; nothing is lost or failed. When a worker comes back it claims the backlog and runs them, which can cause a thundering herd — pool and queue concurrency limits, or pausing the deployment's schedule during a long outage, keep that under control.
  • How do work queues inside a Prefect work pool differ from the pool itself?
    The pool owns the infrastructure type and base job template; queues are subdivisions of its backlog. Each queue has a priority and can carry its own concurrency limit, so a high-priority queue is served first and a bulk queue can be capped. Workers can be started against specific queues to dedicate capacity.

The work pool is a numbered job board with the shop's standard tooling written at the top; the worker is the technician who checks that board, takes the next ticket, and sets up the machine described on it.

saying these in an interview costs you the question

  • Calls a Prefect work pool the same thing as an Airflow pool or executor
  • Thinks the Prefect API pushes runs to workers over an inbound connection
  • Believes a work pool executes runs on its own without a worker
  • Says one work pool per Prefect instance, chosen globally like an executor
  • Assumes flow code and data are sent to Prefect Cloud for execution

context