In Airflow, which long-running processes make up a deployment and what does each do?
answer
- several processes, one shared database
- one decides, another executes
- the UI process does not schedule
- something waits on behalf of deferred tasks
- scheduler, workers, api/web, triggerer
basics
~20 sAn Airflow deployment runs a scheduler that decides what should run, workers that execute task code, a web UI/API process, a triggerer for deferred tasks, and a metadata database that all of them read and write.
solid answer
~40 sAirflow is several cooperating processes around one relational **metadata database**. The **scheduler** parses DAGs (or reads them from a separate **DAG processor**), works out which task instances are ready, and hands them to the configured **executor**. **Workers** actually run task code — with `LocalExecutor` they are subprocesses of the scheduler, with `CeleryExecutor` they are separate `airflow celery worker` processes, with `KubernetesExecutor` they are one pod per task instance. The **webserver** (called the **api-server** in Airflow 3) serves the UI and REST API; it never schedules anything. The **triggerer** runs the async waits of deferrable operators so a deferred task holds no worker slot. Every process's view of state comes from the metadata database, which is why it is the one component you must back up and size properly.
code
bash · 7 linesairflow scheduler
airflow triggerer
airflow celery worker
# Airflow 2:
airflow webserver
# Airflow 3 equivalent:
# airflow api-servergo deeper
Be ready to name the processes — scheduler, worker, web/API server, triggerer — and the metadata database, and say in one line what each is for.
Explain how they coordinate: everything goes through rows in the metadata database, and the executor lives inside the scheduler rather than being its own service.
Show you can reason about partial failure — which pipelines keep running when the scheduler, the web process or the triggerer is down, and what you monitor for each.
Own the deployment topology: database sizing and backup, remote logging, isolating DAG parsing from scheduling, and how many of each process a given workload needs.
## Airflow is not one program A running Airflow installation is a small set of long-lived processes plus one relational database. Understanding which process does what is the prerequisite for almost every operational question about Airflow, because failures present very differently depending on which piece is down. ## The metadata database The metadata database (PostgreSQL or MySQL in production; SQLite only for a laptop) holds DAG runs, task instances and their states, connections, variables, pools, XComs and the serialized DAG structure the UI renders. It is shared state: no Airflow process talks to another by RPC in Airflow 2 — they coordinate through rows in this database. That makes the database the single most important thing to size, monitor and back up. ## The scheduler The scheduler is the brain. In a loop it creates DAG runs for schedules that are due, evaluates each task instance's upstream dependencies, trigger rule, pool availability and concurrency limits, and moves eligible task instances to the `scheduled` and then `queued` state. It also detects **zombie** tasks — task instances whose worker stopped heartbeating — and marks them for retry. The **executor** is not a separate service: it is a pluggable component that lives *inside* the scheduler process and decides how a queued task instance is handed off to something that will run it. ## The workers Where task code executes depends entirely on the executor: - `LocalExecutor` forks subprocesses on the scheduler's own machine. - `CeleryExecutor` publishes a message to a broker (Redis or RabbitMQ); separate `airflow celery worker` processes on other machines pick it up. - `KubernetesExecutor` asks the Kubernetes API to create one worker pod per task instance, which runs `airflow tasks run` and then exits. Whatever the shape, the worker runs the operator's code and writes the resulting state back to the metadata database (in Airflow 3, via the Task Execution API rather than a direct database connection). ## The webserver / api-server The web process serves the UI and the REST API. It is a *reader and requester*: it renders DAG graphs from serialized DAGs, shows logs, and lets you clear or manually trigger tasks by writing rows the scheduler will act on. It never schedules or runs anything itself, so a dead webserver means you are blind, not stopped. Airflow 3 renames this component to the **api-server**, with the UI served as a client of that API. ## The triggerer The triggerer, introduced in Airflow 2.2, runs an asyncio event loop that hosts *triggers* — the small awaitable objects that deferrable operators hand off to. When a deferrable sensor or operator defers, its worker slot is released and its trigger waits inside the triggerer; when the trigger fires, the scheduler re-queues the task to resume. One triggerer can host thousands of waits, which is why long polls should be deferrable rather than occupying workers. ## The DAG processor Parsing DAG files is expensive and runs arbitrary Python. In Airflow 2 the scheduler does it in subprocesses by default, but you can run a **standalone DAG processor** (`airflow dag-processor`) so parsing is isolated from scheduling; Airflow 3 makes that separation mandatory, which also means DAG-authoring code no longer executes inside the scheduler. ## What breaks when each is down - Scheduler down: running tasks continue to completion, but nothing new is scheduled or queued. Nothing is lost — it resumes from the database. - Workers down: task instances pile up in `queued`. - Webserver down: you lose visibility; pipelines keep running. - Triggerer down: deferred tasks stay deferred and never resume. - Metadata database down: everything stops. ## Optional extras Deployments commonly add a broker (Redis/RabbitMQ) and Celery Flower for `CeleryExecutor`, a log store (S3, GCS) since worker-local logs vanish with the pod or container, and a StatsD/OpenTelemetry endpoint for metrics.
- What happens to tasks that are already running if you restart the scheduler?They keep running. The worker owns the running process and writes the final state to the metadata database, so a scheduler restart does not kill anything. What stops is new work: no new DAG runs are created and no eligible task instances move to queued until the scheduler is back. On restart it re-reads state from the database and, depending on version and executor, adopts or re-queues task instances it had already dispatched.
- Why would you run a standalone DAG processor instead of letting the scheduler parse DAG files?Parsing runs arbitrary user Python, so isolating it keeps a heavy or malicious DAG file from starving or crashing the scheduling loop, and lets the scheduler run without access to DAG code or its dependencies. It is optional in Airflow 2 (`airflow dag-processor`) and mandatory in Airflow 3.
- Where do task logs live, and why does that matter for a containerised deployment?By default a worker writes logs to its own local filesystem and the web process fetches them from the worker over HTTP. That breaks when the worker is a pod that has already exited, so production deployments configure remote logging to object storage (S3, GCS, Azure Blob) so logs outlive the process that produced them.
saying these in an interview costs you the question
- Says the webserver schedules or triggers DAG runs
- Thinks the executor is a separate service you deploy
- Cannot name where task state is actually stored
- Believes stopping the scheduler kills running tasks
- Assumes worker-local logs survive a finished pod