In Airflow 2, how can several schedulers run at once without duplicating task instances?
answer
- no leader, no ZooKeeper
- the database arbitrates
- two schedulers must not claim the same row
- it is a locking feature, not a partitioning one
basics
~20 sAirflow 2 supports multiple active schedulers with no leader election. They coordinate through the metadata database, taking row-level locks with SELECT ... FOR UPDATE so only one scheduler can claim and queue a given task instance.
solid answer
~50 sFrom Airflow 2.0 you can run several `airflow scheduler` processes **active-active**. There is no leader, no ZooKeeper, and no static partitioning of DAGs: every scheduler runs the full loop and they arbitrate through the metadata database, taking row-level locks (`SELECT ... FOR UPDATE`, with `SKIP LOCKED` so schedulers do not block each other) on the DAG-run and task-instance rows they are about to act on. Whoever wins the lock queues the task; the others move on. Because the mechanism is database locking, it requires a backend that supports it — PostgreSQL or MySQL 8 — and it does add write and lock pressure to the metadata database, so the database is usually the first thing to run out of headroom. The benefits are failover (losing one scheduler slows scheduling but does not stop it) and lower scheduling latency at high DAG counts. All schedulers must see the same DAG files, or a DAG a scheduler cannot parse will be inconsistently handled.
code
text · 6 linesscheduler A: SELECT ... FROM task_instance
WHERE state = 'scheduled'
FOR UPDATE SKIP LOCKED LIMIT 16;
scheduler B: identical query -> receives a disjoint set of rows,
never blocks on the rows A already holdsgo deeper
Just know that Airflow 2 allows more than one scheduler at a time, unlike Airflow 1.
Explain the mechanism — active-active with row-level database locks rather than leader election — and what database support it assumes.
Reason about the operational side: metadata database pressure, consistent DAG distribution to every scheduler, and what happens to in-flight work when one dies.
Set the topology policy: how many schedulers, how the database is sized and made highly available, and what evidence justifies scaling schedulers rather than workers.
## Before Airflow 2 In Airflow 1.10 the scheduler was a singleton. Running two was unsafe, so high availability meant an active/passive pair with an external supervisor, and scheduling latency for a large deployment was bounded by one process's loop time. Airflow 2.0's scheduler rewrite made active-active operation a supported, first-class deployment mode. ## The mechanism: database row locks There is no leader election and no partitioning of DAGs across schedulers. Every scheduler runs the same loop: find DAG runs that need creating, find task instances whose dependencies are met, queue them. Correctness comes from the metadata database. Before a scheduler acts on a set of rows it takes **row-level locks** with `SELECT ... FOR UPDATE`. Because the locks are held only for the duration of a short critical section and use `SKIP LOCKED`, a second scheduler issuing the same query simply receives a disjoint set of rows rather than blocking behind the first. The consequence is straightforward: exactly one scheduler transitions any given task instance to `queued`, so no task is dispatched twice, and work spreads across schedulers naturally by whoever grabs rows first. ## What that requires of the database Because the guarantee lives in the database, the backend must actually implement the locking semantics — in practice PostgreSQL or MySQL 8. This is the main reason Airflow's documentation is explicit that certain older or alternative backends are not supported for multi-scheduler deployments, and why SQLite is out of the question. It also means the metadata database becomes the scaling bottleneck. Every additional scheduler adds queries, locks and writes. Symptoms of over-subscribing look like scheduling latency going *up* as you add schedulers, lock waits in `pg_stat_activity`, and autovacuum falling behind on the `task_instance` table. Tuning `[scheduler] max_tis_per_query` and giving the database more headroom usually matters more than adding a fourth scheduler. ## What each scheduler must see All schedulers must have access to the same DAG files (baked image, git-sync, shared volume) and the same provider packages. If one scheduler cannot import a DAG, the behaviour is confusing rather than dramatic: that scheduler simply never contributes work for it while the others do, and you get intermittent, hard-to-explain scheduling gaps. The same rule applies to workers, so most deployments distribute one artefact to everything. Running a **standalone DAG processor** (`airflow dag-processor`) separates parsing from scheduling, which makes multi-scheduler setups cleaner: parsing happens once, in its own process, and schedulers work from serialized DAGs in the database. Airflow 3 makes that separation mandatory. ## What it buys you - **Failover.** Killing a scheduler — a node dying, a rolling upgrade — degrades throughput but never stops scheduling. Tasks in flight are unaffected; tasks the dead scheduler had queued are adopted or reset when it comes back or by its peers. - **Latency.** With thousands of task instances per loop, more schedulers means the queue-eligible set is drained faster, cutting the time between "dependencies satisfied" and `queued`. ## What it does not buy you It does not increase execution capacity — that is the executor's and workers' job. It does not remove the metadata database as a single point of failure; the database is still the thing to make highly available. And it does not protect against a bad DAG file: a DAG whose parse takes 60 seconds slows every scheduler that parses it. ## How many to run Two is the usual starting point, purely for availability. Go to three or more only with evidence: a measured scheduling delay (Airflow exposes scheduler loop and task-queueing metrics via StatsD/OpenTelemetry) and a database with the headroom to absorb it.
- Does running more schedulers make your pipelines run faster?Only the scheduling half. More schedulers shorten the delay between a task's dependencies being satisfied and it reaching the queued state, which matters in deployments with thousands of task instances per loop. Actual execution throughput is set by the executor and worker capacity, so if tasks are already piling up in `queued`, adding a scheduler makes things marginally worse by adding database load rather than better.
- What is the first thing to watch when you add a second or third scheduler?The metadata database. Each scheduler adds queries, row locks and writes, so watch lock waits, connection count, CPU, and autovacuum health on the `task_instance` and `dag_run` tables. If scheduling latency rises rather than falls after adding a scheduler, the database is the constraint and the fix is database capacity or query tuning, not more scheduler processes.
- What happens to tasks a scheduler had already queued when that scheduler dies?Nothing kills them — a queued or running task belongs to the executor and worker, and running tasks finish normally. Task instances the dead scheduler had dispatched but whose executor state is lost become orphans; a surviving or restarted scheduler adopts or resets them so they are re-dispatched. This is the same orphan-adoption path used after any scheduler restart.
saying these in an interview costs you the question
- Claims Airflow elects a scheduler leader via ZooKeeper or etcd
- Thinks DAGs are statically partitioned across schedulers
- Believes more schedulers means more task execution capacity
- Ignores the metadata database as the limiting resource
- Assumes any database backend supports multi-scheduler HA