skip to content

In an Airflow sensor, what is the difference between poke mode and reschedule mode?

level: middleimportance: must knowfreq 70%

answer

  1. one holds the slot, one gives it back
  2. think about what the task's state is between checks
  3. up_for_reschedule is not a retry
  4. the trade is slot occupancy versus startup cost per check

basics

~20 s

Poke mode keeps the sensor task running for the whole wait, sleeping between checks and holding a worker slot. Reschedule mode ends the task after each failed check and re-queues it, freeing the slot until the next check is due.

solid answer

~50 s

Both modes call the same `poke()` method; they differ in what happens between calls. With `mode='poke'` (the default) the task instance stays **running** on a worker for the entire wait, sleeping `poke_interval` seconds between checks — one sensor, one occupied slot, for hours if need be. With `mode='reschedule'` the task exits after a failed poke, goes to the `up_for_reschedule` state, and Airflow re-queues it when the next poke is due; between checks it holds nothing. The rule of thumb: waits of seconds with a short `poke_interval` can stay in poke mode, because each reschedule costs a full scheduling round-trip and a fresh task startup. Waits of many minutes or hours should use reschedule mode, or a deferrable sensor. Note that `timeout` still measures the whole wait, not one attempt, and that reschedule mode keeps nothing in memory across pokes — `poke()` must recompute its state every time.

code

python · 20 lines
python
from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor

# Holds a worker slot for the entire wait
bad = S3KeySensor(
    task_id="wait_poke",
    bucket_name="landing",
    bucket_key="vendor/{{ ds }}/data.csv",
    poke_interval=30,
    timeout=60 * 60 * 8,
)

# Releases the slot between checks; total budget unchanged
good = S3KeySensor(
    task_id="wait_reschedule",
    bucket_name="landing",
    bucket_key="vendor/{{ ds }}/data.csv",
    mode="reschedule",
    poke_interval=300,
    timeout=60 * 60 * 8,
)

go deeper

for a junior

Recall that mode is a sensor argument with two values, and that the default keeps the task running the whole time while reschedule lets it stand down between checks.

for a middle

Explain the mechanics both ways: the poke loop inside one execution versus a task that exits and is re-queued, and why the second trades slot occupancy for per-attempt startup cost.

for a senior

Be ready to diagnose the outage: worker capacity full of sleeping sensors while real tasks queue. Talk about poke_interval sizing, timeout as a total budget, soft_fail semantics, and when you would defer instead.

for a principal

Own the standard: what the default sensor configuration is across the fleet, whether long waits are allowed at all, and how you would move teams from polling to event-driven triggering.

## The failure this question is really about A team puts a `S3KeySensor` at the top of forty DAGs, each waiting for a vendor file that usually lands at 09:00 but sometimes at 14:00. All forty sensors start at 06:00 in the default mode. Every one of them is a *running* task instance holding a worker slot for eight hours. The deployment's capacity is gone; the actual load tasks, and every other DAG, sit queued behind sleeping sensors. Nothing has failed, nothing has errored, and the platform is dead. That is the scenario `mode='reschedule'` was introduced for. ## What each mode does mechanically Every sensor implements `poke(context) -> bool`. `BaseSensorOperator.execute()` drives the loop. **`mode='poke'`** — the default. `execute()` calls `poke()`; if it returns `False`, the process sleeps `poke_interval` seconds and calls it again, inside the same task-instance execution. The task state remains `running` from the first check to the last, so the worker slot, and any pool slot the task holds, are occupied for the full duration. Advantages: no scheduling overhead per check, sub-minute polling is cheap, and anything the sensor computed stays in memory. **`mode='reschedule'`** — after a `False` poke, the task raises an internal reschedule signal, the process exits, and the task instance moves to the `up_for_reschedule` state. Airflow records when the next attempt is due (it stores reschedule records so it knows when the wait started) and the scheduler re-queues the task then. Between attempts, the sensor holds no worker capacity at all. Costs: each attempt is a full task startup — scheduler pickup, process launch, DAG file parse, connection setup — so a `poke_interval` of a few seconds turns into constant churn on the scheduler and metadata database. ```python S3KeySensor( task_id="wait_for_vendor_drop", bucket_name="landing", bucket_key="vendor/{{ ds }}/data.csv", mode="reschedule", poke_interval=300, timeout=60 * 60 * 8, ) ``` ## Consequences you should be able to state - **Timeout spans the whole wait.** In reschedule mode `timeout` is not per attempt; Airflow measures from the first attempt so the sensor still gives up after the configured total wait rather than restarting the clock each time. - **No in-memory state survives.** A poke-mode sensor can cache a client, a cursor or a partial result between checks. A reschedule-mode sensor starts a fresh process every time, so `poke()` must be self-contained and cheap — and, crucially, it should not have side effects, because it will run many times. - **Startup cost per poke.** If the DAG file is slow to import, every reschedule pays that import again. Long parse times plus short poke intervals is a bad combination. - **Log shape changes.** A poke-mode sensor writes one log for the whole wait; a reschedule-mode sensor produces a log per attempt, which people occasionally mistake for retries. ## The neighbouring knobs - `soft_fail=True` makes a timed-out sensor go to `skipped` instead of `failed`. Use it when the awaited thing is genuinely optional — but understand that skips propagate to downstream tasks under the default trigger rule. - `exponential_backoff=True` grows the interval between pokes, which pairs well with a long timeout on something that is either early or very late. - `deferrable=True`, where the provider supports it, is the third option and generally the best one for long waits: the wait is handed to the async triggerer process, and hundreds of waits cost one process rather than hundreds of task startups. Reschedule mode frees the slot but still pays a full task launch per check; deferral pays neither. ## Choosing Wait measured in seconds, sub-minute polling, cheap DAG parse: poke. Wait measured in tens of minutes to hours: reschedule with a `poke_interval` of minutes, or deferrable if available. And if the awaited event has a producer you control, the strongest answer is to remove the wait — have the producer trigger the consumer instead of paying anything to poll.

  • What does soft_fail=True change about a sensor that times out?
    Instead of failing, the sensor is marked skipped. That keeps a genuinely optional wait from paging anyone, but under the default all_success trigger rule the skip propagates to everything downstream, so a task that must still run needs an explicit trigger rule such as none_failed_min_one_success.
  • Why is a five-second poke_interval a bad idea in reschedule mode?
    Every check is a full task launch: scheduler pickup, process start, DAG file parse, connection setup, plus metadata database writes. At five-second intervals that overhead dwarfs the work, hammers the scheduler and the database, and can end up costing more than simply holding one slot.
  • How would you tell from the UI that a sensor is in reschedule mode rather than retrying?
    Look at the state between checks: reschedule mode parks the task in up_for_reschedule and does not increment the try number, while retries show up_for_retry after a failure and each attempt is a new try. Reschedule logs read as repeated clean pokes, not as errors.

saying these in an interview costs you the question

  • Says reschedule mode restarts the timeout clock each attempt
  • Uses reschedule mode with a five-second poke interval
  • Thinks a sleeping poke-mode sensor consumes no capacity
  • Confuses up_for_reschedule with up_for_retry after a failure
  • Keeps state in memory across pokes in reschedule mode

context