skip to content

In Airflow, what does setting an sla on a task do, and how does it differ from execution_timeout?

level: middleimportance: should knowfreq 54%

answer

  1. a deadline, not a duration
  2. the clock starts at the interval, not the task
  3. nothing gets killed
  4. the callback lives on the DAG, not the task
  5. its sibling argument is the one that kills

basics

~20 s

In Airflow an sla is a timedelta measured from the DAG run's data interval end; if the task has not succeeded by then, Airflow records an SLA miss and fires sla_miss_callback. It never interrupts the task, unlike execution_timeout, which kills a running attempt.

solid answer

~50 s

In Airflow, `sla` is a `timedelta` you set on a task, measured from the **end of the DAG run's data interval** — not from when the task started. If the task has not reached success by that point, Airflow records an SLA miss in the metadata database, surfaces it in the UI's SLA-miss view, and calls the DAG's `sla_miss_callback` (and sends SLA emails if configured). Crucially it is **purely observational**: the task keeps running, nothing is killed, and no state changes. `execution_timeout` is the opposite — it is measured from when the attempt started and it *kills* the attempt, marking it failed and feeding the retry machinery. Use `sla` to express 'this data must be ready by 6am' and `execution_timeout` to stop a hung task from occupying a worker forever. Note that the SLA feature was reworked in Airflow 3, so state the version you are assuming.

code

python · 24 lines
python
from datetime import datetime, timedelta
from airflow.decorators import dag, task

def on_sla_miss(dag, task_list, blocking_task_list, slas, blocking_tis):
    page(f"{dag.dag_id} late; blocking tasks: {blocking_task_list}")

@dag(
    schedule="@daily",
    start_date=datetime(2024, 1, 1),
    catchup=False,
    sla_miss_callback=on_sla_miss,   # DAG level, not task level
)
def sales():
    @task(execution_timeout=timedelta(minutes=45))   # bounds a hung attempt
    def load_raw():
        ...

    @task(sla=timedelta(hours=6))                    # deadline from interval end
    def publish_mart():
        ...

    load_raw() >> publish_mart()

sales()

go deeper

for a junior

Recall that sla is a timedelta expressing a deadline and that missing it produces a notification and a UI entry rather than stopping the task.

for a middle

Explain the anchoring to the data interval end, that the callback lives on the DAG while sla lives on the task, and how it contrasts with execution_timeout.

for a senior

Show the operational judgment: SLAs cover lateness where failure alerting is blind, they belong on the publishing task not every task, and they produce meaningless noise during backfills.

for a principal

Own the freshness commitment end to end — which outputs carry a stated deadline, whether the orchestrator or an independent freshness check is the source of truth, and how you version-proof this given Airflow 3 reworked the SLA model.

## What an SLA is in Airflow An Airflow `sla` is a `timedelta` argument on an operator. It expresses a **deadline relative to the DAG run's data interval end**, not to the task's start time and not to wall clock. For a `@daily` DAG, the run covering 2024-03-11 has a data interval ending at 2024-03-12 00:00 UTC; an `sla=timedelta(hours=6)` on a task in that run means 'this task should have succeeded by 06:00 UTC on 2024-03-12'. That anchoring is the detail people get wrong. Because the deadline is fixed to the interval, a run that starts late — because the scheduler was down, because an upstream sensor waited, because a backfill queued it — eats into the same budget. This is a feature: it matches how a business deadline actually works ('the mart is ready by 6am'), not how a duration budget works. ## What happens on a miss When the deadline passes with the task not successful, Airflow: 1. Records an **SLA miss** row in the metadata database. 2. Surfaces it in the UI's SLA-miss listing (Browse → SLA Misses in Airflow 2). 3. Invokes the DAG-level **`sla_miss_callback`** if one is defined, passing it the DAG, the list of task lists and blocking task instances, and the SLA-miss records. 4. Sends an SLA-miss email to the `email` addresses if email is configured. And that is all it does. **The task is not stopped, not failed, not retried, and its state is unchanged.** An SLA miss is a notification about lateness, deliberately decoupled from execution control. ```python from datetime import timedelta def on_sla_miss(dag, task_list, blocking_task_list, slas, blocking_tis): page(f"{dag.dag_id} missed its SLA; blocking: {blocking_task_list}") @dag(schedule="@daily", start_date=..., sla_miss_callback=on_sla_miss) def sales(): @task(sla=timedelta(hours=6)) def publish_mart(): ... ``` The `sla_miss_callback` lives on the **DAG**, while `sla` lives on the **task** — a common configuration mistake is defining the callback per task and wondering why nothing fires. ## How it differs from execution_timeout | | `sla` | `execution_timeout` | |---|---|---| | Measured from | data interval end | the attempt's start | | On breach | records a miss, calls the callback | kills the attempt, marks it failed | | Affects task state | no | yes — feeds retries | | Answers | 'is the data late?' | 'is this attempt hung?' | They are complementary and most production tasks want both. `execution_timeout=timedelta(minutes=45)` stops a query that wedged on a lock from holding a worker slot all night. `sla=timedelta(hours=6)` tells you the downstream dashboard will be stale even though nothing has technically failed yet. A third relative is `timeout` on a sensor, which bounds how long the sensor pokes before failing — again an execution control, not an observation. ## Why SLAs are the answer to a specific blind spot Failure alerting cannot see slowness. A DAG that is running six hours behind schedule produces no failed task, no `on_failure_callback`, and no email — everything is green and the data is not there. The SLA mechanism exists precisely to convert 'late' into a signal. That is the point worth making in an interview: retries and failure callbacks cover *broken*, SLAs cover *late*, and a pipeline with a business commitment needs both. The complementary blind spot SLAs do **not** cover is a run that never existed. If the DAG is paused or `catchup=False` skipped the interval, there is no task instance and therefore no SLA evaluation. Absence still needs an external freshness check on the output table. ## Practical cautions - **Do not put an `sla` on every task.** SLA misses on intermediate tasks generate noise proportional to DAG size; put the deadline on the task that publishes the thing someone actually consumes. - **The miss fires per task instance**, so a wide fan-out with SLAs can produce a storm; the DAG-level callback receives them batched, which is one reason to handle them there. - **SLAs interact badly with `catchup=True` backfills**, because historical intervals have deadlines long in the past and will be recorded as missed the moment they run. Teams commonly disable or ignore SLA output during a large backfill. - **Version matters.** The classic `sla` + `sla_miss_callback` behaviour above is Airflow 2.x. Airflow 3 reworked this area in favour of a deadline-alerting model, so if you are answering about a specific deployment, say which major version you mean. Do not assert 2.x mechanics as if they were unchanged in 3. ## The judgment layer A good answer connects the mechanism to the commitment. Someone downstream expects a table by a time; that expectation is a *freshness* commitment on the output, and the `sla` on the publishing task is how you encode it inside the orchestrator. Encoding it as an `execution_timeout` instead would be wrong on both axes: it would kill work that was going to succeed, and it would stay silent when a fast task simply started four hours late.

  • A task with sla=timedelta(hours=2) takes 30 minutes but the DAG run started 3 hours late. Is the SLA missed?
    Yes. The deadline is measured from the DAG run's data interval end, not from when the task started, so a late start consumes the same budget. That is intentional: the SLA encodes 'the data is ready by this time', which is exactly what a late start breaks even when every task ran quickly.
  • Where do you define sla and where do you define sla_miss_callback?
    `sla` goes on the task/operator; `sla_miss_callback` goes on the DAG. Putting the callback on a task is a frequent misconfiguration that produces recorded misses in the UI with no notification. Receiving misses at the DAG level is also better behaviour for wide DAGs, since they arrive batched rather than one message per task.
  • Why do SLAs behave badly during a large catchup backfill?
    Every historical data interval has a deadline that is already long past, so each backfilled run records an SLA miss the instant it runs. The misses are technically correct and operationally meaningless. Teams typically suppress SLA notification during a planned backfill rather than let the noise train people to ignore the channel.
  • If a task must not run longer than 45 minutes, is sla the right argument?
    No — `sla` never interrupts anything. `execution_timeout=timedelta(minutes=45)` is the argument that bounds an attempt: it kills the running task, marks it failed, and hands it to the retry logic. Use `sla` alongside it if you also want to know when the output is late for a reason other than a hang.

saying these in an interview costs you the question

  • Thinks an SLA miss kills or fails the task
  • Measures the SLA from when the task started running
  • Puts sla_miss_callback on the task instead of the DAG
  • Uses sla where execution_timeout is meant, to bound runtime
  • Assumes SLA alerting catches a DAG that never ran at all

context