In Airflow, how does declaring a DAG with the @dag decorator differ from the DAG() context manager?
answer
- one style attaches, the other manufactures
- the with block adopts what it contains
- a factory you never call builds nothing
- dag_id defaults to the function name
- both serialize to the same thing
basics
~20 sBoth produce the same DAG. The DAG() context manager registers every operator built inside the with block; the @dag decorator turns a function into a DAG factory that you must call at module level, or nothing is registered.
solid answer
~50 sThey are two authoring styles over the same object. `with DAG(dag_id=..., schedule=..., start_date=...) as dag:` opens a context that automatically attaches any operator constructed inside it. `@dag(...)` decorates a function whose body builds the tasks; the decorated function is a factory, and **you have to invoke it at module level** — `my_pipeline()` — for the DAG to exist in the parsed file. Forgetting that call is the single most common reason a new file shows no DAG in the UI. The decorator is the TaskFlow-native style: the function's parameters become DAG params, and calling it more than once with different arguments is a clean way to generate a family of DAGs. There is also a third, older style: pass `dag=dag` to every operator explicitly. All three serialize identically, so the choice is readability, not capability.
code
python · 21 linesimport pendulum
from airflow.decorators import dag, task
@dag(
schedule="@daily",
start_date=pendulum.datetime(2024, 1, 1, tz="UTC"),
catchup=False,
tags=["sales"],
)
def daily_sales():
@task
def extract():
return [1, 2, 3]
@task
def load(rows):
print(len(rows))
load(extract())
daily_sales() # required: without this the module defines no DAGgo deeper
Be able to write a minimal DAG in either style from memory, including dag_id, start_date, schedule and catchup, and remember to call the decorated function.
Explain the mechanics: the with block pushes a context that operators attach to, while the decorator produces a factory whose body runs at parse time when invoked.
Show judgment about which style suits which pipeline, and be able to run the not-appearing-in-the-UI checklist in order rather than guessing.
Own the convention across a repo — one authoring style, shared default_args, a naming scheme for dag_id and tags — so that hundreds of DAGs stay navigable and no two files can claim the same id.
## Same DAG, two syntaxes Airflow needs each DAG file, when parsed, to leave one or more `DAG` objects reachable in the module's global scope. Everything else is sugar over that requirement. ### The context manager ```python with DAG( dag_id="daily_sales", schedule="@daily", start_date=pendulum.datetime(2024, 1, 1, tz="UTC"), catchup=False, ) as dag: extract = PythonOperator(task_id="extract", python_callable=pull) load = PythonOperator(task_id="load", python_callable=push) extract >> load ``` Entering the `with` block pushes the DAG onto a context stack; every operator constructed while that context is active attaches itself to it, so you never write `dag=dag`. The DAG object is bound to a module-level name, which is what the parser picks up. ### The decorator ```python @dag( schedule="@daily", start_date=pendulum.datetime(2024, 1, 1, tz="UTC"), catchup=False, ) def daily_sales(): @task def extract(): return pull() @task def load(rows): push(rows) load(extract()) daily_sales() # <- without this line the file defines no DAG ``` `@dag` converts the function into a factory. The `dag_id` defaults to the function name, and the decorated function's body runs *at parse time* when you call it, building the task objects. The trailing call is not optional: a decorated-but-never-called function leaves no DAG object behind, the file parses cleanly, and the UI simply shows nothing. It is worth memorising because it looks like nothing is wrong. ### The explicit style ```python dag = DAG(dag_id="daily_sales", schedule="@daily", start_date=START, catchup=False) extract = PythonOperator(task_id="extract", python_callable=pull, dag=dag) ``` Still valid, still occasionally clearer when tasks are built by a helper function called from several places, but repetitive enough that most codebases have moved off it. ## When each style earns its place **Reach for the decorator** when the pipeline is mostly plain Python, because it pairs naturally with `@task`: dependencies come from calling one function with another's output, so you rarely write an arrow at all. It also gives you parameterisation for free — arguments of the decorated function become DAG params with defaults, exposed for a triggered run — and calling the factory in a loop is a tidy pattern for generating one DAG per region, tenant or source system, each with a distinct `dag_id` supplied through `@dag(dag_id=...)` or an `.override(dag_id=...)` call. **Reach for the context manager** when the DAG is a wiring of heavyweight provider operators — transfer operators, Kubernetes pod operators, SQL operators — where there is no Python function to decorate anyway, or when a large existing codebase already reads that way. The two styles also mix freely: a `with DAG(...)` block can contain `@task` functions, and a decorated DAG can contain classic operators, joined with `>>` or via the operator's `.output` attribute. ## Parameters that matter either way Whichever syntax you use, the same constructor arguments decide behaviour: `dag_id` (unique across the whole deployment — two files defining the same id is undefined behaviour and shows up as a mysterious flip-flopping DAG), `start_date`, `schedule`, `catchup`, `max_active_runs`, `max_active_tasks`, `default_args` (a dict merged into every task, typically `owner`, `retries` and `retry_delay`), `tags` for UI filtering, and `params` for run-time inputs. Note the naming shift: `schedule_interval` was superseded by `schedule` in Airflow 2.4 and removed in Airflow 3, so seeing `schedule_interval` in code is a quick way to date it. ## Debugging "my DAG is not in the UI" The checklist is the same for both styles, in order: is the file under the configured DAGs folder; is the decorated factory actually called; did the file raise on import (import errors surface in the UI and via `airflow dags list-import-errors`); is the file excluded by `.airflowignore`; has the parse interval elapsed. Only after those does it get interesting.
- You add a new file using the @dag decorator and no DAG appears in the UI, but there is no import error. What do you check first?Whether the decorated function is actually invoked at module level. `@dag` only creates a factory; without a trailing `daily_sales()` call the module defines a function and no DAG object, so the file parses successfully and contributes nothing. After that, check the DAGs folder path, `.airflowignore`, and whether the parse interval has elapsed.
- How would you generate one DAG per source system from a list of names?Loop at module level and either call a `@dag`-decorated factory once per name with a distinct `dag_id`, or construct a `DAG` object per name and assign each to a unique global. Both work because the parser collects every DAG object in the module. Keep the driving list cheap to obtain — reading it from an API on every parse is the classic top-level-code mistake.
- What does the `default_args` dictionary do that per-task arguments do not?It is merged into every task in the DAG, so shared settings — `owner`, `retries`, `retry_delay`, `on_failure_callback` — are declared once instead of repeated. Task-level arguments win over it, so a single task can opt out. It is not magic scoping: it is applied when each task object is constructed.
saying these in an interview costs you the question
- Thinks the @dag decorator registers the DAG without calling the function
- Believes the two styles produce different scheduling behaviour
- Uses schedule_interval and cannot say it was replaced by schedule
- Reuses one dag_id across two files
- Assumes the decorator cannot host classic operators