In Luigi, what do the requires(), output() and run() methods of a Task define?
answer
- a task declares, it is not registered
- what it needs, what it leaves behind
- the graph is walked, not written down
- self.input() mirrors one of the three
basics
~20 sA Luigi Task declares its upstream dependencies in requires(), the Target it will produce in output(), and the actual work in run(). Luigi walks requires() to build the graph and checks output() to decide whether run() is needed.
solid answer
~40 sA Luigi pipeline is just Python classes inheriting from `luigi.Task`, and three methods carry the whole contract. `requires()` returns the upstream **task instances** — constructed with their parameters, like `ExtractOrders(date=self.date)` — and Luigi walks that recursively to discover the dependency graph. `output()` returns one or more `Target` objects (`luigi.LocalTarget` for a file, contrib targets for S3, HDFS, databases); a Target is a handle that can answer `exists()`, not the data itself. `run()` is the body: it reads through `self.input()`, which mirrors the shape of `requires()` with each task replaced by its `output()`, and writes through `self.output()`. Parameters (`luigi.Parameter`, `DateParameter`) make instances distinct, so `AggregateOrders(date=...)` for two dates are two nodes, and identical instances are deduplicated into one.
code
python · 25 linesimport luigi
class ExtractOrders(luigi.Task):
date = luigi.DateParameter()
def output(self):
return luigi.LocalTarget(f"/data/raw/orders_{self.date:%Y-%m-%d}.csv")
def run(self):
with self.output().open("w") as f:
f.write(fetch_orders(self.date))
class AggregateOrders(luigi.Task):
date = luigi.DateParameter()
def requires(self):
return ExtractOrders(date=self.date)
def output(self):
return luigi.LocalTarget(f"/data/agg/orders_{self.date:%Y-%m-%d}.csv")
def run(self):
with self.input().open("r") as src, self.output().open("w") as dst:
dst.write(summarize(src))go deeper
Be ready to write a two-task Luigi pipeline from memory: a parameter, requires(), output() returning a LocalTarget, and run() reading self.input() and writing self.output().
Explain how the graph is discovered by recursively calling requires() on constructed task objects, and why parameters have to be threaded down the chain by hand.
Show the habits that keep such a pipeline operable: pure, cheap requires() and output(), deterministic paths, no side effects outside run(), and ExternalTask for data you do not own.
Frame the design tradeoff — declaring the graph in Python at instantiation time gives real expressiveness but couples parameters through every level, which is exactly what later DAG-object orchestrators separated out.
## The unit of work Luigi, open-sourced by Spotify, is a Python library for batch pipelines. There is no separate DAG file and no registry you must add a task to: a pipeline is a set of classes inheriting from `luigi.Task`, and the graph is *discovered* by asking one task what it needs. Three methods carry the entire contract, and almost every Luigi question circles back to them. ## requires(): what must exist first `requires()` returns the task instances this task depends on — a single task, a list, or a dict of them. Crucially it returns *objects*, not names or strings, and you construct them yourself with their parameters: ```python def requires(self): return ExtractOrders(date=self.date) ``` Luigi calls `requires()` on the task you asked for, then on every task that returns, recursively, until it reaches tasks with no requirements. That traversal *is* the graph — it is built bottom-up at runtime rather than declared in one place. A direct consequence is that parameters must be threaded explicitly down the chain: if a leaf task needs the date, every task between it and the root has to pass it along. ## output(): what the task leaves behind `output()` returns one or more `Target` objects. `luigi.LocalTarget` wraps a path on the local filesystem; contrib modules add targets for S3, HDFS, and databases. A Target is not the data — it is a handle that can answer `exists()` and usually open a stream. Two rules matter. First, `output()` must be **deterministic** for a given set of parameter values: the same parameters must yield the same path, so a path built from `datetime.now()` is a bug. Second, `output()` does no work; it only describes where the result will live. Luigi calls it before, during, and after the run, and uses it to decide whether the task has anything left to do at all. ## run(): the work `run()` is the body, and the only place side effects belong. It reads through `self.input()`, which mirrors the shape of `requires()` with each task object replaced by that task's `output()` — one task in, one Target out; a list in, a list out. It writes through `self.output()`. Idiomatically you write with `self.output().open('w')` rather than opening the raw path, because `LocalTarget` writes to a temporary file and moves it into place when the stream closes. ```python import luigi class AggregateOrders(luigi.Task): date = luigi.DateParameter() def requires(self): return ExtractOrders(date=self.date) def output(self): return luigi.LocalTarget(f"/data/agg/orders_{self.date:%Y-%m-%d}.csv") def run(self): with self.input().open("r") as src, self.output().open("w") as dst: dst.write(summarize(src)) ``` ## Parameters make instances A task class is a template; a task *instance* is the class plus a specific set of parameter values, declared as class attributes such as `luigi.Parameter()`, `luigi.DateParameter()` or `luigi.IntParameter()`. That pair is the node's identity. Two instances of the same class with the same values are the same node, so a diamond-shaped graph where two tasks both require `ExtractOrders(date=d)` runs the extraction once, not twice. Parameters are also how the command line addresses a task: `luigi --module pipelines.orders AggregateOrders --date 2026-08-20`. ## Two special shapes `luigi.ExternalTask` defines `output()` but no `run()`: it represents data produced by something outside the pipeline, and the graph simply waits for or fails on its existence. `luigi.WrapperTask` is the mirror image — it defines `requires()` and no output, and exists to be a convenient root that pulls in a batch of real tasks. ## Where beginners go wrong The most common mistake is doing upstream work inside `run()` — calling the extraction function directly instead of declaring `ExtractOrders` in `requires()`. That produces a task that always reruns everything and cannot be resumed. Others: putting side effects in `requires()` or `output()` (both are called repeatedly and must be cheap and pure), building a non-deterministic output path, and confusing `self.input()` (upstream Targets) with the task's own parameters. ## What an interviewer is checking That you understand a Luigi task describes itself rather than being registered somewhere, that the graph comes from recursive `requires()` calls, and that the output Target is a first-class part of the definition rather than an implementation detail of `run()`. That last point is the doorway to every deeper Luigi question, because the Target is also how Luigi decides whether the task is already done.
- What does self.input() return when requires() returns a list of three tasks?A list of three entries, in the same order, each holding whatever that task's `output()` returned — a single Target if the task has one output, or a list of Targets if it has several. `self.input()` always mirrors the *shape* of `requires()` with tasks swapped for their Targets, which is why a dict-returning `requires()` gives you a dict of Targets keyed the same way.
- Why must output() be deterministic for a given set of parameters?Because Luigi calls it separately to decide whether the task needs to run, to hand Targets to downstream tasks via `self.input()`, and to write the result. If the path embeds `now()` or a random suffix, those three calls disagree: the completion check never finds the file, downstream tasks look in the wrong place, and every invocation redoes the work.
- What is a luigi.ExternalTask used for?It models data the pipeline does not produce — a vendor drop, another team's export. It defines `output()` but no `run()`, so Luigi can only check whether the Target exists. Depending on one lets the rest of the graph express "this file must be here first" without pretending your pipeline can create it.
Each task is a recipe card that names the ingredients it needs, the dish it puts on the shelf, and the cooking steps — and the kitchen plans the whole meal by reading ingredient lists backwards.
saying these in an interview costs you the question
- Calls the upstream function inside run() instead of declaring it in requires()
- Thinks output() writes the file rather than describing where it will be
- Expects a decorator or DAG object to register the task somewhere
- Confuses self.input() with the task's own parameter values
- Builds an output path from the current timestamp