In Luigi, what determines that a task is already complete and can be skipped?
answer
- state lives outside the scheduler
- ask the storage, not the run log
- one method, and it defaults to existence
- delete it to make it run again
basics
~20 sLuigi calls the task's complete() method, which by default returns True only when every Target returned by output() reports exists(). Completion is a property of the data on storage, not of any run history, so reruns skip whatever already exists.
solid answer
~50 sLuigi asks the task itself: `complete()` is called before scheduling, and its default implementation returns True only if **every** Target from `output()` answers `exists()`. There is no run-history database consulted for correctness — the state lives in the storage layer, in the files or tables the pipeline produced. That is the whole idempotency story: rerun the same command with the same parameters and Luigi rebuilds the graph, skips every task whose output is already there, and executes only the missing ones, so a pipeline that died at step 7 of 10 resumes at step 7. The flip side is that *existence is not correctness*. Editing a task's code does not invalidate yesterday's output, and a truncated or empty file still counts as done — to force a rerun you delete the Target, or override `complete()` with a stricter check.
code
python · 19 linesimport luigi
# default completion: every Target from output() must exist
class DailyExtract(luigi.Task):
date = luigi.DateParameter()
def output(self):
return luigi.LocalTarget(f"/data/raw/{self.date:%Y-%m-%d}/rows.csv")
def run(self):
with self.output().open("w") as f:
f.write(pull(self.date))
# stricter completion: only a _SUCCESS marker counts as done
class DailyExtractChecked(DailyExtract):
def complete(self):
marker = luigi.LocalTarget(f"/data/raw/{self.date:%Y-%m-%d}/_SUCCESS")
return marker.exists()go deeper
Remember the rule in one line: a Luigi task is done when the files its output() names already exist, so rerunning the pipeline only fills in what is missing.
Be able to state the default complete() implementation, explain that all Targets must exist when there are several, and show how deleting an output is what forces a rerun.
Demonstrate that you know where existence lies — partial writes, empty results, changed code, late-arriving source data — and what you do about each in a pipeline you operate.
Own the tradeoff: data-as-state removes an entire class of metadata drift but gives up run history, cascade invalidation and clear-and-rerun, which is a real cost once auditing and reprocessing policy matter.
## Completion is a property of the data Most orchestrators keep a record of runs: a metadata database says task X for run Y succeeded, and that record is what decides whether to run it again. Luigi does the opposite. It asks the **data**. Before scheduling anything, Luigi calls `complete()` on each task in the graph, and the default implementation is essentially: ```python def complete(self): outputs = flatten(self.output()) return all(t.exists() for t in outputs) ``` So a task is done when the artifacts it promised are present — every one of them, if `output()` returns several. Nothing else is consulted: not the exit code of a previous attempt, not a log, not the central scheduler's memory. This is the single idea the leaf exists for, and it is the idea later tools inherited in various forms. ## Why this makes reruns cheap Because the answer is recomputed from storage every time, re-invoking the pipeline is safe and self-healing. Luigi rebuilds the graph by walking `requires()`, evaluates `complete()` on each node, and runs only the incomplete ones. A ten-task pipeline that failed on step seven, rerun an hour later, does exactly three tasks' worth of work. There is no "clear the task state" ritual and no metadata store to keep in sync with reality — if you copy the outputs to a new machine, the pipeline there also considers them done, because the truth travels with the data. Parameters extend this per instance. `AggregateOrders(date=2026-08-19)` and `AggregateOrders(date=2026-08-20)` have different output paths, so each date is independently resumable, and running the task for a list of dates naturally fills only the gaps. ## Forcing a rerun The corollary trips people up: to make a task run again you delete its output. There is no "clear" command that marks a completed task as pending, because Luigi has nothing to clear. Rerunning after a logic change therefore means removing the affected artifacts, and if downstream tasks should also be rebuilt, removing theirs too — Luigi will not cascade an invalidation for you, since a downstream task whose file still exists is still complete. This is also why some teams put a content or code version into the output path (`/data/agg/v3/orders_2026-08-20.csv`): bumping the version changes the Target, so everything downstream becomes incomplete on its own. ## Where existence lies to you Existence is a proxy for correctness, and it is a leaky one: - A crashed task that wrote half a file leaves a Target that exists. The next run skips it and downstream tasks consume garbage. This is why Luigi's file targets write to a temporary path and move on close, and why writing directly to `self.output().path` is discouraged. - Empty output is indistinguishable from correct-but-empty output. A query that legitimately returns no rows and a query that failed after creating the file look identical. - Changed code with unchanged output paths means stale results silently persist. - Late-arriving source data does not invalidate anything; yesterday's aggregate stays "complete" even though the inputs it summarised have since grown. ## Overriding complete() `complete()` is an ordinary method, so you can make it stricter — check a row count, check that a `_SUCCESS` marker sits beside the data, check a checksum, or compare a stored input fingerprint. Two built-in shapes already override it. `luigi.WrapperTask` has no output at all and defines completion as "all my requirements are complete", which makes it a usable root node. `luigi.ExternalTask` has an output but no `run()`, so it is either complete because someone else produced the data or it blocks the graph. Overriding is powerful and easy to abuse: `complete()` is called often and must stay cheap and side-effect free. A completion check that runs an expensive warehouse query on every scheduling pass will dominate the runtime of a wide graph. ## Contrast worth stating out loud Run-history orchestrators can express things Luigi cannot — "this interval ran, even though it produced no rows", "this attempt failed three times", "clear and rerun without touching the data". Luigi trades all of that for a model with no state to corrupt: the pipeline's notion of progress and the data are the same object. Being able to name both sides of that trade is what a middle-level answer looks like; asserting only that "Luigi is idempotent" is not. ## Interview framing Expect the follow-up "so how do you rerun a task after fixing a bug?" and the scenario "the task failed but the next run skipped it — why?". Both are testing whether you actually internalised that Luigi never asks what happened; it only asks what is there.
- You fixed a bug in a task's run(). How do you make Luigi execute it again?Delete the Target it produced, and the Targets of every downstream task that should be rebuilt — Luigi has no state to clear, so the only lever is the data. Teams that do this often encode a version in the output path (`/data/agg/v3/...`), so bumping the version instantly makes the task and everything below it incomplete without hand-deleting files.
- When is overriding complete() the right move, and what is the risk?Override it when existence is a bad proxy for done — check a `_SUCCESS` marker, a row count, or a checksum instead. The risk is cost and purity: Luigi calls `complete()` on every node on every scheduling pass, so an expensive query or a side-effecting check will dominate a wide graph's runtime and can even mutate the state it is meant to observe.
- Does running with the central scheduler instead of --local-scheduler change how completion is judged?No. The central scheduler coordinates workers, prevents two of them running the same task instance, and drives the visualiser, but completion is still decided by calling `complete()` on the task, which by default inspects the Targets. Scheduling topology changes who runs what and when, never whether a task counts as done.
It is like deciding whether dinner is cooked by looking at the counter for the finished dish, rather than by reading the kitchen log to see whether anyone claims to have cooked it.
saying these in an interview costs you the question
- Says Luigi keeps a run-history database marking tasks as succeeded
- Assumes editing a task's code invalidates its existing output
- Believes a half-written or empty output file counts as incomplete
- Thinks the central scheduler versus local scheduler changes completion
- Claims there is a clear-state command to force a rerun