skip to content

When should a pipeline generate its tasks dynamically at runtime instead of declaring a fixed graph?

level: seniorimportance: nice to knowfreq 36%

answer

  1. you do not know the work until you start
  2. one unit fails, redo only that unit
  3. the graph is a different size each day
  4. the scheduler and the target both feel it
  5. cap it, batch it, record what expanded

basics

~20 s

Only when the set of work units is genuinely unknown until the run starts and each unit deserves its own retry and visibility — for example one task per tenant discovered at run time. Otherwise a fixed graph is cheaper to read, compare and operate.

solid answer

~50 s

Dynamic generation means the number of tasks is computed when the run starts, from a list of files, partitions or tenants that is not known when the pipeline is written. It is worth it when each unit needs to fail, retry and be observed independently — one bad tenant should not force reprocessing the other forty. It costs you a graph whose shape changes between runs, which makes run-to-run comparison and historical views harder, and it puts real load on the scheduler and on whatever the tasks call. So bound it: cap the fan-out, batch many units into each task rather than one task per row, and prefer generating from the run's own parameters over an expensive discovery query. If the units are tiny or number in the thousands, push the parallelism down into a data engine and keep one task. Also record the generated unit list, or a rerun of an old window will expand differently than it did originally.

code

python · 7 lines
python
# Bounded expansion derived from the run's own window
def plan_units(interval_start, max_units=200, batch=50):
    units = list_partitions(interval_start)          # not from now()
    if len(units) > max_units:
        raise RuntimeError(f"unexpected expansion: {len(units)} units")
    record_run_artifact(interval_start, units)       # reproducible rerun
    return [units[i:i + batch] for i in range(0, len(units), batch)]

go deeper

for a junior

Know the idea: sometimes the number of tasks depends on data only known when the run starts, such as one task per file found — and that a fixed, declared graph is the normal case.

for a middle

Explain the tradeoff: per-unit retry and per-unit visibility bought at the price of a graph whose shape changes between runs, plus per-task scheduling overhead that adds up quickly.

for a senior

Show the operational bounds you would impose — a cap on width, batching units per task, throttling against downstream limits, one aggregated alert — and explain why the unit list must come from the run's own window and be recorded.

for a principal

Decide where parallelism belongs: orchestrator fan-out buys observability and retry granularity, an engine buys throughput. Set the platform rule for which units justify their own node, and keep expansion bounded so one bad upstream cannot launch a herd.

## What dynamic generation is Most pipelines are declared statically: the tasks and edges are written down, and every run has the same shape. Dynamic generation instead computes the node set at run time — list the objects under a prefix, query the active tenants, read the partitions that changed — and creates one task per element. The graph is still acyclic and still schedulable; the difference is only that the number of nodes varies per run. ## The case for it The justification is always **independent failure and independent visibility**. If forty tenants are processed in one task, one tenant's malformed file fails the task and a retry redoes all forty. With one task per tenant you get: - a retry scoped to the tenant that failed; - a per-tenant duration and log, so "tenant 17 has been getting slower" is visible; - true parallelism the scheduler can place across workers; - a rerun that touches only what needs redoing. That is a real benefit and it is why the pattern exists. ## The case against it **The graph shape stops being stable.** Yesterday's run had forty nodes, today's has sixty-three. Comparing runs, reading history, and reasoning about "the pipeline" all get harder, because there is no single pipeline shape any more. A node that existed yesterday and not today is not a failure — but it looks like an absence. **Scheduling cost is per task.** Every node carries queueing, state transitions, metadata rows and log streams. A run expanding to thousands of tiny tasks can make the scheduler itself the bottleneck while the actual work is seconds. **Fan-out hits things downstream.** Two hundred simultaneous tasks means two hundred concurrent warehouse sessions or API callers. The graph will happily generate a thundering herd against a system with a much smaller safe concurrency. **Alert volume scales with the fan-out.** One bad source can produce a hundred failed tasks and a hundred notifications, all with the same root cause. ## Where the unit list comes from, and when The list should be derived from the run's own inputs — its data interval and parameters — or produced by an explicit upstream task whose output is recorded. Two rules follow: - **Do not compute the list from wall-clock `now()`.** Rerunning an old window would then expand over today's units instead of the ones that window actually had, so the rerun silently does different work than the original. - **Record the expansion.** Write the unit list as an artifact or run metadata. It makes a rerun reproducible and it answers "was tenant 17 in scope on the 3rd?" months later. Computing the list when the pipeline definition is *parsed* rather than when a run *starts* is worse still: the query runs on every parse, and the graph you inspect may not be the one that ran. ## How to bound it - **Cap the width.** Refuse to expand beyond a limit, or fail loudly when the list is unexpectedly huge — an upstream glitch producing fifty thousand units should page, not launch. - **Batch.** One task per *group* of units — per hundred files, per region — keeps independent retry at a useful granularity without one node per row. - **Throttle.** Constrain how many of the generated tasks run at once, sized to what the downstream system tolerates, not to what the workers can do. - **Aggregate the outcome.** Follow the fan-out with a single task that summarises how many units succeeded and enforces a floor, so alerting fires once with a count instead of once per unit. ## The alternatives **Loop inside one task.** Simple, one retry unit, one log, no scheduler pressure — but a failure at unit thirty-nine of forty redoes everything unless the task tracks its own progress. Right when units are small and numerous. **Push it to the engine.** Hand the whole prefix or partition set to a processing engine that parallelises natively. The orchestrator keeps one node and the engine does what it is built for. Right when there are thousands of units. **Keep a static graph over a stable list.** If the tenants change twice a year, a declared list in configuration is more legible than run-time discovery, and changes go through review. ## A rule of thumb Generate dynamically when the units are **few enough to read** (dozens to low hundreds), **coarse enough to matter** (minutes each), **genuinely independent**, and **unknown until the run**. Break any of those and prefer batching, an internal loop, or the engine. ## The downstream half of the pattern A fan-out almost always wants a reduce step: one task that runs after all generated tasks, checks how many succeeded, enforces a completeness floor, and performs the single publish. Without it, the pipeline has no place to say "this run covered thirty-eight of forty tenants", and consumers cannot distinguish a complete day from a partial one. The fan-in policy discussion applies here in full — the aggregator should be strict about failures unless partiality is explicitly modelled. ## Interview framing The strong answer is not "dynamic generation is powerful". It is: independent retry and visibility are the benefit; variable graph shape, scheduler load, downstream herd and alert noise are the costs; and here are the bounds — cap, batch, throttle, aggregate, record the list — that make it safe.

  • How do you keep a dynamically generated fan-out from overwhelming the system the tasks call?
    Throttle how many generated tasks run at once, sized to what the target tolerates rather than to available workers, and batch units so each task covers a group instead of a single item. Add a hard cap on the expansion itself so an upstream glitch producing tens of thousands of units fails loudly instead of launching.
  • What breaks when you rerun an old window of a dynamically generated pipeline?
    If the unit list is discovered from the current state of the world, the rerun expands over today's units rather than the ones that window originally had — different tasks, different work, no reproducibility. Derive the list from the run's own interval and parameters, and record the expansion as run metadata so you can see and repeat exactly what was in scope.
  • When is looping inside a single task better than generating one task per unit?
    When units are small and numerous, when they share expensive setup, or when per-task overhead would dominate the actual work. You give up per-unit retry and visibility, so it suits cases where redoing the whole batch is cheap — or where the task tracks its own progress and skips completed units on a retry.
  • One upstream glitch causes 150 generated tasks to fail. How do you keep alerting useful?
    Do not alert per task. Route the fan-out into a single aggregation task that counts successes and failures, enforces a completeness floor, and emits one notification carrying the count and a sample of causes. Page on the run's overall state and on output freshness, not on each generated node.

saying these in an interview costs you the question

  • Generates one task per row or per tiny file
  • Discovers the unit list from current time instead of the run's window
  • Ignores that fan-out width hits downstream concurrency limits
  • Never records which units a run expanded to
  • Alerts once per generated task on a shared root cause

context