skip to content

Engine or Orchestrator

Both are called a DAG, a graph of steps with no cycles - but in one an edge carries records inside a single job, in the other it only means after. Confusing them is the classic design error.

on this pageshow

questions

4

A workflow scheduler and a cluster execution engine both draw a graph of steps with no cycles - what does an edge mean in each?

level: juniorimportance: must knowfreq 68%

answer

  1. same shape, different arrow
  2. does anything travel along it?
  3. success signal against records in flight
  4. delete the edge: what is lost?

basics

~20 s

In a workflow scheduler's graph an edge means only 'start after the previous step reported success'; nothing travels along it. Inside one program handed to a cluster execution engine, an edge carries the records themselves from one step to the next.

solid answer

~50 s

Both pictures are a set of steps with ordering edges and no cycles, and that shared shape is all they have in common. A workflow scheduler exists only to start other programs in a required order and record whether each finished; it moves no records itself, so its edge is a happens-after relation and anything one step produces for the next must be written somewhere both can reach and found there by a name you chose. A cluster execution engine is handed one whole program as a single submission and splits its work across machines; inside that submission an edge is data in flight, and the engine decides where the handoff physically lives - a buffer in memory, a file on a worker's own local disk, or durable shared storage between rounds in the oldest disk-to-disk designs. The test on any arrow: delete it, and do you lose records or only ordering?

go deeper

for a junior

Recall the one-line difference: one arrow means 'start after this finished', the other means 'these records go there'. If you can say which of the two moves bytes, you have the question.

for a middle

Explain the mechanics that follow from each meaning: who creates the intermediate, who names it, who deletes it, and why the engine's handoff can live in memory, on a worker's local disk or in durable storage depending on the design.

for a senior

Show the judgment by auditing a real picture arrow by arrow - which boundaries genuinely need a named durable intermediate, and which ones exist only because someone drew the whole pipeline in the scheduler.

for a principal

Frame it as a platform question: the two graphs answer to different owners and different review habits, and where you put a boundary decides who can see the pipeline's shape and what it costs to change it.

## Two pictures that share a word A **workflow scheduler** and a **cluster execution engine** both get drawn as boxes and arrows, and both graphs are acyclic: a set of steps with ordering edges and no cycles, a directed acyclic graph. The shape is genuinely the same. The shape is also the only thing that is the same, because the word says the arrows never loop back and says nothing at all about what an arrow *is*. Spelling the two systems out: - a **workflow scheduler** is a system whose only job is to start other programs in a required order and record whether each one finished; it moves no records itself; - a **cluster execution engine** is a system you hand a whole program to, which splits that program's work across many machines, runs the pieces and puts the results back together. Handing it that program - packaged with the resources it is asking for - is one *submission*, and the engine runs that submission end to end. ## The scheduler's edge means 'after' An edge from step A to step B in the workflow graph is a statement about *when*: do not start B until A has reported success. That is the whole of its content. Nothing travels along the arrow. Each step is usually its own process, started and then watched - sometimes it is a unit run inside a shared worker instead, but either way it has its own memory and that memory is gone when it exits. So the consequence candidates miss is immediate: **if B needs anything A produced, A must write it somewhere B can independently find, and B must be told where.** The scheduler supplies neither the storage nor the name. A table, a path in the shared storage every machine can read, a message somewhere - a real durable location, chosen, named and eventually cleaned up by you. ## The engine's edge carries records Inside one submission that the engine runs end to end, an edge between two steps means the records the upstream step produces become the input of the downstream step. The engine owns that handoff, and where it physically happens is its choice rather than yours - and this is exactly where real engines differ from each other: - some keep the handoff in memory and run adjacent steps as a single pass over the data; - some write it to a file on the producing worker's own local disk and have the consuming side fetch it across the network; - the oldest disk-to-disk designs in this family write each round's whole output to durable shared storage before the next round reads it, so a chain of rounds pays a full write and a full read between every pair. What is constant is not the mechanism but the **ownership**: you never name the intermediate, never create it, never delete it, and nothing outside the submission can see it. A claim like 'the engine never writes intermediates' is true of one design and false of another; a claim like 'the engine chooses where the intermediate lives' is true of all of them. ## Side by side | | edge in the workflow graph | edge inside one submission | |---|---|---| | what it asserts | B starts after A succeeded | records flow from A into B | | what travels along it | nothing | the records themselves | | the intermediate | you choose, name and delete it | the engine chooses; it is unnamed | | what sits at each end | a whole program | an operator inside one run, which starts no process of its own | | who can see it | anything that can read the location | nothing outside the submission | | what crossing costs | materialising the data, plus a second start-up | whatever handoff the engine picked | ## The test that separates them Put one question to every arrow in your picture: **if I deleted this edge, would records be lost, or only ordering?** 1. Records would be lost - the arrow is carrying data, so it belongs inside one submission where an engine can own the handoff. 2. Only ordering would be lost, and the two sides are separate programs - the arrow belongs in the workflow graph. 3. Both, because the downstream step reads a file the upstream step wrote - you have a scheduler edge with a data dependency bolted to the side of it, which is legitimate when that file is a real product other things read, and an accident when it exists only because the picture was drawn in the wrong system. ## Why the confusion is expensive The two mistakes are mirror images. Treat an engine edge as a scheduler edge and every intermediate becomes a named, written, read-back dataset with a start-up between the halves. Treat a scheduler edge as an engine edge and you get one enormous submission containing work that has no data relationship at all, where the boundaries that mattered to people are invisible. Neither is a scale problem and neither is fixed by a bigger cluster: they are both the same category error about what an arrow means.

  • If the scheduler's edge moves no data, how does a downstream step get the upstream step's output?
    Somebody writes it to a durable location both processes can reach - a table, a path in object storage - and the downstream step is told where to look. Choosing that location, agreeing the name and the encoding, and deciding who deletes it afterwards is real design work, and the edge supplies none of it.
  • Can a boundary inside one submission ever behave like an ordering-only edge?
    In effect, yes: many engines have points where everything upstream must finish before anything downstream starts. But the engine still owns that boundary, the intermediate stays unnamed, and nothing outside the submission can see it or start from it. A boundary the outside world can observe, point another program at, or begin from is a scheduler edge.
  • Why do engineers confuse the two graphs so consistently?
    Because both are boxes and arrows, both are acyclic, both get called the same three-letter word, and the arrow is drawn identically in each. The word describes the shape of the graph, not the semantics of an edge - and the semantics of an edge is the entire difference.

A relay race hands a baton along: the runners are connected by the object being passed, and if the handoff fails nothing arrives. A day's calendar is also an ordered chain of boxes, but nothing physical moves from the nine o'clock meeting to the ten o'clock one - if notes must reach the second meeting, a person has to write them down somewhere and carry them. The scheduler's graph is the calendar; the engine's graph is the relay.

saying these in an interview costs you the question

  • Says both graphs are the same thing drawn at different zoom levels.
  • Assumes the scheduler hands one step's output to the next step automatically.
  • Believes data moves between scheduled steps because the picture has arrows.
  • Calls both pictures the same graph without saying which edges carry records.
  • Thinks the difference is size: big graphs go to the scheduler, small ones to the engine.
  • Claims an engine never writes an intermediate between two steps.
open as a page

Two steps a workflow scheduler starts in order pass a 400 GB intermediate through object storage - what does that boundary cost?

level: middleimportance: should knowfreq 58%

basics

~20 s

A full write and a full read of 400 GB, plus a second program's start-up and a path someone must name, encode and eventually delete. An edge between two steps a scheduler starts carries nothing, so whatever crosses it has to be materialised.

open as a page

A workflow scheduler is given one step per input file and there are 40,000 files - what breaks, and where does that fan-out belong?

level: middleimportance: should knowfreq 52%

basics

~20 s

Fixed per-step costs dominate: 40,000 program launches and 40,000 rows of scheduler bookkeeping for a few seconds of real work each, throttled by whatever concurrency limit the scheduler enforces. Fan-out over data belongs inside one submission, as parallelism.

open as a page

Your platform's default is one submission per pipeline - which boundaries should become scheduler edges instead, and what does each one cost?

level: principalimportance: should knowfreq 42%

basics

~20 s

Promote a boundary only when something other than the next step needs it: a different program or runtime, an intermediate real consumers read, an external wait, or two halves wanting very different capacity. Each promotion costs a materialised named artefact, a second start-up, and a second picture to read.

open as a page