In CI, why can a pipeline whose phases run as strict sequential barriers finish later than the same work expressed as a dependency graph between jobs?
answer
- sum of slowest per phase
- versus the longest path
- waiting on work you never consume
- fan-out from build, fan-in at the gate
- only pays off with free runner capacity
basics
~20 sA barrier makes every job in a phase wait for the slowest job in the previous phase, so total time is the sum of those slowest jobs. A dependency graph lets each job start as soon as its own inputs are ready, so total time is the longest actual chain.
solid answer
~50 sWith strict phase barriers, wall-clock time is the sum, over phases, of the slowest job in each phase — even when a job in phase three depends on only one job from phase two. With a dependency graph, each job starts the moment its own predecessors finish, so the run costs the longest *path* through the graph, the critical path. The gap is exactly the idle time jobs spend waiting on work they never needed. A graph also lets you fan out from one build into many independent test jobs and fan back in to a single gate that requires all of them. The tradeoffs are real, though: a graph is harder to read and reason about, every hand-off must be an explicit artifact or output, and the speedup only materialises if you have runner capacity free to actually run those jobs concurrently. Barriers are still the right choice where the ordering is genuinely total, such as never deploying before all tests have passed.
go deeper
Know that a phase barrier means everything in the next phase waits for the slowest job in the current one, and that a dependency graph lets a job start as soon as its own inputs are ready.
Be able to do the arithmetic out loud: sum of per-phase maxima versus the critical path, with a small worked example, and name what a job was idling on.
Show you have diagnosed a real pipeline — reading per-job queue and run times, separating barrier waste from runner starvation, and deciding where a deliberate fan-in gate belongs.
Own the tradeoff between speed and legibility across many teams: when a hand-maintained graph becomes a liability, how you keep declared dependencies honest, and what wall-clock target actually justifies the complexity.
## The two models **Phase barriers.** Jobs are grouped into named phases and the platform enforces one rule: no job in phase *n+1* starts until every job in phase *n* has finished. Ordering is expressed by which bucket a job sits in. **Dependency graph (DAG).** Each job names the jobs it needs. A job becomes runnable the instant its own predecessors succeed, regardless of what else is still running. ## The arithmetic With barriers, the run takes ``` T_barrier = Σ over phases of ( max job duration in that phase ) ``` With a graph, it takes ``` T_dag = max over all paths of ( sum of job durations along that path ) ← the critical path ``` A worked example. Phase 1: lint 1 min, unit tests 9 min, build 4 min. Phase 2: integration tests 6 min (needs only the build), publish docs 2 min (needs only lint). - Barriers: phase 1 costs 9 (the unit tests), phase 2 costs 6 → **15 minutes**. - Graph: the critical path is build → integration = 4 + 6 = 10; unit tests run alongside and finish at 9 → **10 minutes**. The five lost minutes are integration tests idling on unit tests they do not consume. Notice the pathology: adding one slow job to an early phase penalises *everything* downstream of the barrier, whether or not anything depends on it. Teams feel this as "the pipeline got slower when we added a job that runs in parallel". ## Fan-out and fan-in A graph makes two shapes explicit. **Fan-out**: one job (usually the build) is followed by many independent consumers — unit, integration, contract, browser, security scan — each depending only on the build. This is where a graph earns most of its wall-clock savings. **Fan-in (join)**: a single job depends on many predecessors, and it starts only when the slowest of them finishes. A fan-in job is the natural place for the deploy or the "all checks passed" gate — and it is a barrier, deliberately, in one specific spot rather than across the whole pipeline. ``` ┌─ unit ────────┐ build ──────┼─ integration ─┼──► deploy (fan-out then fan-in) └─ browser ─────┘ ``` ## What the graph costs you 1. **Concurrency you may not have.** The graph's advantage is theoretical if your runner pool can only execute two jobs at once; the extra runnable jobs simply queue. Measure queue wait before blaming the topology. 2. **Explicit hand-offs.** Because jobs are isolated, every edge in the graph corresponds to an artifact upload plus download, an output value, or a registry push. Fan-out from a build means the build's output is fetched N times — sometimes enough to eat the savings. 3. **Comprehensibility.** A ten-node graph with hand-written edges is genuinely harder to review than four ordered phases, and a missing edge is a real bug: a job that reads something it never declared a dependency on will pass by luck and fail intermittently. 4. **Skip and failure propagation.** When an upstream job fails or is skipped, its dependents are normally skipped too, and a fan-in gate has to be written so that a *skipped* prerequisite does not read as success. This is where "green pipeline that ran nothing" bugs come from. 5. **Cost is unchanged.** The graph reduces wall-clock, not job-minutes. You pay for the same compute, just compressed. ## When barriers are the right answer - The ordering really is total: build → test → deploy, where deploy must never start early. - You want cheap fail-fast gating: putting a 20-second lint job alone in the first phase deliberately blocks the expensive phase behind it and saves spend on obviously broken commits. Here the barrier is the feature. - The pipeline is small, and legibility beats two minutes. The mature answer is usually a hybrid: a fast gating phase up front, a wide fan-out in the middle expressed as real dependencies, and one fan-in gate before anything is deployed. ## Diagnosing it Do not guess at the topology. Take one representative run, list each job's queued time, start time and duration, and find the job that ends last — then walk backwards through what it waited for. If what it waited for is not what it consumes, you have found barrier waste. If the waiting was queue time, you have a capacity problem and rewiring the graph will change nothing.
- You rewired the pipeline into a dependency graph and the wall-clock time barely moved. What would you check first?Runner capacity. A graph only helps if the newly-runnable jobs can actually start; if the pool is saturated, they queue and the critical path is set by queue wait rather than dependencies. Compare each job's queued time against its run time across a few runs. The second suspect is artifact hand-off: fanning out from a build means every consumer downloads it, which can consume the savings.
- What breaks when a job reads a file produced by another job it never declared a dependency on?It works by accident whenever the scheduler happens to order them that way, and fails intermittently when it does not — the classic "flaky on a busy day" pipeline. Because jobs are isolated, the missing edge also means the file may simply not be downloaded. Declare every input; treat an undeclared read as a bug even while it passes.
saying these in an interview costs you the question
- Says a DAG reduces the total compute cost
- Assumes jobs run in parallel regardless of runner capacity
- Thinks a barrier only delays the job that is slow
- Believes stage ordering is required for correctness everywhere
- Ignores that every graph edge needs an explicit artifact hand-off