skip to content

Pipeline Anatomy & Triggers

How a pipeline is decomposed into stages, jobs and steps, and what causes one to start. Interviewers open with this because the shape of the dependency graph — what fans out, what fans back in — decides both feedback speed and how much you pay per merge.

on this pageshow

questions

5

CI pipelines are usually described as stages that contain jobs that contain steps. What is each of those three units actually responsible for?

level: juniorimportance: must knowfreq 78%

answer

  1. three different kinds of unit
  2. ordering, allocation, execution
  3. only one of them owns a machine
  4. steps share a filesystem; jobs do not
  5. a stage is only a barrier

basics

~20 s

A stage is a unit of ordering — a named phase that acts as a barrier before the next one starts. A job is a unit of allocation — one machine and one workspace, and the thing that runs in parallel. A step is a unit of execution — one command inside a job.

solid answer

~50 s

They are three different kinds of unit, not three sizes of the same thing. A **stage** orders work: it groups jobs that may run together and says nothing in the next stage starts until everything in this one finishes. It owns no machine and no filesystem — it is purely a barrier. A **job** is what gets scheduled onto a runner: it gets a machine or container, a fresh working directory, and it is the unit of parallelism, of isolation and usually of retry and billing. A **step** is one command or one packaged action inside a job; steps run in order, share the job's working directory and environment, and a failing step normally ends the job. The practical consequence is that anything one job wants from another must be published explicitly — as an artifact, an output value, or an image in a registry — because a job's filesystem does not survive it.

go deeper

for a junior

Be able to say plainly that a step is one command, a job is one machine running a list of steps, and a stage groups jobs that run before the next group starts. Knowing that jobs get a clean workspace is the part interviewers listen for.

for a middle

Explain the mechanics: why state does not cross a job boundary, what an artifact or output is for, why a failing step ends a job, and what fixed setup cost each extra job pays.

for a senior

Show you have tuned a real pipeline — how you decided what deserves its own job, where artifact hand-off cost outweighed the parallelism, and how job granularity affects retries and the bill.

for a principal

Own the tradeoff at scale: the granularity conventions you set across many repositories, how per-job setup cost multiplies across hundreds of runs a day, and when you deliberately accept a slower pipeline for a simpler one.

## Why the vocabulary matters Every mainstream CI platform describes a run with roughly these three nested words, and candidates often treat them as "big, medium, small" containers of the same substance. They are not. Each answers a different question: *when may this run?* (stage), *where does it run?* (job), *what actually executes?* (step). Almost every avoidable pipeline slowness comes from putting work at the wrong level. ## A stage is a unit of ordering A stage — some platforms call it a phase or a section — is a label attached to a set of jobs, plus one rule: nothing in the next stage begins until *everything* in this stage has finished. That trailing rule is the entire content of the concept. A stage owns no machine, no filesystem, no process and no environment variables. If you deleted the stage labels and instead recorded, for each job, which other jobs it depends on, you would lose nothing except the barrier semantics. That barrier is sometimes exactly what you want ("do not deploy until every test job has passed") and sometimes pure waste (a two-minute integration job waiting on an unrelated eight-minute browser job that it does not depend on). ## A job is a unit of allocation A job is what the scheduler hands to a runner. It gets a machine, a VM or a container; it starts from a clean workspace; it checks out the source itself; and it disappears when it ends. This makes the job simultaneously: - **the parallelism unit** — jobs are what run at the same time on different runners; - **the isolation unit** — one job cannot see another's filesystem, processes or shell state; - **the transfer unit** — anything crossing a job boundary must be uploaded as an artifact, emitted as a declared output value, or pushed to a registry or cache; - **the retry and accounting unit** — you generally re-run a failed job, and you are generally billed per job-minute. ## A step is a unit of execution A step is one command, script or packaged action. Steps inside a job run sequentially, in the declared order, in the same working directory, and they inherit the environment the job set up. Installing dependencies in step two and using them in step five works precisely because the two steps share a machine. By default the first failing step ends the job, and later steps are skipped unless they are explicitly marked to run regardless — which is why cleanup logic written as an ordinary trailing step often never executes. ``` checkout ─► install deps ─► lint ─► unit tests ─► upload report (one job, 5 steps, one machine) ``` ## The two symmetrical mistakes **Everything in one job.** Lint, unit tests, integration tests and the image build all become steps of a single job. The result is a fully serial run whose wall-clock time is the sum of every part, no ability to use more than one runner, and a retry that redoes all of it to re-test the one thing that failed. **Every step as its own job.** Each job now pays the fixed setup cost again — provisioning a runner, pulling a container image, checking out the repository, restoring dependencies — and each hand-off becomes an artifact upload followed by a download. For a thirty-second task with a ninety-second setup, splitting it out makes the pipeline slower *and* more expensive while looking more parallel on the graph. The usable rule: keep work in one job when it shares a working directory and is short; split it into separate jobs when the parts are genuinely independent and each is long enough to repay its own setup cost. ``` ┌─ unit tests ─┐ checkout+build ─┤ ├─ package └─ lint ───────┘ (3 jobs in parallel-capable positions; each pays its own setup) ``` ## Where the boundary bites in practice - **State does not travel.** A file written in the build job is invisible in the test job unless it was published. Teams that learn CI by writing one long job hit this the first time they split. - **Failure granularity follows jobs.** If you want to re-run only the flaky browser suite, it has to be its own job. - **Cost follows jobs too.** Ten jobs of one minute usually cost more than one job of ten minutes, because you pay the setup ten times. - **A stage is not an environment.** "Staging" as a deployment target and "stage" as an ordering phase are unrelated ideas that unfortunately share a word.

  • If a job's filesystem does not survive it, how does a later job get the binary an earlier job built?
    It has to be published deliberately: uploaded as a build artifact the later job downloads, pushed to a container or package registry and pulled by digest, or — for small values like a version string — emitted as a declared job output. Nothing crosses a job boundary implicitly, and that explicit hand-off is what lets you deploy the exact bits you tested.
  • Why does a cleanup command written as the last step of a job often fail to run?
    Because steps are sequential and the first failure normally aborts the job, so a trailing step is skipped in exactly the case you wrote it for. Cleanup has to be marked to run regardless of outcome — every platform has a mechanism for that — or moved into a separate job that depends on the first and runs whatever its result.
  • When is splitting work into more jobs actually the wrong optimisation?
    When each part is short relative to the fixed per-job setup — runner provisioning, image pull, checkout, dependency restore — or when the parts must exchange large files, so you trade compute time for artifact upload and download. You also pay more, since billing is per job-minute. Split when the pieces are independent and long; keep them together when they are chatty and quick.

saying these in an interview costs you the question

  • Says a job and a step are basically the same thing
  • Assumes files written in one job are visible in the next
  • Thinks steps inside a job run in parallel
  • Believes more jobs is always faster
  • Confuses a pipeline stage with a deployment environment

context

open as a page

What are the common ways a CI/CD pipeline run gets started, and how do they differ in what they cost and what feedback they give?

level: middleimportance: must knowfreq 68%

basics

~20 s

Runs start from a branch push, a pull/merge request, a tag, a schedule, a manual button, an upstream pipeline, or an external API call. Each differs in how often it fires, who is waiting for the result, and how much compute it burns per merge.

open as a page

In CI, why can a pipeline whose phases run as strict sequential barriers finish later than the same work expressed as a dependency graph between jobs?

level: middleimportance: should knowfreq 58%

basics

~20 s

A barrier makes every job in a phase wait for the slowest job in the previous phase, so total time is the sum of those slowest jobs. A dependency graph lets each job start as soon as its own inputs are ready, so total time is the longest actual chain.

open as a page

Your verification suite has grown to 40 minutes and currently runs in full on every push. How would you decide which work runs on which trigger?

level: principalimportance: should knowfreq 44%

basics

~20 s

Tier the work by who is waiting for it: fast checks on every proposal update, the full gate where changes land, and slow or drift-detecting suites on a schedule. Each move trades earlier feedback for compute, or compute for later discovery of a failure.

open as a page

A busy branch receives several pushes a minute and the CI queue keeps growing. What does grouping runs under a concurrency key and cancelling superseded runs buy you, and where is that dangerous?

level: seniorimportance: nice to knowfreq 42%

basics

~20 s

Grouping runs by a key such as the branch or proposal, and cancelling the older in-flight run when a newer one arrives, keeps at most one active run per key. It cuts queue depth and spend, since only the newest commit's result is usually wanted — but cancelling deploys is unsafe.

open as a page