CI pipelines are usually described as stages that contain jobs that contain steps. What is each of those three units actually responsible for?
answer
- three different kinds of unit
- ordering, allocation, execution
- only one of them owns a machine
- steps share a filesystem; jobs do not
- a stage is only a barrier
basics
~20 sA stage is a unit of ordering — a named phase that acts as a barrier before the next one starts. A job is a unit of allocation — one machine and one workspace, and the thing that runs in parallel. A step is a unit of execution — one command inside a job.
solid answer
~50 sThey are three different kinds of unit, not three sizes of the same thing. A **stage** orders work: it groups jobs that may run together and says nothing in the next stage starts until everything in this one finishes. It owns no machine and no filesystem — it is purely a barrier. A **job** is what gets scheduled onto a runner: it gets a machine or container, a fresh working directory, and it is the unit of parallelism, of isolation and usually of retry and billing. A **step** is one command or one packaged action inside a job; steps run in order, share the job's working directory and environment, and a failing step normally ends the job. The practical consequence is that anything one job wants from another must be published explicitly — as an artifact, an output value, or an image in a registry — because a job's filesystem does not survive it.
go deeper
Be able to say plainly that a step is one command, a job is one machine running a list of steps, and a stage groups jobs that run before the next group starts. Knowing that jobs get a clean workspace is the part interviewers listen for.
Explain the mechanics: why state does not cross a job boundary, what an artifact or output is for, why a failing step ends a job, and what fixed setup cost each extra job pays.
Show you have tuned a real pipeline — how you decided what deserves its own job, where artifact hand-off cost outweighed the parallelism, and how job granularity affects retries and the bill.
Own the tradeoff at scale: the granularity conventions you set across many repositories, how per-job setup cost multiplies across hundreds of runs a day, and when you deliberately accept a slower pipeline for a simpler one.
## Why the vocabulary matters Every mainstream CI platform describes a run with roughly these three nested words, and candidates often treat them as "big, medium, small" containers of the same substance. They are not. Each answers a different question: *when may this run?* (stage), *where does it run?* (job), *what actually executes?* (step). Almost every avoidable pipeline slowness comes from putting work at the wrong level. ## A stage is a unit of ordering A stage — some platforms call it a phase or a section — is a label attached to a set of jobs, plus one rule: nothing in the next stage begins until *everything* in this stage has finished. That trailing rule is the entire content of the concept. A stage owns no machine, no filesystem, no process and no environment variables. If you deleted the stage labels and instead recorded, for each job, which other jobs it depends on, you would lose nothing except the barrier semantics. That barrier is sometimes exactly what you want ("do not deploy until every test job has passed") and sometimes pure waste (a two-minute integration job waiting on an unrelated eight-minute browser job that it does not depend on). ## A job is a unit of allocation A job is what the scheduler hands to a runner. It gets a machine, a VM or a container; it starts from a clean workspace; it checks out the source itself; and it disappears when it ends. This makes the job simultaneously: - **the parallelism unit** — jobs are what run at the same time on different runners; - **the isolation unit** — one job cannot see another's filesystem, processes or shell state; - **the transfer unit** — anything crossing a job boundary must be uploaded as an artifact, emitted as a declared output value, or pushed to a registry or cache; - **the retry and accounting unit** — you generally re-run a failed job, and you are generally billed per job-minute. ## A step is a unit of execution A step is one command, script or packaged action. Steps inside a job run sequentially, in the declared order, in the same working directory, and they inherit the environment the job set up. Installing dependencies in step two and using them in step five works precisely because the two steps share a machine. By default the first failing step ends the job, and later steps are skipped unless they are explicitly marked to run regardless — which is why cleanup logic written as an ordinary trailing step often never executes. ``` checkout ─► install deps ─► lint ─► unit tests ─► upload report (one job, 5 steps, one machine) ``` ## The two symmetrical mistakes **Everything in one job.** Lint, unit tests, integration tests and the image build all become steps of a single job. The result is a fully serial run whose wall-clock time is the sum of every part, no ability to use more than one runner, and a retry that redoes all of it to re-test the one thing that failed. **Every step as its own job.** Each job now pays the fixed setup cost again — provisioning a runner, pulling a container image, checking out the repository, restoring dependencies — and each hand-off becomes an artifact upload followed by a download. For a thirty-second task with a ninety-second setup, splitting it out makes the pipeline slower *and* more expensive while looking more parallel on the graph. The usable rule: keep work in one job when it shares a working directory and is short; split it into separate jobs when the parts are genuinely independent and each is long enough to repay its own setup cost. ``` ┌─ unit tests ─┐ checkout+build ─┤ ├─ package └─ lint ───────┘ (3 jobs in parallel-capable positions; each pays its own setup) ``` ## Where the boundary bites in practice - **State does not travel.** A file written in the build job is invisible in the test job unless it was published. Teams that learn CI by writing one long job hit this the first time they split. - **Failure granularity follows jobs.** If you want to re-run only the flaky browser suite, it has to be its own job. - **Cost follows jobs too.** Ten jobs of one minute usually cost more than one job of ten minutes, because you pay the setup ten times. - **A stage is not an environment.** "Staging" as a deployment target and "stage" as an ordering phase are unrelated ideas that unfortunately share a word.
- If a job's filesystem does not survive it, how does a later job get the binary an earlier job built?It has to be published deliberately: uploaded as a build artifact the later job downloads, pushed to a container or package registry and pulled by digest, or — for small values like a version string — emitted as a declared job output. Nothing crosses a job boundary implicitly, and that explicit hand-off is what lets you deploy the exact bits you tested.
- Why does a cleanup command written as the last step of a job often fail to run?Because steps are sequential and the first failure normally aborts the job, so a trailing step is skipped in exactly the case you wrote it for. Cleanup has to be marked to run regardless of outcome — every platform has a mechanism for that — or moved into a separate job that depends on the first and runs whatever its result.
- When is splitting work into more jobs actually the wrong optimisation?When each part is short relative to the fixed per-job setup — runner provisioning, image pull, checkout, dependency restore — or when the parts must exchange large files, so you trade compute time for artifact upload and download. You also pay more, since billing is per job-minute. Split when the pieces are independent and long; keep them together when they are chatty and quick.
saying these in an interview costs you the question
- Says a job and a step are basically the same thing
- Assumes files written in one job are visible in the next
- Thinks steps inside a job run in parallel
- Believes more jobs is always faster
- Confuses a pipeline stage with a deployment environment