skip to content

How do you draw a build-and-deploy pipeline as a data-flow diagram, and where do its trust boundaries go?

level: middleimportance: must knowfreq 66%

answer

  1. the pipeline is the system, not scenery
  2. actors, jobs, stores, flows
  3. trust changes at contribution and at deploy
  4. the runner's disk is a data store
  5. every path ends at the deploy identity

basics

~20 s

Treat the pipeline as the system: external entities (contributors, approvers, upstream registries), processes (trigger, build job, deploy job), stores (source repo, secret store, artifacts, runner disk). Put trust boundaries where trust actually changes — untrusted contribution entering the build, and the build reaching production credentials.

solid answer

~50 s

I make the delivery path the system under analysis instead of one box labelled CI. External entities are the people and services I do not control: contributors (anonymous ones on a public repo), reviewers and approvers, upstream package sources. Processes are the trigger that decides what runs, the build job and the deploy job — kept separate because they hold different privileges. Data stores are the source repository, the secret store, the artifact store, the build cache and the runner's own workspace. Flows are the change event, the checkout, the secret read, dependency pulls, the artifact push, the credentialed call into production and the logs coming back out. Boundaries go at the contribution edge, at the secret read, between jobs sharing a runner, and at the deploy call. Elevation of privilege dominates, because every path ends at an identity that can change production.

go deeper

for a junior

Be ready to name the parts of a delivery path and say which are outside your control: contributors, the source repository, the build job, the secret store, the deploy step. Knowing that a trust boundary marks a change in trust, not a change in component, is the recall you need.

for a middle

Explain the mechanics: which elements are processes versus stores, why the build job and deploy job are drawn separately, and what actually crosses each boundary. An interviewer expects you to place the contribution boundary and the secret boundary without prompting.

for a senior

Demonstrate that you have looked at a real pipeline and found the boundary that was not there — a shared runner, a workflow file editable in the same change, a credential broader than the job needs. Show the shortest path from attacker-writable content to the deploy identity.

for a principal

Own the argument that the delivery path deserves a model of its own with a named owner, and decide how much of it belongs inside each application's model versus one platform model. The tradeoff is duplication across teams against a central model nobody keeps current.

## Why the pipeline deserves its own diagram Most application threat models draw the delivery path as a cloud on the edge of the page, or leave it off entirely. That hides the uncomfortable fact that the highest-privilege actor in the system is a program that runs whatever a merged change tells it to run, on a schedule nobody watches, holding a credential that can rewrite production. Modeling it means turning the notation on the delivery path itself: same elements, same questions, different subject. ## The elements **External entities** — anything outside your control, drawn outside every boundary: - contributors, including anonymous ones if the repository is public - reviewers and release approvers (people are external entities; they are not part of your process) - upstream package sources and base images the build pulls from - the operator of the CI service itself, if it is managed rather than self-run **Processes** — keep them separate, because they carry different privilege: - the trigger/orchestrator that decides which job runs on which event - the build and test jobs - the deploy job **Data stores:** - the source repository, including branch-protection state and the pipeline definition that lives inside it - the secret store or identity provider the jobs read from - the artifact store and the build cache - the runner's own workspace and disk — the store people forget **Data flows:** the change event into the trigger; the checkout into the job; the secret read; dependency downloads; the artifact push; the credentialed call into production; logs and job output flowing outward. ## Where the boundaries actually go Four lines earn their place. A boundary marks a change in trust or privilege, not a change in component. 1. **The contribution boundary** — between a proposed change nobody has vetted and the thing that decides to build it. Everything a contributor authors crosses here: application source, tests, build scripts, dependency manifests and lockfiles, and often the pipeline definition itself. The question to ask at this line is: does content from the untrusted side get to choose what the trusted side executes? 2. **The secret boundary** — between job execution and the credential store. Crossing it means code running in that job can read the credential, whatever the code turns out to be. 3. **The runner boundary** — between one job and the next on the same executor. On an ephemeral runner this is real. On a long-lived shared host it is imaginary: a persistent workspace, shared caches and the host's own identity put every team's jobs in one trust zone. 4. **The deploy boundary** — pipeline into production. In most organisations this is the only write path into production with no human in the request path. ## What per-element threat enumeration yields Walking STRIDE over the elements gives a first pass: | Element | Categories that dominate | | --- | --- | | Contributor (external entity) | Spoofing (authentication) | | Trigger / orchestrator (process) | Tampering, Elevation of Privilege | | Build job (process) | Tampering with the artifact, Information Disclosure of secrets, Elevation of Privilege | | Secret store (data store) | Information Disclosure | | Source repository (data store) | Tampering, Repudiation | | Deploy flow (data flow) | Tampering, Elevation of Privilege | Elevation of privilege — the category whose violated property is authorization — dominates the whole model, because every interesting path ends at an identity that can change production. The single most productive question in a pipeline model is: what is the shortest path from something an attacker can write to something the deploy credential executes? ## The mistakes that make the model useless - **Drawing the pipeline as one opaque box.** You then cannot see the secret read or the deploy flow inside it, which is where the threats live. - **Putting developers inside the trust boundary.** On a public repository the contributor is anonymous. Even internally, "employee" is not the same trust level as "the identity that can deploy", and collapsing the two is how a code review becomes an unwitting approval of a production change. - **Forgetting the runner host.** Its disk is a data store shared across jobs unless something actively destroys it. - **Modeling the artifact and not the credential.** The artifact is one asset; the standing deploy identity is usually the bigger one, because it outlives the build and can be used again. - **Treating the pipeline definition as configuration.** It is code that determines what runs, it lives in the repository, and a contributor may be able to change it in the same change you are reviewing for application logic. ## Scope The model stops where the pipeline stops. What happens inside the platform it deploys into, and the catalogue of published supply-chain attacks and their controls, are separate subjects with separate models; here the delivery path is the system, and the deploy credential is the prize.

  • Which element of a pipeline diagram do people most often leave off entirely?
    The runner host and its workspace. Everyone draws the jobs and the repository but treats the executor as scenery, which hides shared caches, leftover files from the previous job, and the host's own identity. Once you draw it as a data store with its own lifetime, the question of what survives between jobs becomes impossible to skip.
  • Is the pipeline definition file a process or a data store in your model?
    Both, and that dual role is the point. It sits in the repository as a store any contributor can propose changes to, and it determines what the trusted process executes. Modeling it only as configuration hides the fact that editing it is equivalent to editing the behaviour of the highest-privileged process in the system.
  • How does the diagram change if runners are ephemeral rather than long-lived?
    The line between consecutive jobs becomes a real trust boundary rather than a wish. With a fresh executor per job, the persistent workspace and cache disappear as shared stores, and one job's leftovers can no longer be read or executed by the next. The host identity question remains: the runner still authenticates to something.

It is the difference between diagramming a bank's vault and diagramming the armoured van that visits it nightly with its own key.

saying these in an interview costs you the question

  • Draws CI as one box outside the trust boundary
  • Puts all developers inside the trusted zone
  • Models the artifact but never the deploy credential
  • Treats the pipeline definition as config, not attacker-writable code
  • Assumes every job starts on a clean isolated machine

context