skip to content

Environments, Gates & Approvals

Modelling dev/staging/prod as first-class objects with protection rules, and deciding which gates are automated and which need a human. Interviewers ask because approvals are where compliance requirements meet delivery speed and one of the two usually loses.

on this pageshow

questions

5

In a CI/CD system, what does it mean to model a deployment target such as staging or production as a first-class environment object rather than just a set of variables?

level: middleimportance: must knowfreq 60%

answer

  1. a named target, not a string
  2. credentials scoped to the target
  3. rules checked before the job starts
  4. history answers what is live now
  5. one deployment at a time

basics

~20 s

A first-class environment is a named deployment target the delivery platform owns: it carries its own scoped credentials, protection rules that decide who and what may deploy to it, and a recorded deployment history that serves as the audit trail.

solid answer

~50 s

In the weak version, an environment is just a string plus a block of variables: a job runs a deploy script with the argument `production`, and the pipeline file is the only thing that knows what production means. Modelling it as a first-class object moves that into the platform. The environment becomes a named record with four things attached: configuration and credentials scoped so only jobs targeting it can read them; protection rules such as required approvers, allowed source branches or tags, and wait timers, evaluated *before* the job starts; a deployment history that answers "what version is live, from which commit, approved by whom, deployed when"; and a serialization rule so two deploys to the same target cannot race. The payoff is that authorization and evidence stop being a convention inside a script and become something the platform enforces and records.

go deeper

for a junior

Know that dev, staging and production are separate targets with separate configuration, and that a deploy to production normally needs more permission than a deploy to staging. Be able to name what distinguishes them.

for a middle

Explain the mechanics: secrets scoped to the environment, protection rules evaluated before the job starts, and a deployment history recording artifact, actor and approver. Contrast that with an if-check inside a deploy script.

for a senior

Show the operational payoff — answering what is live right now, rolling back to a recorded artifact, and serializing concurrent deploys so an older build cannot win by finishing last. Point out that the record is only trustworthy if nothing can deploy outside the pipeline.

for a principal

Own the model across many teams: how many long-lived tiers actually earn their cost, how ephemeral per-change environments are templated and torn down, and how environment records become the compliance evidence instead of a parallel manual process.

## Two ways to represent a deployment target Almost every pipeline starts with the environment as a string. A job runs `./deploy.sh production`, and "production" is an argument plus a block of variables — a hostname, an API key, a cluster name. Nothing in the delivery platform distinguishes the job that deploys to a scratch sandbox from the job that deploys to the system customers use. Anyone who can edit the pipeline file, or who can persuade a job to run with the right variables in scope, can reach production. Modelling the environment as a **first-class object** moves that knowledge out of the script and into the platform. The environment becomes a named record the platform owns, and four capabilities hang off it. ## 1. Scoped configuration and credentials Secrets attach to the environment rather than to the repository or the pipeline as a whole. A job only receives production's deploy credential if the job declares that it is deploying to production — and, crucially, only after that environment's rules have passed. A pull-request job, a nightly lint job, or a staging deploy never has the value in its process at all. This is a containment property, not a secrecy property: the credential is not merely hidden from other jobs, it is never injected into them, so a compromised or malicious job in the same repository has nothing to steal. ## 2. Protection rules evaluated before the job starts The rules a platform typically supports are: required human approvers (and often, who may *not* approve — the author of the change); allowed sources, so only the default branch or a signed tag may deploy; a wait timer that forces a soak or a cooling-off period; and required upstream checks. The important word is *before*. A guard written as an `if` inside the deploy script runs on a machine that has already been allocated, with the repository already checked out and the secrets already injected. A platform-level protection rule holds the run in a pending state where nothing has started. The difference matters both for cost (no agent is tied up during a two-hour approval wait) and for security (the failure mode of the script guard is "the credential is on a box that decided not to use it"). ## 3. Deployment history and current state An environment object records deployments as events: which artifact, which commit, which run, which identity performed it, who approved it, when, and whether it succeeded. That single table answers the two questions teams cannot otherwise answer under pressure — *what is running in production right now?* and *who authorized it?* It is also what makes "roll back" a concrete operation: the previous successful record names an exact artifact to re-deploy, rather than someone reconstructing it from memory. This history is the audit trail auditors actually want. It is generated as a by-product of deploying, which is why it is trustworthy; evidence assembled afterwards from screenshots and chat messages is not. ## 4. Serialization and lifecycle Because the environment is an object, the platform can enforce one deployment at a time against it. Without that, two runs overlap and the *older* artifact can win simply by finishing last. Newer runs are typically queued or supersede the pending one. The object model also has to cope with two very different shapes: a small number of long-lived environments (dev, staging, production, plus per-region production instances), and a potentially unbounded number of ephemeral ones — one per pull request, created on open and destroyed on merge or after an idle timeout. Ephemeral environments are cheap to create precisely because the environment is a record with a lifecycle rather than hand-maintained infrastructure. ## What it does not give you An environment object is metadata about a target; it is not the target. The cluster, the database and the DNS entry still have to exist and are usually managed separately as infrastructure. It is also not a deployment strategy — how traffic moves onto the new version is a different concern entirely. And it is easy to have the form without the substance. An environment with no protection rules, whose secrets are duplicated at repository scope anyway, and which every branch may deploy to, is a variable group with a nicer name. The value comes from the rules and the record, not from the label.

  • Why is a protection rule enforced by the platform stronger than the same check written at the top of the deploy script?
    By the time a script check runs, the platform has already allocated an agent, checked out the repository and injected the environment's credentials. The guard is being enforced by the very job it is supposed to restrain. A platform rule holds the run pending: nothing starts and no credential is issued until the rule passes.
  • Your environment object records every deployment, but a service was updated without any record appearing. What does that tell you?
    That something deployed outside the delivery path — a human or a script holding credentials to the target directly. The record is only complete if the pipeline's identity is the *only* principal with write access to that environment. Fixing the gap means removing standing human credentials, not adding another gate to the pipeline.
  • How would you model dozens of short-lived per-pull-request environments without drowning in objects?
    Treat them as a class rather than as individually configured targets: one template that names the shared credentials, a naming scheme derived from the pull request, and an automatic teardown on merge, close, or idle timeout. Their protection rules are deliberately weak because they hold no production data — that is the point of separating them from the long-lived tier.

saying these in an interview costs you the question

  • Says an environment is just a bag of variables
  • Thinks protection rules are checked inside the deploy script
  • Sees no problem sharing production credentials with staging jobs
  • Cannot say what version production is running or who deployed it
  • Treats an approval as a conversation rather than a recorded event

context

open as a page

In a deployment pipeline, what is a manual approval gate, and what is the pipeline actually doing while it waits for a human to approve?

level: juniorimportance: should knowfreq 58%

basics

~20 s

A manual approval gate pauses a run at a boundary before a protected stage until a permitted person approves it. A platform-level gate holds the run as pending state with no agent allocated, and expires after a configured timeout.

open as a page

A pipeline blocks merges whenever total test coverage falls below 80 percent. What makes a threshold gate like that useful or useless, and what would you gate on instead?

level: middleimportance: should knowfreq 48%

basics

~20 s

A whole-repository coverage threshold is dominated by legacy code, so one change barely moves it and the gate rarely fires on the change that deserves it. Gate on coverage of the lines the change touched, or ratchet the number so it can never fall.

open as a page

Your compliance regime requires a documented human approval before every production change, and the team wants to deploy twenty times a day. How would you satisfy both?

level: principalimportance: should knowfreq 40%

basics

~20 s

Separate the control objective from its implementation. The requirement is that someone other than the author authorized the change and that evidence survives; it is not that a person clicks a button per deploy. Peer review bound to the deployed artifact usually satisfies it.

open as a page

A change reached production without passing the pipeline's required security-scan gate. How do you work out how it got there, and how would you design gates so that route is closed?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Start from the target's deployment record and identify which principal performed the deploy. Either it was the pipeline, and the gate did not apply or did not fail closed, or it was some other identity deploying outside the pipeline entirely — which no in-pipeline gate could ever have stopped.

open as a page