skip to content

A change reached production without passing the pipeline's required security-scan gate. How do you work out how it got there, and how would you design gates so that route is closed?

level: seniorimportance: nice to knowfreq 34%

answer

  1. ask which identity deployed
  2. no run, no pipeline involvement
  3. the gate may not cover this trigger
  4. the change can edit its own gate
  5. break-glass should be loud, not silent

basics

~20 s

Start from the target's deployment record and identify which principal performed the deploy. Either it was the pipeline, and the gate did not apply or did not fail closed, or it was some other identity deploying outside the pipeline entirely — which no in-pipeline gate could ever have stopped.

solid answer

~50 s

Work backwards from the environment's deployment record, not from the pipeline. It names the identity that performed the deploy and, ideally, the run that produced it. If no run maps to it, someone deployed with their own credentials and the gate was never in the path — the fix is removing standing human access to the target, not adding a check. If a run does map to it, ask why the gate did not stop it: the gate may only be wired to one trigger or branch so the hotfix path skips it; the step may continue on error, or lose its exit status through a shell pipe; the gate may be defined in the repository's own pipeline file, which the same change is free to edit; or a break-glass approval was used and looks indistinguishable from a normal one. Gates belong in platform policy the change cannot edit, must fail closed, and the emergency path should be loud and separately recorded.

go deeper

for a junior

Understand that a gate only constrains changes that travel through the pipeline, and that a person with direct credentials to the target can deploy without touching it at all.

for a middle

Trace the mechanics: which trigger or branch the gate was attached to, whether the step failed open, and whether the deployed artifact is the same one the gate examined.

for a senior

Run the investigation from the target's deployment record outward, name the specific bypass classes, and land the structural fix — pipeline-only write access, policy enforced outside the repository, gates that fail closed.

for a principal

Treat the bypass rate as the real signal: instrument break-glass usage, decide organisation-wide where gate policy is defined and who may change it, and fix gates that people route around rather than adding controls on top.

## Start at the target, not at the pipeline The instinct is to open the pipeline and read the YAML. Start one step later instead: at the deployment record on the environment itself, or failing that at the target's own audit log — the cloud provider's control-plane log, the cluster's audit trail, the registry's push history. That record answers the first question that actually partitions the investigation: **which identity performed this deployment?** If the identity is not the delivery pipeline's, the pipeline was never involved. Someone with standing credentials pushed the change directly. No gate written inside a pipeline can constrain a path that does not go through the pipeline, and adding another check is a non-fix. This case is more common than teams expect, because break-glass credentials issued during a past incident are rarely revoked. If the identity *is* the pipeline's, you have a run to inspect, and the question becomes why the gate did not stop it. ## The ways an in-pipeline gate fails to apply **It was not wired to this path.** The gate is attached to the pull-request workflow but the release runs on a tag; or it is conditioned on the default branch and the hotfix went out from a release branch; or it runs only for one directory in a monorepo and this change lived elsewhere. This is the single most common cause: the gate exists and works, on a route this change did not take. **It ran and failed open.** The step is set to continue on error so tool outages do not block releases; or the scanner's exit code is discarded by a shell pipeline whose status is its last command's; or the command exits 0 while writing findings to a report nothing reads. The step is green and the gate is decorative. **It was self-modifiable.** If the gate lives in the repository's own pipeline definition, a change can remove or weaken the gate in the same commit that needs to pass it — the change under review is also the change to the control reviewing it. This is why gate definitions belong in platform-level or organisation-level policy, evaluated outside the repository's control, or at minimum in a file requiring separate ownership approval to edit. **It was bypassed with sanctioned machinery.** A break-glass or emergency-deploy path exists, was used, and produced a record that looks like any other approval. If emergency use is not distinguishable from routine use at query time, nobody can tell you how often it happens — and in practice it becomes the fast lane. **The artifact was not the scanned artifact.** The gate scanned a build, and the deploy stage rebuilt from source or pulled a mutable tag that had since moved. The scan was truthful about something that is not what is running. ## Making the finding durable The design conclusions follow directly from the causes. *Only the delivery identity may write to production.* Humans hold no standing credentials to the target; the pipeline's identity does, and it is short-lived. This is the control that makes the deployment record complete, and therefore the one that makes every other gate meaningful. Without it your gates are advice offered to people who can ignore them. *Enforce gates where the change cannot reach them.* Organisation or platform policy, environment protection rules, or an admission control point in front of the target — anywhere that is not a file the change can edit. *Fail closed everywhere.* A gate whose tool failed must block. Assert exit codes explicitly rather than trusting a pipeline's default, and treat "the scanner did not run" as a failure result rather than an absence of a failure result. *Bind the gate to an immutable artifact identity.* The thing deployed must be the exact thing scanned, referenced by content digest or an equivalently immutable identifier, so a rebuild cannot silently substitute for it. *Make the emergency path loud.* Break-glass should be a distinct, obviously-labelled mechanism that notifies a channel, opens a record automatically, and is reviewed afterwards. It should stay available — removing it just pushes people back to direct credentials — but its usage rate is a metric someone reads. ## The organisational half One bypass is an incident; a *rate* of bypasses is a design signal. If the emergency path is used weekly, the gate is too slow, too flaky, or asks for something teams cannot supply, and the correct response is to fix the gate rather than to tighten the exception. Investigations that end with "we reminded the team of the policy" close the ticket without closing the path.

  • The deployment record shows a human identity, not the pipeline. What is the fix?
    Remove standing human write access to the target and leave the pipeline's identity as the only principal that can deploy. Humans get temporary, approved, recorded access through a break-glass path instead. Adding another pipeline gate does nothing here — the path being used never touches the pipeline.
  • Why is a gate defined in the repository's own pipeline file weaker than one enforced by platform policy?
    Because the change being gated can also change the gate. A single commit can delete the step, relax its threshold, or add a skip condition, and it passes on its own terms. Platform-level policy is evaluated outside the repository, so the same change cannot both fail the control and rewrite it.
  • Should you remove the emergency bypass path after an incident caused by it?
    No — removing it drives people to direct credentials, which are worse because they leave no correlated record. Keep it, make it obviously distinct from a normal approval, have it notify and open a record automatically, and treat its usage rate as a metric. Frequent use means the normal gate needs fixing.

saying these in an interview costs you the question

  • Investigates the pipeline before checking who actually deployed
  • Assumes every production change passed through the pipeline
  • Adds another gate to close a path that bypasses gates
  • Keeps gate definitions in a file the gated change can edit
  • Responds to repeated bypasses by restricting the exception rather than fixing the gate

context