skip to content

You own a GitOps manifest repository serving dev, staging and production for a dozen services. How would you design the promotion path — what is automated, where a human belongs, and how do you stop environments diverging?

level: principalimportance: nice to knowfreq 36%

answer

  1. automate the commit, vary the approver
  2. gates should check evidence, not vibes
  3. enforce the ordering with a check
  4. allowlist what may differ
  5. measure out-of-band edits

basics

~20 s

Automate the bump into dev, automate staging behind passing tests, and make production an automated pull request that a human approves. Keep environments from diverging by enforcing that only an allowlisted set of keys may differ between overlays, checked on every pull request.

solid answer

~50 s

I would make the promotion commit itself always automated and vary only who has to agree to it. A bot writes the new digest into the dev overlay on every successful build; staging is bumped automatically once the post-deploy checks pass; production gets an opened pull request that a code owner approves, so the human decision is an approval rather than a hand-edited YAML file. I would enforce ordering — a check that refuses a production digest that never ran in staging — and forbid direct edits to the production overlay outside that flow. Divergence is controlled by rendering all overlays in CI and failing if they differ outside an allowlist such as image, replicas, hostnames and limits. The signals I would watch are promotion lead time, how often production is edited outside the automated path, and how often a promotion is reverted.

go deeper

for a junior

Understand that the promotion commit can be written by automation and that production usually still requires a person to approve the pull request before it merges.

for a middle

Be able to describe the tiering — automatic to dev, gated to staging, approved to production — and explain why the commit content is generated rather than hand-edited.

for a senior

Demonstrate enforcement rather than intent: a check that the production digest already ran in staging, protected paths, rendered-overlay diffs against an allowlist, and a rehearsed revert whose image is guaranteed still pullable.

for a principal

Own the whole policy: which gates carry real signal, how to keep approvals from decaying into rubber stamps, one uniform mechanism across services for auditability, and the metrics — lead time split by wait cause, out-of-band edits, revert rate — that tell you whether the design is holding.

## The design question underneath "How much human approval?" is the wrong framing, because approvals are cheap to add and quietly become rubber stamps. The better framing is: **what does each gate actually check, and would a human notice if it were wrong?** Design the path so that machines check the things machines are good at — did this exact artifact pass the tests, did it run in staging, does the diff touch only what a promotion should touch — and reserve humans for the judgment machines cannot make, which is usually timing and business risk rather than correctness. ## A path that works for a dozen services **Dev: fully automatic.** Every successful build writes the new digest into the dev overlay. No approval, no ceremony. The value of dev is that it is always the newest thing, so it finds integration breakage early. **Staging: automatic, but gated on evidence.** The bump lands only after the artifact's tests pass, and ideally after dev has been healthy for some minimum period. This is where the artifact earns its promotion record. **Production: automated pull request, human approval.** A bot opens the PR with the digest, the source commit, and the changelog between the current production digest and the new one. A code owner approves. The important detail is that the *content* is generated: a human editing production YAML by hand is how a version bump quietly becomes a replica-count change. The uniform shape matters more than the specific gates. Twelve services sharing one flow means one thing to teach, one thing to audit, and one place to change policy. ## Enforce the ordering, do not merely intend it If promotion is "copy the digest from the staging overlay", make it a check rather than a convention. A CI job on the production pull request can assert that the proposed digest is the one currently in the staging overlay, or appears in a promotion ledger of digests that have run in staging. This kills the most common escape — a hotfix built and pushed straight at production, skipping every check, at exactly the moment judgment is worst. ## Keep divergence measurable The reason environments drift is that nothing fails when they do. Render every overlay in CI and diff them, then fail the build when they differ outside an explicit allowlist — image reference, replica count, hostnames, resource limits, and a named handful of flags. A team adding a genuinely new per-environment difference must extend the allowlist in a reviewed commit, which is exactly the conversation you want to force. Everything else belongs in the shared base. ## Constrain who can write where The production overlay path should be protected: code-owner review required, no direct pushes, the automation account allowed to *open* pull requests but not to merge them. This is the practical meaning of "the repository is production" — write access to that path is deployment access, and it should be reviewed as carefully as cluster credentials. ## Rollback as a first-class path Rolling back must be faster than rolling forward or people will not use it. Revert of the promotion commit, opened by the same automation, with the approval requirement kept but the review deliberately trivial because the diff is a digest the environment previously ran. Rehearse it, and make sure the previous image is still present in the registry — an aggressive registry retention policy silently removes the ability to roll back, which teams discover at the worst moment. ## What I would measure - **Promotion lead time** from merged code to production, split by waiting-on-machine and waiting-on-human. If humans dominate, the gate is queueing, not deciding. - **Out-of-band changes**: commits to the production overlay that did not come from the automation. This should be near zero; every one is either an emergency worth reviewing or a hole in the flow. - **Revert rate and time-to-revert**, which say more about delivery health than deployment frequency alone. - **Allowlist growth**, as a proxy for environments quietly forking. ## What I would deliberately not do I would not add an approval to staging just because production has one — an approval nobody can meaningfully refuse trains people to click through the one that matters. I would not let each of the twelve services invent its own promotion mechanism, because the audit story then requires reading twelve pipelines. And I would not treat the automated pull request as the safety mechanism: it is a *record*. The safety comes from the artifact having run somewhere real first, and from the ability to revert quickly.

  • How do you prevent the production approval from becoming a rubber stamp?
    Make the pull request carry decision-grade content — the digest, the source commit range, the changelog, and the check results — so approving means reading something specific. Keep approvals rare enough to matter by not requiring them where nothing can be refused, and track how often an approver actually blocks. An approval that has never once been withheld is a queue, not a gate.
  • What breaks a revert-based rollback even when the manifest history is intact?
    The image being gone. Registry retention that prunes untagged or older digests removes exactly what the revert points at, and the sync then fails on image pull. Retention has to keep anything currently or recently referenced by any environment overlay. Stateful changes are the other class — a migration that ran does not un-run because the manifest went back.
  • With a dozen services, would you give each its own promotion pipeline?
    No — one shared mechanism, parameterized per service. The reason is auditability and change cost: one place to add a check, one flow to teach, one report of who approved what. Per-service variation should be data — which environments exist, who owns the production approval — not twelve separately maintained pipelines that drift apart the way the environments would.

saying these in an interview costs you the question

  • Adds approvals everywhere and calls it governance
  • Lets an urgent build go straight to production
  • Relies on convention that prod uses staging's digest
  • Assumes overlays stay aligned without a check
  • Never verifies the old image is still pullable

context