What test evidence justifies letting a dependency bump merge without a human reviewing it?
answer
- green is not the same as covered
- the build runs upstream code first
- scope the rule to your evidence
- blast radius, not the version number
- cheap rollback substitutes for review
basics
~20 sOnly evidence that actually exercises the changed dependency. A green build proves the covered paths still work and nothing about the rest, so the unattended-merge rule must stay narrow enough that your tests cover the risk it takes.
solid answer
~50 sStart by separating two decisions people merge into one. The upstream release's own build and install code already runs when CI builds the proposal, before any merge decision, so that job must hold no deploy role, no signing key and no long-lived token -- otherwise a compromised release reaches your credentials whether or not anyone approves. The unattended-merge question is narrower: what reaches the release train unread. Tests are evidence only for the paths they execute, so a green run on an uncovered call path is absence of evidence, not safety. A defensible rule is scoped on four axes: the dependency's scope (build or test tooling versus something in the request path), its blast radius (a formatter versus an auth, crypto or deserialisation library), the size of the change, and whether a bad outcome is cheaply reversible via staged rollout and automated rollback. Anything failing those gets a human.
go deeper
Know that a passing build only tells you the tests that ran, passed, and that nothing was checked in code the tests never touch. Say that plainly rather than treating green as safe.
Explain when a dependency's own build code executes -- during the proposal's CI run, before any merge -- and why that makes the privileges of that job a separate concern from the merge rule.
Produce an actual rule with named axes: scope, blast radius, change size and reversibility. Show how a staged rollout with automatic rollback lets you widen automation safely, and which packages you would never automate.
Own the trade between review capacity and delivery speed across an estate, and decide what the organisation must build -- coverage, canaries, rollback -- before automation is a bet worth taking. Be ready to say where you would spend first.
## Two decisions that get confused The question sounds like one decision -- may this merge itself? -- and is really two. **Decision one: what is allowed to execute a brand-new upstream release, and with what privileges?** In most ecosystems, resolving and building a dependency executes code the publisher wrote: install hooks, build plugins, code generation, native compilation steps. That happens **when CI builds the proposal**, before any merge decision exists. Getting this direction wrong is the single most common mistake on this topic: auto-merge is not what causes the new code to run. So the build that installs a freshly proposed release must be treated as running untrusted code -- no production deploy role, no signing key, no long-lived registry credential, minimal token scope, restricted egress, and no cache or workspace it can poison for a job that *is* trusted. Get that wrong and a compromised release lands your cloud deployment credentials no matter how carefully a human reviews the diff afterwards. **Decision two: what may reach the release train without anyone reading it?** That is the actual policy question, and it turns on evidence. ## Tests are evidence only where they execute A green build means every assertion that ran, passed. If the suite covers 40% of the service and the bumped library is called from the other 60%, the run asserted nothing about the change. It is not weak evidence -- it is *no* evidence about the risk you are taking, dressed as a green check. The rule that follows is uncomfortable but simple: **an unattended-merge policy is only as broad as the evidence behind it.** A thin suite does not forbid automation; it narrows what may be automated. That also means coverage is not a global gate. What matters is whether the specific calling paths into the specific bumped package are exercised, in a test that runs on the proposal. A service with 40% coverage may still have excellent coverage of its payment path and none of its admin console; the first can be automated and the second cannot. ## The four axes of a defensible rule **Scope.** Build-time and test-only tooling never reaches a customer, so its failure mode is a broken build -- loud, immediate, cheap. Runtime dependencies fail in production. These deserve different rules. Note the asymmetry though: build tooling that runs in a privileged job has a *worse* compromise story than a runtime library, even though it has a better bug story. Scope decides review; privilege decides where it runs. **Blast radius.** A formatter and a deserialisation library are not the same bet. Anything on the authentication, authorisation, cryptography, deserialisation or templating path should be read by a human regardless of how small the version change looks, because the failure is silent rather than loud. The test suite that would catch a subtle change in these is rarely the suite you have. **Change size.** A lockfile-only movement of a transitive package is a smaller claim than a direct dependency crossing a major. Note carefully what the version number is and is not: it is the publisher's assertion about their own change, not evidence produced by you. Your tests are the evidence. **Reversibility.** This is the axis that actually buys the automation. If the change flows through a staged rollout with health-based automatic rollback, a bad merge costs minutes and a fraction of traffic, and automation is a cheap bet. If it flows straight to all users with a manual, hour-long rollback, the same merge is an expensive bet and deserves a human even with good tests. **Cheap reversal is a substitute for expensive review.** ## Writing it as a rule someone can apply at 2am A policy that requires judgment is not a policy. Something like: *unattended merge is permitted when all four hold -- the package is build or test scope, or the calling paths into it are covered by tests that ran on this proposal; the change does not cross a major; the diff touches only manifest and lockfile entries; and the deployment it feeds is staged with automatic rollback. Otherwise a named human reads it.* An on-call engineer can apply that at 2am without a meeting, and everything outside it fails closed to review rather than failing open. ## The parts people forget **Audit truth.** Unattended does not mean unrecorded. Every automated merge needs an attributable identity, a retained diff and retained build logs, so that *what changed in production last night* is one query. A policy that cannot answer that has traded away non-repudiation for convenience. **The bot is not the trust anchor.** The proposal machinery is trustworthy; it faithfully proposes whatever the publisher released, including a release from a hijacked maintainer account. Automating merge is a statement about your evidence and your reversibility, never about trusting the source. **Fail closed.** When a test is flaky, when coverage data is missing, when the change touches something unclassified -- the rule must send it to a human. An unattended-merge policy that treats *unknown* as *permitted* is how the entire class becomes automated by accident.
- Your suite covers 40 percent of the service. Does that rule out unattended merges entirely?No, it narrows them. Automate build and test-scope packages, plus runtime packages whose calling paths sit inside the covered 40 percent, and send everything else to a human. Alternatively buy evidence another way -- a staged rollout with automatic rollback makes a bad merge cheap enough that thin test evidence is tolerable.
- What should the CI job that installs a freshly proposed upstream release be allowed to touch?Nothing a compromise could monetise or reuse. No production deploy role, no signing key, no long-lived registry token, narrow permissions, restricted egress, and no shared cache or workspace that a trusted job later consumes. Assume the code in that job is attacker-controlled and ask what it could reach.
- How do you keep unattended merges auditable?Give the automation an attributable identity rather than a shared human account, retain the diff and the build logs for every automated merge, and make deployments traceable back to them. The test is whether you can answer what changed in production last night in one query, without reconstructing it from memory.
- Which dependencies would you exclude from unattended merge no matter how good the tests are?Anything on the authentication, authorisation, cryptography, deserialisation or templating path. Their failure mode is silent -- a subtly weakened check still passes a functional suite -- so the evidence a test run provides does not match the risk. Those are read by a person, every time, however small the version change looks.
saying these in an interview costs you the question
- Treats a green build as proof the change is safe
- Thinks auto-merge is what causes upstream code to run
- Runs untrusted dependency builds in a job holding deploy credentials
- Applies one blanket rule to every dependency class
- Trusts the version number as evidence rather than a claim
- Leaves unclassified changes to merge by default
- Automates merges with no attributable record of what landed