Your shared release template has a flag that skips the signing stage, and twelve services set it. What do you do?
answer
- The hatch is evidence, not just a bug
- Ask the twelve why first
- Removing it moves traffic off-road
- Owner, reason, expiry, report
- Explicit unsigned marker beats silence
basics
~20 sFind out why the hatch is used before removing it: it usually marks a gap in the paved road. Make every use attributed, dated and reported, close the gaps, then narrow the hatch to an approved exception and delete it.
solid answer
~50 sDeleting the flag on Monday is tempting and usually the wrong first move. An escape hatch twelve teams reach for is evidence the paved road has a gap: the signing stage is slow, or cannot get a key in a fork-triggered run, or does not handle a legacy artifact type. Remove it without closing the gap and those teams do not start signing; they fork the template or hand-roll a pipeline, and you lose the visibility you had. So instrument first: every use attributed to a person, dated, given an expiry, and reported, with the release record carrying an explicit unsigned marker rather than a silent absence. Then fix the gaps, convert the hatch to an approved exception, and remove it. The property at stake is attribution: a release nobody vouched for cannot be traced back later.
go deeper
Understand that skipping the signing stage means the released artifact has nobody vouching for it, so it cannot be attributed later, and that an exception needs a recorded owner and reason.
Be able to explain the mechanics of instrumenting an override: attribution in the change, an expiry that fails the build, a report, and an explicit unsigned marker in the release record instead of silent absence.
Show that you would diagnose before removing, name the concrete gaps that drive hatch use, and sequence the work so usage falls before the hatch narrows. Explain why the naive removal loses visibility.
Own the argument that a template is convenience and the acceptance point downstream is the control, and decide what level of attributed, reviewed exception the organisation should accept rather than chasing zero.
## Read the hatch as a signal, not just a defect A one-line override that skips the signing and attestation stage did not appear by accident. Somebody on the platform team added it because a real team was blocked and the release had to go out. Twelve services using it is data: the paved road has a gap wide enough that a dozen teams walked around it. The first task is diagnosis, and it is cheap — ask the twelve. Typical answers are mundane and fixable. The signing stage adds minutes to a hot-fix path that people run under pressure. A build triggered from a fork cannot reach the key material, so the stage fails and the flag is the only way to get a green run. An older artifact shape is not supported by the stage. A team inherited the flag from a copied config and never knew it was set. ## Why removing it first backfires The control you actually care about is not the flag; it is whether releases are attributable. The template is one of several places that could enforce that, and it is the weakest, because the team that runs the pipeline can edit its own configuration. If you delete the hatch while the underlying reason still exists, the affected teams have three options: block their release, patch around the template, or leave the template entirely. Two of those three make your estate less observable than it was. You will have removed the *evidence* of non-compliance rather than the non-compliance. This is the general shape of paved-road work: the road holds traffic only while it is the easiest route. Any change that makes the road impassable for a real use case pushes that traffic off-road, where you cannot see it. ## Instrument before you enforce Make the hatch loud rather than quiet: - **Attribute it.** Setting the flag requires a named owner and a stated reason recorded in the change, not just a boolean in a config file. - **Expire it.** The flag carries a date after which the build fails. An exception without an expiry is a permanent policy change made by one team on a Thursday afternoon. - **Report it.** A recurring report lists every service releasing without a signature, and it goes to the same people who get told the control coverage number — so the coverage figure and the exception list can never disagree. - **Make the absence explicit.** The release record should say *this artifact was released unsigned, by this person, for this reason*, rather than simply lacking a signature. Silence and a deliberate exception look identical to a later investigator, and only one of them is an accountable decision. Instrumentation alone often reduces usage sharply, because most hatch use is habit or inheritance rather than necessity, and nobody wants their name on a monthly report. ## Then close the gaps, then narrow the hatch With the reasons known, the fixes are ordinary engineering: make the stage fast enough that it is not worth skipping, give fork-triggered runs a path that does not need the key, support the missing artifact shape. Each fix should visibly retire one column of the report. Once usage is down to genuine edge cases, change the flag from self-service to a request that a named approver grants for a period. Finally delete it, when the report shows nobody would notice. ## The insider angle There is a sharper reading of this scenario. A silent one-line flag in a service's own configuration is a very quiet way for someone with legitimate commit access to ship an artifact that carries no attestation of who produced it — and if the downstream consumer accepts unsigned artifacts, that is the whole attack. This is why the reporting matters more than the flag's existence, and why the durable enforcement point is downstream of the pipeline, at the place artifacts are accepted, rather than inside a template the releasing team can edit. The template's job is to make the right thing effortless; making it mandatory is somebody else's control, and the two are not substitutes. ## What a strong answer sounds like Diagnose, instrument, close the gap, narrow, remove — in that order, with a sentence on why the naive order fails and a sentence on what property you are protecting. Candidates who lead with policy language, or who propose an immediate ban, generally have not run a platform team through this.
- What do you record when a team legitimately needs the hatch for one release?A named owner, the reason, an expiry date, the artifact digest, and an explicit statement in the release record that this artifact went out unsigned. The point is that a later investigator can tell a considered exception from an accident. An absent signature with no explanation reads identically to a control that quietly failed.
- Where should the requirement that a release is signed actually be enforced?Downstream of the pipeline, where artifacts are accepted for deployment, because the team that runs the pipeline can always change its own template configuration. The template's job is to make signing the effortless default; the acceptance point is what makes it non-optional. Treating the template as the enforcement mechanism confuses convenience with a control.
- Instrumentation shows usage dropped to two services and both have valid reasons. Now what?Stop pushing on those two and fix the reason instead, or grant a standing exception with a review date and an owner who is not the requesting team. Two known, attributed, reviewed exceptions are a healthy end state. Chasing them to zero for the sake of a round number is how a paved road starts losing the traffic that matters.
saying these in an interview costs you the question
- Deleting the flag immediately without diagnosing why it is used
- Treating the twelve teams as non-compliant rather than blocked
- Leaving the hatch silent because usage is low
- Assuming the template is where the requirement gets enforced
- Accepting an exception with no owner or expiry date