skip to content

Your team replaced creation on first use with a reviewed request, and engineers are now routing around it — what did the gate get wrong?

level: seniorimportance: should knowfreq 48%

answer

  1. a control competes with its alternative
  2. the removed path took seconds
  3. reuse of an existing stream hides the damage
  4. pre-agreed bands, humans on exceptions
  5. publish a time-to-stream service level

basics

~20 s

The gate is slower than the path it replaced, so avoiding it is the cheapest option. A reviewed path only holds if it delivers a stream faster than the workaround does; otherwise engineers reuse an unrelated stream and the estate degrades invisibly.

solid answer

~50 s

A provisioning gate does not compete with nothing; it competes with the path it removed, which delivered a stream in seconds. If the reviewed path takes days, the rational move for a blocked engineer is to avoid it: publish an unrelated payload onto a stream that already exists, ask somebody with broad grants to create one by hand, or build in a lower environment and never declare the production stream properly. Each of those is worse than the original problem, because the damage is now invisible — the estate looks governed while its names and grants quietly stop meaning anything. The fix is to make the declared path faster than the workaround: pre-agreed bands that a pipeline applies automatically, a human involved only for requests outside them, and time-to-stream measured as a service the platform team owns.

go deeper

for a junior

The takeaway is that a rule people can avoid gets avoided. If getting a stream properly takes days and getting one another way takes minutes, most engineers under deadline will take the second route.

for a middle

Explain the mechanics of the bypass: which alternatives stay open after the fast path is closed, and why reusing an existing stream leaves every estate metric looking healthy while the names and grants stop meaning anything.

for a senior

Show operating judgement: publish a time-to-stream target, automate everything inside agreed bands, and reserve human review for exceptions. Say plainly that a gate slower than its workaround is not a control at all.

for a principal

Take the policy position: the platform team owns provisioning as a service with a service level, not as an approval queue, and the organisation should measure streams lacking a declaration rather than streams created.

## A gate competes with what it replaced The moment a platform team stops streams appearing on a client's say-so, it takes on an obligation it may not have noticed: **something has to create streams, and it has to do so at a speed the people who need them will accept.** The baseline is brutal, because the path just removed took no time at all. This is the single most common way a well-intentioned provisioning policy fails. The policy is right, the values it asks for are the right values, and the estate still degrades — because the policy is slower than the alternatives it did not close. ## What routing around it actually looks like The workarounds are rarely defiant. They are what a reasonable engineer does at five o'clock with a release waiting: 1. **Reuse an existing stream for an unrelated payload.** The most damaging one. Nothing new appears in any list, so the estate looks unchanged, while the stream's name now lies about its contents, its owning team is answerable for data it has never seen, and the grants written against it now expose a second payload to everyone who could read the first. 2. **Find a human with broad grants.** A stream gets created by hand, outside the declared definitions, so nothing in version control describes it and the next pipeline run does not know it exists. 3. **Build in a lower environment and promote nothing.** The production stream is created at the last minute, by whatever means works, with no declaration behind it. 4. **Multiplex inside the payload.** Several logical flows are stuffed into one stream and separated by a field, which pushes the governance problem down into application code where the platform team cannot see it. All four convert a visible queue of pending provisioning requests into an invisible mess. The metric that looked good — few new streams — is exactly the wrong one. ## What the reviewed path must be measured on | Design | Time to a usable stream | What it costs the estate | |---|---|---| | Creation on first use | Seconds | Unowned streams with values nobody chose | | A slow reviewed path | Days | The four workarounds above, invisibly | | A fast declared path | Minutes | A review only where the request is unusual | The target is not "a review exists". It is **time to stream**, treated as a service level the platform team publishes and is held to. If that number is worse than the workaround, the policy is decorative. ## Designing a gate that survives contact The usable shape is narrow, and it is broadly the same everywhere: - **Declared, not requested.** The definition — name, owning team, and the values whose wrong setting is expensive on your platform — lives in version control and is applied by a pipeline. The artefact is the source of truth, not a conversation. - **Bands, not approvals.** Agree in advance the ranges that need no discussion. A request inside them applies automatically on merge. Only a request outside them — an unusually large parallelism count, a retention far beyond the norm, a stream in a space it does not belong to — waits for a person. - **Review the values, not the existence.** Almost nobody creates a stream that should not exist. What review adds is catching a value that will hurt later, and a name that will mislead, which is why the human step belongs on the exceptions. - **Fail loudly, early and locally.** A malformed request should be rejected by the pipeline in seconds, in the requester's own change, not discovered in a weekly meeting. - **Make the declared path the easy one.** A template, a single command, and an example to copy beat any amount of policy text. ## The judgement being tested An interviewer asking this is not checking whether you approve of governance. They are checking whether you have watched a control fail by being avoided rather than by being wrong. The senior answer names the competition explicitly — the gate is racing the path it replaced — accepts that engineers responding to a slow gate are behaving rationally, and puts the burden on the design: **if the reviewed path is not faster than the workaround, the workaround is the real policy.**

  • Why is reusing an existing stream for an unrelated payload worse than the typo problem the gate was meant to solve?
    Because it is invisible. A stray stream at least appears in a list and can eventually be investigated. A reused stream leaves every count unchanged while its name stops describing its contents, its recorded owning team answers for data it has never seen, and every grant written against it now covers a second payload.
  • What single number would you put on a dashboard to tell whether the gate is working?
    Median and worst-case time from a declared request to a usable stream, tracked against the number of streams that exist without a matching declaration. The first tells you whether the path is competitive; the second tells you whether anyone is bypassing it. Either alone is easy to make look good.
  • Does adding review necessarily mean adding a person?
    No, and treating those as the same thing is what makes gates slow. Most of what review catches — a missing owning team, a name that breaks the convention, a value outside the agreed band — is mechanical, so a pipeline can enforce it on merge. A person is worth waiting for only on the genuine exception.

saying these in an interview costs you the question

  • Blames engineers for lacking discipline rather than the gate's speed
  • Counts fewer new streams as proof the policy is working
  • Thinks every provisioning request deserves a human reviewer
  • Treats reuse of an existing stream as harmless resourcefulness
  • Assumes removing the fast path automatically produces a governed estate