skip to content

How do you run the exception queue for an ingest gate when three engineers serve 400 developers?

level: principalimportance: should knowfreq 40%

answer

  1. the real risk is bypass
  2. most requests should skip humans
  3. risk owner approves, platform owns policy
  4. decide what happens on timeout
  5. every exception expires

basics

~20 s

Design so most requests never reach a human, delegate the rest to the teams that own the risk, define what happens when nobody responds, and make every exception expire. A gate's real failure is bypass, not a bad approval.

solid answer

~50 s

Start from the failure you actually fear: not a wrong approval, but engineers routing around the feed because the queue is slow, which turns a controlled gate into a blind spot. So tier the traffic. Low-risk classes — a patch of an already-ingested package from the same publisher with a clean diff — should auto-approve, leaving humans only the residual. Delegate approval for the rest to the service-owning team, since they carry the risk; the three-person platform team owns the policy, the evidence shown to approvers, and the record, not each decision. Publish an SLA and, critically, define what happens on timeout, because the holiday case decides whether the gate survives contact with reality. Every exception is time-boxed, records who and why, and expires. Then measure the gate by queue latency, expiry follow-through and detected bypasses, never by how many things it blocked.

go deeper

for a junior

Understand that an exception is a deliberate, recorded decision to allow something the gate would normally hold, and that it should have an owner and an end date rather than being a quiet permanent hole.

for a middle

Explain why a slow queue is a security problem: people work around it, and the workaround is invisible. Be able to name a class of request that should not need a human at all.

for a senior

Show the operating design — tiers, evidence shown to approvers, SLA with a defined timeout behaviour, expiry enforcement — and how you would detect bypass in your own estate.

for a principal

Own the org-level tradeoff: where decision authority sits, what rigour you trade for adoption, how the platform team stays out of the critical path, and which metrics you would take to leadership to defend the gate's existence.

## Frame the problem correctly first The naive framing is *how do three people review enough requests*. That framing loses, because the arithmetic never works and no amount of process fixes it. The correct framing is: **an ingest gate's dominant failure mode is bypass, not bad approval.** A gate that is slow does not make teams wait; it makes them vendor the code into the repository, copy files by hand, add a second source, build in an unmanaged environment, or ship a container image with dependencies baked in from somewhere else. Every one of those converts a visible, controlled path into an invisible, uncontrolled one — and it happens quietly, so you keep believing the gate works. Every design decision below follows from that. ## Push the volume away from humans Most requests are not interesting. Define classes that resolve without a person: - A newer version of a package already in the feed, from the same publishing identity, with no anomalous change in the release. - A transitive dependency that is itself already ingested. - Anything covered by an existing standing decision for that package. What should remain for a human is genuinely small: a package the organisation has never used, a release that tripped an anomaly check, or a request to skip a hold-back window. If the human queue is not small, the automation rules are wrong, and that is where to spend the effort — not on hiring reviewers. ## Put the decision where the risk is Three platform engineers cannot hold context on 400 developers' problems, and they do not carry the consequences of a bad dependency choice. The service-owning team does. So delegate the approval to them, with the platform team owning three things instead: 1. **The policy** — what classes exist, what evidence each requires, what is never allowed. 2. **The evidence** — the approver sees the diff, the publisher, the age of the release and the existing inventory without going and finding it. An approver with no evidence in front of them rubber-stamps by definition. 3. **The record** — who approved what, when, why, and until when. This is a deliberate trade: you accept more approvers of variable rigour in exchange for a queue that actually moves and decisions made by people with context. The centralised alternative is more rigorous per decision and worse in aggregate, because it is the one that gets bypassed. ## Define the timeout, not just the SLA Publishing a target response time is easy. The question that separates a real design from a slide is **what happens when nobody responds** — the approver is on holiday, it is a Friday night, the request is the only thing between an on-call engineer and a fix. There is no universally right answer, but there is a required one: it must be **decided in advance and cheap to invoke**. A common shape is a break-glass path any senior engineer can take, which grants the exception immediately, notifies loudly, and creates a mandatory review item. Loud and reviewed beats a locked door, because the locked door is what teaches people to stop using the door. What you must not ship is an approval path with a single named human and no alternative. That design fails every holiday. ## Time-box everything An exception with no expiry is a policy change made by whoever was quickest to ask. Every grant carries an expiry and a reason, and when it lapses the state returns to the default — the package must be properly ingested, or the exception re-argued. This matters most for the hold-back-window skip: it should be an event with a lifetime, not a permanent property of that package. The corollary is that expiries need follow-through. An expiry nobody enforces is worse than none, because it creates a false record of control. ## Measure the right things Gates get justified with block counts, which measure activity, not safety. Better instruments: - **Queue latency**, p50 and p95. This is the leading indicator of bypass. - **Auto-approval share.** Falling share means humans are absorbing volume the rules should handle. - **Exceptions granted, and how many expired without follow-up.** - **Detected bypasses** — dependencies in shipped artifacts that never appeared in the feed, or build runners talking to public indexes. Finding these is a sign your telemetry works, not that your teams are hostile. ## What to say about the three-person team Be explicit that the team's job is to build and tune the path, not to sit in it. The moment their throughput is the constraint on 400 engineers, the design has already failed, and adding a fourth engineer buys a few months at most.

  • What is your default when the approver does not respond inside the SLA?
    It must be decided in advance and cheap to invoke. My preference is a break-glass path any senior engineer can take: the exception is granted immediately, the notification is loud, and a mandatory review item is created with an expiry. A silent auto-approve teaches people the queue is theatre, and a hard stop with one named approver fails every holiday and trains teams to route around the feed.
  • How would you detect that teams are bypassing the feed rather than using it?
    Compare what ships against what the feed holds. Emit resolved dependency sources from every CI build and alert on anything that did not come from the feed; inspect built images for components with no feed record; watch egress from build runners to public indexes. Also watch for the tell-tale shapes of avoidance — third-party code committed directly into repositories, or dependencies pinned to files rather than to the feed.
  • An auditor asks how you can prove a hold-back window was skipped safely. What do you show?
    The record: the specific package and version, who requested it and who approved, the stated reason, the evidence they were shown, the expiry, and what happened at expiry. The claim is not that the skip was risk-free — it is that a named person with the authority to accept that risk did so knowingly, and that the exception did not silently become the new default.
  • Would you ever accept a lower bar for approvals to keep the queue moving?
    Yes, deliberately and by class. A gate that everyone uses at moderate rigour beats a rigorous gate half the organisation avoids, because the second one has no visibility into the half that left. What I would not relax is the record and the expiry, since those are what let you go back and re-examine decisions later, and they cost the requester almost nothing.

saying these in an interview costs you the question

  • Requires a security review for every ingest request
  • Names a single approver with no fallback path
  • Grants exceptions with no expiry or recorded reason
  • Measures the gate by how many requests it blocked
  • Treats bypass as a discipline problem rather than a design signal
  • Answers by asking to hire more reviewers

context