skip to content

Your policy program's human-review queue is the bottleneck — how do you decide which decisions stay with people?

level: principalimportance: nice to knowfreq 33%

answer

  1. Review capacity is a fixed budget
  2. Reduce demand before adding reviewers
  3. Rank by the cost of being wrong
  4. Name what you do not review
  5. An unread queue is silent risk acceptance

basics

~20 s

Treat review capacity as a fixed budget. Automate every decision the artifact fully determines, spend the remaining review on the few classes where necessity genuinely matters, and explicitly accept the rest rather than queueing work nobody will do.

solid answer

~50 s

The mistake is treating review as unlimited because it is not a machine. It is the scarcest control you have, so allocate it deliberately. First, push everything decidable into rules: if the answer is fully determined by the artifact, a person adding nothing but latency should not be in that path. Second, reduce demand rather than supply — make the approved option the easiest one, so most teams never generate a review at all. Third, rank what remains by what a wrong answer costs: a bespoke base for a payment service earns a design review; the same deviation in an internal dashboard probably earns a recorded note. Fourth, and hardest, say out loud which classes you are not reviewing and why, rather than letting a queue depth make that decision silently. An unbounded queue is an unstated risk acceptance made by whoever stops reading first.

go deeper

for a junior

Understand that a human review is a limited resource, not a free fallback, and that sending everything to it makes it slower for the cases that really need it.

for a middle

Be able to spot a review whose outcome is already determined by the artifact — that one belongs in a rule, and moving it frees capacity for real judgement calls.

for a senior

Show you would triage by consequence and make the approved path easy, rather than treating queue depth as a staffing problem to solve later.

for a principal

Own the allocation explicitly: which classes get judgement, which get a recorded decision, which are accepted with a detective control instead, and who agreed to that split.

## Why this is a resource problem, not a policy problem Once you accept that some decisions cannot be encoded — necessity, intent, whether a deviation was warranted — the natural response is to route them to people. Do that across an estate of any size and you discover the second constraint: human judgement does not scale, does not queue gracefully, and degrades invisibly. A review that takes four days is not a slower control than one that takes four hours; past some threshold it is a different control, because people start routing around it or approving without reading. So the lead-level question is not "what can policy not decide?" but "given that people can absorb roughly N judgement calls a week, which N?" ## Four moves, in order **1. Automate everything the artifact determines.** Audit what reviewers actually do. Any review whose outcome is a function of fields already in the artifact is a rule that has not been written yet, and every one of those you leave in the queue is capacity stolen from a decision that genuinely needs a person. This is the cheapest capacity you will ever recover. **2. Reduce demand before you add supply.** Most deviations are not principled; they happen because the approved path was harder to find or harder to use than the alternative. A well-maintained set of approved bases that actually covers the languages and runtimes teams need eliminates whole categories of review without anyone deciding anything. Adding reviewers is the expensive answer to a question that better defaults answer for free. **3. Rank the remainder by the cost of being wrong.** Not every judgement deserves the same depth. A useful three-tier split: a *design review* for deviations on systems where being wrong is expensive and hard to reverse; a *lightweight recorded decision* — one named person, one written reason — for the middle; and *automatic acceptance with a record* for the low-consequence tail. The tail is where most of the volume lives and where review adds the least. **4. State what you are not reviewing.** This is the move most programs skip, and it is the one that distinguishes a decision from a drift. If you cannot review the tail, say so explicitly, name the class, and say what you rely on instead — usually a detective control that catches the consequence rather than the cause. An unreviewed queue is the same risk acceptance, made by the person who stopped reading, recorded nowhere, owned by nobody. ## The tensions you will be asked about **Consistency versus judgement.** A rule decides identically every time; two reviewers decide differently. When teams compare notes and find the same deviation approved in one place and refused in another, trust in the whole program drops. Written precedent — the reasons previous deviations were accepted — is what keeps human decisions from looking arbitrary, and it costs almost nothing to keep. **Latency versus rigour.** Every review you add is time on someone's critical path, and the cost is paid by teams that mostly did nothing wrong. If the review cannot be fast, it must be rare. **The temptation to automate the last mile.** There is always pressure to encode the judgement anyway, because a rule is cheaper than a reviewer. This is where a program starts writing clauses that encode past verdicts, and the result is a predicate nobody can read that is still wrong about the next novel case. Resist it explicitly: automating a judgement does not make it decidable, it makes it silently wrong. **The reviewer's own limit.** A person given fifty near-identical cases per week stops judging and starts pattern-matching. At that point you have the cost of review with the discrimination of a bad rule. Queue volume is therefore not a scheduling detail — it determines whether the control works at all. ## What a good answer sounds like A lead who has done this can say, for their estate: here is what the rules decide alone; here is the small class we route to a person and who that person is; here is what we deliberately do not review and what covers it instead; and here is how we noticed when the queue grew past what the reviewers could absorb. The weakest answer is that everything undecidable goes to review — that is not an allocation, it is the absence of one.

  • A stakeholder proposes hiring more reviewers instead of changing anything else. What do you say?
    That it is the most expensive fix and usually the last one. First I want the reviews whose outcome is already determined by the artifact moved into rules, and the approved path made easy enough that most deviations stop happening. If demand is still above capacity after that, more reviewers is a legitimate answer — but bought against a known volume and a known cost of being wrong, not to drain a queue we never triaged.
  • How do you keep human decisions from looking arbitrary across teams?
    Keep written precedent. Every accepted deviation records the reason it was accepted, and reviewers read the prior ones before deciding. It does not make judgement mechanical, and it should not — but it means two teams with the same situation get the same answer, and a reviewer refusing something previously allowed has to say what is different. Inconsistency, more than strictness, is what makes a program lose credibility.

saying these in an interview costs you the question

  • Routing every undecidable case to review and calling it a plan
  • Adding reviewers before removing reviews rules could make
  • Letting queue depth decide what goes unreviewed
  • Encoding past verdicts as clauses to save reviewer time
  • Ignoring that overloaded reviewers stop actually judging

context