A moderation team can hand-label 20,000 posts a week while the platform receives 4,000,000. Which posts should reach a human?
answer
- capacity is 0.5 percent of the volume
- cheap rules take the confident bulk
- humans get the uncertain band only
- set band edges from capacity
- reserve a random slice for audit
basics
~20 sHumans can touch 0.5 percent of the volume, so spend it where it buys most: high-precision programmatic rules decide the bulk, uncertainty sampling sends the genuinely ambiguous band to annotators, and a small random slice re-reviews the automatically decided stream.
solid answer
~50 sStart from the gap: 4,000,000 against 20,000 is 200 to one, so annotation capacity, not model choice, is the binding constraint on the design. Route in a cascade. High-precision programmatic rules, such as hash matches on known-bad media and narrow pattern rules, label the large confident bulk with no human involved. A score then auto-decides the confident bands at either end. Only the uncertain band in the middle reaches annotators, selected by uncertainty sampling, and the band edges are tuned so its weekly volume matches capacity rather than the other way round. Reserve part of the budget, say 2,000 of the 20,000 touches, for a stratified random re-review of the automatically decided stream. That reserve is the only slice that gives an unbiased read of the automatic decisions, because the uncertainty-selected queue over-represents ambiguous posts by construction.
code
pseudocode · 19 linesfor each incoming post:
r = programmatic_rules(post) // hash match, narrow patterns, repeat offender
if r.matched and r.precision_tier == 'high':
decide(post, r.verdict)
audit_pool.add(post) // eligible for random re-review
continue
s = risk_score(post)
if s >= high_edge or s <= low_edge:
decide(post, verdict_from(s))
audit_pool.add(post)
continue
human_queue.add(post, priority = uncertainty(s)) // the middle band only
weekly:
send_to_annotators(top_n(human_queue, 18000))
send_to_annotators(stratified_random(audit_pool, 2000))
tune(high_edge, low_edge) so that size(human_queue per week) ~= 18000go deeper
Recall that humans can only see a tiny fraction of the volume, so most posts must be decided automatically and the human queue has to be chosen deliberately rather than taken in arrival order.
Explain the cascade: cheap rules for the confident bulk, score bands for most of the rest, uncertainty sampling for the ambiguous middle, and why the band edges are derived from annotation capacity.
Show the consequences you would defend in review: the human-labelled set is a selection and not a sample, so a random audit reserve is carved out of the same budget and never spent on more uncertain-band verdicts.
Own the allocation as a spending argument. Split the budget across redundancy, the uncertain band and audit by the cost of a wrong decision per policy category, and state the coverage the platform is knowingly buying.
## Annotation capacity is the binding constraint The first number to say out loud is the ratio. 4,000,000 posts a week against 20,000 human touches is **200 to one**, or 0.5 percent coverage. No annotator-pool expansion closes a 200x gap, and the design question is therefore not *how do we label everything* but **where does a scarce human verdict buy the most**. Everything below follows from that framing, and an answer that reaches for a bigger pool before it fixes triage has not engaged with the constraint. ## The routing cascade 1. **Programmatic rules first.** Exact and perceptual hash matches against already-removed media, narrow high-precision patterns, and reputation rules on repeat offenders. These are cheap, deterministic and auditable. Say they confidently decide 3,680,000 posts, about 92 percent. 2. **A score decides the confident bands.** Of the remaining 320,000, the posts scoring very high or very low are auto-decided. Around 302,000 fall in those bands. 3. **The uncertain band reaches humans.** That leaves roughly 18,000 posts in the middle - and the band edges are chosen so this volume **matches capacity**, which is the actual operational trick. The threshold is set by the annotator budget, not chosen first and lamented afterwards. 4. **The audit reserve.** The last 2,000 touches are a stratified random sample **drawn from the automatically decided stream**, re-reviewed by the same annotators against the same guideline. The totals reconcile: 3,680,000 plus 302,000 gives 3,982,000 automatic decisions, 18,000 go to humans as first verdicts, and 2,000 automatic decisions are re-reviewed, for 20,000 human touches in all. ## Why uncertainty sampling, and what it costs **Uncertainty sampling** picks the items the current model is least sure about, on the reasoning that a verdict on a post the model already scores confidently teaches it almost nothing, while a verdict on a borderline post moves the decision boundary. Per verdict bought, it is the highest-value selection available. It has a cost that must be said in the same breath: **the human-labelled set is now a selection, not a sample**. It over-represents the ambiguous middle by construction. Any rate computed on it - the violation rate, the disagreement rate, the reviewer's own accuracy - describes the uncertain band and not the platform. This is exactly why the random audit reserve is not optional: it is the only slice whose composition matches the traffic it was drawn from. ## Comparing the label sources | source | unit cost | coverage | characteristic failure | |---|---|---|---| | programmatic rules | near zero | the confident bulk | correlated: they miss a new pattern in the same way at once, and carry no estimate of their own error | | the score's confident bands | near zero | most of what rules leave | inherits whatever the model already believes, including its blind spots | | human verdicts on the uncertain band | high | a fraction of a percent | a selection, not a sample; nothing measured on it generalises to the platform | | random audit of automatic decisions | high | a tiny random slice | small counts, so a rare category needs stratification to produce any usable number | ## Weak and programmatic labelling in more detail Rules and heuristics that emit labels are what weak supervision, or data programming, formalises: many imperfect labelling functions vote on an item, their disagreements are used to estimate how much each should be trusted, and the result is a large, cheap, noisy label set. In a moderation operation this is how the bulk gets labelled at all. Two honest limits: the functions are written by people looking at yesterday's abuse, so they are **correlated** and go wrong together when posting behaviour shifts; and they are silent about their own accuracy unless something independent measures it. The audit reserve is that something. ## What the design has to state - The **ratio** and the coverage it implies, before any mechanism is proposed. - **Who decides what**: rules, score bands, humans, and where each boundary sits. - That the band edges are set **from capacity**, so the queue clears each week rather than growing an unbounded backlog. - That a fixed slice is spent on random re-review, and that dropping it to buy more uncertain-band verdicts destroys the only unbiased estimate the operation has of how the other 3,982,000 decisions went. - That the uncertain band is where **new** abuse appears first, so it is also the early-warning surface, not merely the expensive one.
- What does uncertainty sampling do to the human-labelled set itself?It makes it a selection rather than a sample. The set over-represents the ambiguous middle by construction, so the violation rate, the disagreement rate and the reviewer accuracy measured on it describe that band and not the platform. Quoting any of them as a platform-wide figure is a reporting error, and the random audit slice exists precisely because its composition does match the stream it came from.
- The programmatic rules confidently label 92 percent of posts. Why not stop there?Because rules encode the patterns someone already wrote down. They are correlated, so a new abuse pattern slips past all of them at once, and they carry no estimate of their own error until something independent measures it. The uncertain band is where genuinely new behaviour surfaces first, which makes the human queue an early-warning surface as well as a label source.
- Would hiring twice as many annotators change the design?It moves the band edges, not the shape. At 40,000 touches the uncertain band can widen and the audit slice can grow, but coverage goes from 0.5 to 1 percent and the cascade is still what decides 99 percent of posts. Triage is the design; pool size is a dial on one of its stages, and doubling it before triage works simply buys more verdicts on posts that were already being decided correctly.
- What happens if the uncertain band's weekly volume exceeds capacity?An unbounded backlog forms and the age of a pending post grows without limit, which in moderation is itself a harm. The band edges have to be tuned back so the band clears each week, and the posts pushed out of the band are auto-decided with the conservative verdict for their category. That trade is stated explicitly rather than discovered as a queue that never empties.
saying these in an interview costs you the question
- Spends the whole human budget on a uniform random sample of posts
- Sends every post the score calls a violation to a human reviewer
- Treats programmatic rule outputs as equivalent to human verdicts
- Drops the random audit slice because uncertain posts teach more
- Answers a 200x gap by hiring more annotators before fixing triage
- Quotes a rate measured on the uncertain band as a platform-wide figure