A trust-and-safety team can work 600 seller-abuse cases a day; how do you place the review band's lower bound?
answer
- capacity first, score second
- productive hours, not shift length
- leave headroom under full utilisation
- percentile of yesterday, recomputed daily
- concurrency equals arrival rate times wait
basics
~20 sWork backwards from capacity. Pick the daily case volume the shift can sustain, take that many of the day's flagged candidates highest score first, and read the score at the bottom of that set as the bound, recomputed daily.
solid answer
~40 sThe model gives an ordering of candidates; the shift tells you how far down it you can afford to go. Ten analysts at 7.5 productive hours and 8 cases an hour is 600 nominal, but plan the standing band at about 80% — 480 a day — because arrivals are bursty and a queue run at full utilisation pays for the last few per cent in time to decision. With 9,000 scored candidates a day, 480 is the top 5.3%, so the bound is roughly the 95th percentile of yesterday's score distribution. Recompute it daily rather than freezing a score. Keep an expected-loss floor that forces a case into review even on a full day.
code
pseudocode · 20 linesanalystsOnShift = 10
productiveHoursEach = 7.5
casesPerAnalystHour = 8
targetUtilisation = 0.80
nominalCapacityPerDay = analystsOnShift * productiveHoursEach * casesPerAnalystHour // 600
plannedReviewVolume = nominalCapacityPerDay * targetUtilisation // 480
scoredCandidatesToday = 9000
reviewFraction = plannedReviewVolume / scoredCandidatesToday // 0.053
bandLowerBound = scoreAtPercentile(yesterdayScores, 1 - reviewFraction) // ~95th pct
// Little's Law sanity check, in analysts
arrivalsPerHour = plannedReviewVolume / productiveHoursEach // 64
serviceHours = 1 / casesPerAnalystHour // 0.125
busyAnalysts = arrivalsPerHour * serviceHours // 8 of 10
for each candidate in today:
if candidate.score >= bandLowerBound or expectedLoss(candidate) >= lossFloor:
enqueueForReview(candidate)go deeper
Recall the direction of the calculation: the shift's daily case count comes first, and the score bound is read off the bottom of that many candidates.
Do the arithmetic out loud — analysts x productive hours x cases per hour, a utilisation target, then a percentile of the day's scores — and check it with Little's Law.
Defend the headroom and the exception path: bursty arrivals, uneven case sizes, appeals eating the same shift, and a loss floor that overrides volume on the worst cases.
Treat the bound as the joint staffing and risk decision it is, and say what you would buy first — headcount, throughput tooling, or a wider automated challenge tier.
## The bound is a score chosen from a headcount The review band's lower bound is expressed as a score, but nothing about the model chooses it. The model supplies an **ordering** of the day's flagged candidates; the shift supplies how far down that ordering you can afford to go. So you never pick a score and see how many cases it produces — you pick a volume you can actually work and read the score off the bottom of it. ## The capacity arithmetic 1. **Nominal capacity** = analysts on shift x productive hours x cases per analyst-hour. Ten analysts, 7.5 productive hours each, 8 cases an hour = **600 cases a day**. "Productive" is not shift length: handover, breaks, escalations and training are already removed. 2. **Planned volume** = nominal x target utilisation. At 80% that is **480 cases a day** standing. 3. **Review fraction** = planned volume / scored candidates. With 9,000 candidates above the cheap screen, 480 / 9,000 = **5.3%**. 4. **Lower bound** = the score at the **94.7th percentile** — call it the 95th — of yesterday's score distribution, recomputed each day rather than frozen as a constant. ## Why 80% and not 100% - **Arrivals are bursty.** Flags do not trickle in at a constant rate; waiting time climbs steeply as utilisation approaches one, so the last few per cent of nominal capacity are paid for in time to decision, not in cases closed. - **Cases are not equal.** A linked-account cluster is one queue entry and an hour of work; the 8-an-hour figure is an average over a mix that changes. - **The same people absorb the other work**: appeals, re-reviews, escalations and absence all come out of the same 600. ## Little's Law as the sanity check Little's Law says **concurrency = arrival rate x time in system**. Run it both ways on these numbers: - 480 cases spread over 7.5 productive hours is **64 an hour**. A case takes 1/8 hour = **7.5 minutes** of analyst time, so the offered load is 64 x 0.125 = **8 analysts busy on average out of 10** — the 80% utilisation again, arrived at from the other direction. - With a **4-hour** time-to-decision target, expect about 64 x 4 = **256 cases open** at any moment. If the queue dashboard shows thousands open, either the bound is too low for the staffing or the cases-per-hour figure is fiction. ## What moves the bound | Input that changes | Effect on the bound | Why | |---|---|---| | Flag volume rises | Bound rises | The same 480 is a smaller slice of a bigger pool | | Headcount falls | Bound rises | Fewer cases fit in the planned volume | | Cases per hour rises (better evidence tooling) | Bound falls | The same shift reaches further down the ordering | | The model improves | Bound roughly unchanged | Reordering changes **which** cases fill the band, not how many | | Time-to-decision target tightens | Bound rises | Less work in progress is allowed, so a smaller standing band | That last row is the one candidates miss: a better model does not buy a wider band, it buys a better-filled one. ## The floor that capacity must not override Sizing from headcount alone silently auto-allows the worst case on the worst day: if a single account has a large pending payout and a high score, it should reach a person even when the day is full. Keep an **expected-loss floor** — an absolute level above which a case enters review regardless of volume — and let breaching it trigger surge (overtime, borrowed reviewers, a temporarily wider challenge tier) rather than a quiet fall-through. Capacity sets the routine volume; the floor sets the exception. ## Recomputing, and what to log Recompute the percentile daily from the previous day's scores so the band tracks the flag pool instead of drifting with it. Log, per day: the bound in score terms, the realised volume, cases closed, time to decision at p50 and p95, and how many cases hit their deadline unworked. Those five numbers tell you whether the band is sized to the shift or to wishful thinking.
- Why express the bound as a daily percentile rather than a fixed score?Because the flag pool moves. A frozen score meets a growing pool by sending more cases than the shift can work, and a shrinking pool by leaving analysts idle. A percentile recomputed from the previous day's scores holds the volume steady, which is the quantity staffing is actually bounded by. Pair it with a loss floor so the worst cases are not rationed.
- Your dashboard shows 3,000 cases open against a 4-hour target. What does that mean?Work in progress is far above what the arrival rate and the target imply — about 256 at 64 an hour. So either intake exceeds what the shift clears and a backlog has been accumulating, or the assumed cases-per-hour is wrong. Check closed-per-day against nominal capacity before touching the bound; if intake genuinely exceeds capacity, the bound is too low.
- The team buys tooling that raises throughput from 8 to 11 cases an hour. What changes?Planned volume rises from 480 to about 660 at the same 80% utilisation, so the band's lower bound falls to roughly the 93rd percentile of a 9,000-candidate pool and more mid-band cases reach a person. Time to decision improves too, because the same arrivals now face more service capacity. Nothing about the model or the score distribution changed.
saying these in an interview costs you the question
- Picking a score first and discovering the volume later
- Planning the standing band at 100% of nominal capacity
- Using shift length instead of productive hours per analyst
- Freezing the bound as a constant score across seasons
- Assuming a better model lets the shift review more cases
- Sizing from headcount with no floor for high-loss cases