In a marketplace's fake-review defence, why send a middle band of risk scores to human reviewers instead of one allow/remove cutoff?
answer
- the score is not the action
- two actions, both expensive
- a cheap tier between allow and remove
- analysts hold evidence the model lacks
- band volume comes from the shift
basics
~20 sA single cutoff forces every borderline listing into an automatic allow or an automatic removal. The middle band is where the score is genuinely uncertain, so routing it to an analyst buys a cheap decision instead of a costly wrong one.
solid answer
~40 sThe score is not the decision; the decision is the action the platform emits. With one cutoff there are only two actions, so every uncertain listing gets one of the two most expensive answers: leave abuse up, or remove an honest seller. Banding adds cheaper actions in the middle — a challenge such as identity verification or a held payout, which costs no analyst time and the seller can clear themselves, and above that a queued case review, where an analyst sees evidence the feature set never had: linked accounts, reused photos, message threads. The reviewer is a bigger evidence budget, not a better classifier. The band's volume is set by what the shift can clear, with a loss floor that can override it.
go deeper
Recall that a risk score is an input to a decision, not the decision. Name the tiers in order: allow, challenge or limit, human review, automatic removal.
Explain why the mid-band exists mechanically: score overlap between honest and abusive behaviour, an analyst's larger evidence budget, and a challenge tier that scales while a staffed queue does not.
Show that a queued case is already emitting an action while it waits, and that the band's volume is bounded by the shift with an explicit default when a case reaches its deadline.
Frame the tiering as where the cost asymmetry between wrongful removals and taken loss is made operational, and name who signs off the auto-allow region below the band.
## What the score decides, and what it does not A fake-review model returns one number per listing, review or seller: an estimate of how likely it is that the content or the account behind it is manipulated. That number is **not** a decision. The decision is the **action the marketplace emits** — leave it up, make the seller do something, put a person on it, or take it down — and the mapping from score bands to actions is a separate policy layer that you choose. A single cutoff collapses that layer to two actions. Everything above the line is removed automatically and everything below it is allowed automatically, so every genuinely borderline case receives one of the two answers you would least like to give on a coin flip. ## The action tiers | Score band | Emitted action | Who pays when it is wrong | |---|---|---| | Low | Allow silently | Buyers, later, when abuse that slipped through surfaces as complaints | | Lower middle | Challenge or limit: verify identity, hold the payout, demote the listing | An honest seller absorbs friction and some lost sales | | Upper middle | Queue a case for a trust-and-safety analyst | Review labour, plus the seller waits for a verdict | | High | Automatic removal or account suspension | An honest seller loses their storefront | ## Why the middle band goes to a person - **The overlap is real.** Honest sellers with unusual patterns — a first large sale, a bulk buyer, a burst of genuine reviews after a promotion — land in the same score range as coordinated abuse. Model work moves that overlap; it does not delete it. - **An analyst sees evidence the feature set never had.** Message threads, reused product photography, a cluster of accounts sharing one payout destination, the wording of the reviews themselves. - **The two automatic errors are asymmetric.** Wrongly removing an established seller is visible, appealed, sometimes regulated, and effectively permanent because the buyers move on. Leaving one fraudulent listing up costs a bounded amount for a bounded time. - **A few minutes of analyst time is cheap** relative to either wrong automatic answer on a high-value account. ## The challenge tier is not the review tier A challenge costs no analyst time and the seller clears it themselves, which makes it the right action for the lower half of the uncertain range — elevated score, not worth a person. Two consequences follow: - Widening the challenge tier is the cheapest way to absorb extra risk in a spike, because it scales with traffic while a staffed queue does not. - A challenge an honest seller cannot complete — no documents, a payout account legitimately in another name — quietly becomes a decline. Track completion rate per tier, or the friction tier hides false declines that never appear in the removal numbers. ## What the review tier costs - **Labour**, roughly fixed per shift, which does not scale with a traffic spike. - **Time to decision.** A queued case is in limbo: the listing is usually still live, or throttled, while it waits. A slow queue is itself a decision to allow or to hold. - **A hard ceiling.** The queue absorbs only what the shift can work; past that, cases reach their deadline and fall to whatever the default action is. That default has to be chosen in advance, not discovered during a sale. ## How the band gets its width 1. Start from what the shift can clear in a day, not from the model. 2. Take that many of the day's flagged candidates, highest score first. 3. Read the score at the bottom of that set: that is the band's lower bound today. 4. Confirm the business has accepted the expected loss of the auto-allow region below it, and keep a loss floor that forces a review even on a full day. The width is an operating parameter that moves with staffing and with flag volume; it is not a verdict on model quality. Narrowing it is not automatically an improvement — it simply moves cases from people to one of the two automatic answers. ## What to watch once the tiers are live - The share of the day's candidates landing in each tier, tracked daily: a base-rate move shows up here first. - The proportion of reviewed cases that end in a takedown, per slice of the band. If the top of the review band almost always ends in takedown, that slice is being reviewed for nothing and the automatic line can move down over it. - Time to decision at p50 and p95 against the target. - Appeal and reinstatement rate out of the automatic removal tier — the only direct read you have on wrongful removals.
- What does the challenge tier give you that the review tier cannot?It scales with traffic. A verification prompt or a held payout costs no analyst time and the seller resolves it themselves, so the tier absorbs a volume spike that a staffed queue cannot. The price is friction on honest sellers, and a challenge an honest seller cannot complete is a decline in disguise — so completion rate has to be measured per tier.
- If the model improves, should the review band get narrower?Not automatically. A better model reorders the candidates, so the same review volume catches more real abuse — that is the gain. Narrowing the band is a separate decision to spend less labour and accept more automatic answers; make it because the top of the band now almost always ends in takedown, not because the model got better in general.
- A listing sits in the queue for three days. Has a decision been made?Yes, in practice. While the case waits, the listing is live, demoted or held, and that state is the action the platform is emitting. Time to decision is therefore part of the policy, not an operational detail, and every case needs a deadline with a pre-agreed action when it expires.
saying these in an interview costs you the question
- Assuming a good enough model makes one cutoff sufficient
- Treating the uncertain middle band as a defect to engineer away
- Believing a removal and a wrongful removal cost about the same
- Reaching for removal in the uncertain band because it is the tidier state
- Confusing an automatic challenge with a staffed analyst review
- Assuming a queued case has had no action taken on it