Why do user reports and appeal reversals, used as moderation training labels, teach a model only one side of its errors?
answer
- feedback only follows one decision
- appeals exist only where something was removed
- wrongly kept content emits the same silence
- appellants are not a random sample
- audit the approved stream to see the rest
basics
~20 sEach signal exists on only one side of the decision. An appeal can follow a removal, so reversals expose wrongly removed posts, while a correctly approved post and a wrongly approved post both generate exactly the same silence.
solid answer
~40 sImplicit signals are produced by the decision, not independently of it. An appeal is only possible where something was removed, so an appeal reversal is evidence of over-removal and nothing else. A user report points the other way but is a nomination rather than a verdict: it covers only what someone happened to see and chose to flag, which skews to popular surfaces and motivated reporters. Everything auto-approved and never reported emits no signal at all, and silence there is indistinguishable from correctness. A model trained on these alone becomes calibrated against the errors that generate complaints and blind to the ones that do not. There is a second bias inside the half you do hear: appellants are self-selected, so the reversal rate among appeals is not the error rate among removals.
go deeper
Recall that a user only appeals something that was taken down, so appeals say nothing about posts that were left up, and no news about an approved post is not good news.
Explain the coverage map: which decision each signal can follow, which error it can reveal, and why the approved stream produces identical silence whether it was handled right or wrong.
Show the operational handling: a report routes rather than labels, a reversal is recorded as an adjudicated verdict with its guideline version, and a stratified audit of approved posts supplies the side no complaint channel reaches.
State what the organisation is allowed to claim from complaint data, and fund the audit reserve that supports the claim it cannot otherwise make about content it chose to keep.
## Where each implicit signal comes from An implicit label is a by-product of someone else's action rather than a verdict commissioned by the operation. In a moderation system there are three of them, and the crucial property of all three is that **the decision itself determines whether the signal can exist**. | signal | which decision can produce it | which error it can reveal | its bias | |---|---|---|---| | a user report | a post that stayed up **and was seen** | possibly a wrongly kept post | a nomination, not a verdict; only what someone saw and bothered to flag | | an appeal | a removal only | possibly a wrongly removed post | only users motivated and able to appeal | | an appeal reversal | an appealed removal only | a confirmed wrongly removed post | doubly filtered: removed first, then appealed | | silence on an approved post | every approved post | nothing at all | absence of a signal is not evidence of a correct decision | Read the table as a coverage map. Removals generate a complaint channel by design. Approvals generate one only when a human happens to encounter the post and act, which is why the two sides of the error surface are not observed on remotely equal terms. ## The half you never hear about A post wrongly left up produces, in the overwhelming majority of cases, exactly what a post correctly left up produces: **nothing**. No appeal, because nothing was removed. Often no report, because no one saw it, or saw it and scrolled on. The result is that a training set assembled from complaints has dense, high-quality evidence of false positives and almost none of false negatives. The consequence is directional and predictable. A model trained on it learns, over cycles, to be cautious about removing and unconcerned about keeping, because only one of those two mistakes ever came back with a correction attached. The feedback loop rewards the error nobody complains about. ## Self-selection inside the half you do hear Even on the removal side the signal is not a random sample of removals. Appeals come disproportionately from people with an audience to lose, from repeat posters who know the process exists, and from users who face no language or interface barrier to filing one. So: - The **reversal rate among appeals** is not the false-positive rate among removals. It is a rate over a self-selected subset, and it can sit far above or below the true figure depending on who bothers. - The reversed posts themselves are a skewed sample of wrongly removed content, concentrated on whatever kinds of posts motivate an appeal. Both points matter when a reversal verdict is used as a training label, and they matter even more when a reversal rate is quoted as a quality figure. ## Using these signals anyway They are genuinely valuable, provided each is used at the strength it supports: - **A user report is a nomination for review, not a label.** Reports are noisy and gameable: coordinated reporting can bury a competitor, and a report often means disagreement rather than a policy violation. Treat a report as a strong prior for routing a post into the human queue, and let the verdict come from the adjudication path. - **An appeal reversal is a real verdict**, because a reviewer produced it against the written guideline. Stamp it with the guideline version and record it like any other adjudicated verdict, so it can be interpreted later. - **The missing half is bought, not collected.** A stratified random audit of the auto-approved stream, sized per policy category according to the harm of a miss, is the only source of evidence about wrongly kept content that does not depend on someone complaining. It is deliberately sampled from decisions nobody objected to. - **Keep the channels separate in the record.** A verdict from a commissioned audit, a verdict from an appeal and a report from a user are three different kinds of statement, and collapsing them into one label column destroys the ability to reason about any of them afterwards. ## What you may and may not claim With complaint-driven signals alone you can say how often removals were contested and overturned among people who contested them. You cannot say how much violating content stayed up, and you cannot say the approval decisions were sound because nobody objected. The only statement that survives scrutiny about the approved stream comes from the audit sample, which is exactly why that reserve is protected in the annotation budget even when the uncertain band is hungry for the same capacity.
- What makes the appealing population unrepresentative of all removals?Filing an appeal takes knowledge, motivation and often language or interface access. Creators with an audience, repeat posters and process-literate users appeal at far higher rates than a one-off poster who never returns. So the appealed set skews toward particular content and particular people, and the reversal rate measured on it cannot be reported as the error rate over removals as a whole.
- How do you obtain evidence about content that was wrongly left up?Commission it. Draw a stratified random sample from the auto-approved stream, sized per policy category by the harm of a miss rather than by volume, and put it through the same adjudication path against the same guideline version. Those verdicts are comparable to every other verdict in the operation, and unlike reports they exist for posts nobody complained about, which is the whole point.
- Why is a user report treated as a nomination rather than a label?Because it records that someone objected, not that the guideline was violated. Reports carry disagreement, retaliation and coordinated campaigns as well as genuine violations, and they only cover what was seen at all. Used as a routing prior they are extremely useful; used as a training label directly they teach the model to predict which posts attract complaints.
saying these in an interview costs you the question
- Treats a user report as a training label rather than a nomination for review
- Reads the reversal rate among appeals as the error rate over removals
- Assumes silence on approved content means those decisions were right
- Trains only on appealed items because they were already adjudicated
- Thinks encouraging more reports fixes the missing side of the feedback
- Collapses reports, appeals and audit verdicts into one label column