In a moderation annotation queue, what do seeded gold questions with known answers measure that an agreement score cannot?
answer
- known answer hidden in the normal queue
- agreement is relative, gold is absolute
- scores a person, not an item
- a whole pool can agree and be wrong
- stratify gold across policy categories
basics
~20 sGold questions score one annotator against a verdict already known to be right. An agreement score only compares annotators with each other, so a pool that has all drifted onto the same misreading still scores as agreeing.
solid answer
~40 sA gold question is an ordinary-looking post whose correct verdict under the current guideline was already settled by the adjudication path, dropped into the normal queue so the annotator cannot tell it from real work. It yields a per-annotator accuracy against a recorded answer, which agreement structurally cannot: agreement is relative, so a pool that has all absorbed the same wrong reading of a clause looks healthy while producing a uniformly wrong training set. Operationally the gold stream drives three decisions: qualify a new annotator before their verdicts count, hold a rolling score that triggers coaching or retirement, and accept or re-queue a batch on its gold accuracy. The gold set has to be stratified across policy categories and re-derived after a guideline edit, or it only measures the easy half of the work.
go deeper
Recall that some items in the queue already have a known correct verdict, that they look exactly like real work, and that they exist to score the person rather than the post.
Explain why a relative measure cannot catch a pool that drifted together, and name the three decisions a known-answer score drives: qualification, a rolling accuracy, and accepting or re-queuing a batch.
Demonstrate the operational care around it: stratify gold per category, re-derive answers after a policy edit, keep the pool large enough to resist memorisation, and turn a retirement into a re-review of that person's past verdicts.
Weigh the overhead. Gold slots produce no new labels, so the rate is a quality-versus-volume trade made per category against the cost of shipping a bad batch into the training set.
## What a gold question is A **gold question** (equivalently a known-answer item) is a real post whose correct verdict under the **current** written guideline has already been decided by the adjudication path and recorded. It is injected into an annotator's ordinary queue at a low rate and must be indistinguishable from real work: same surface, same interface, same absence of any marker. The moment an annotator can recognise one, the measurement changes from *how accurately do you work* to *how carefully do you work when watched*, which is not the number the operation needs. ## Why an agreement score cannot do this job Agreement is a **relative** measure. It says how often annotators land on the same verdict as one another, and it is silent about whether that shared verdict is the one the guideline demands. This matters because the most common systematic failure of a labeling operation is not random carelessness, it is a pool trained together or calibrated by the same shift lead converging on one misreading of a clause. | what you need to know | agreement across annotators | seeded known-answer items | |---|---|---| | is this particular item contested | yes | no | | is this particular annotator accurate | no, only how often they match others | yes, against a recorded verdict | | has the whole pool drifted the same way | no, a shared misreading scores as agreement | yes, accuracy falls while agreement holds | | can a brand-new annotator be trusted yet | only once enough overlap has accumulated | yes, from their first shift | The two measures are complements, not alternatives. Agreement finds the **items** the guideline does not resolve; gold finds the **people**, and the pool-wide drift, that agreement is blind to. ## The three decisions a gold score drives 1. **Qualification.** A new annotator works a queue with a raised gold rate and their verdicts are held out of the training set until their accuracy clears a stated bar per category. This is cheap compared with discovering months later that a cohort's verdicts were wrong. 2. **A rolling per-annotator score.** Accuracy is tracked as a moving figure, not a one-off exam. Falling below the bar triggers coaching first, then removal from that category, then retirement from the pool. When someone is retired, the verdicts they already produced become a re-review candidate list, not a silent inheritance. 3. **Batch acceptance.** A batch of labels whose gold accuracy came in under the bar is re-queued rather than shipped to the training set. This is the only gate that catches a bad shift before the labels are consumed downstream. ## Where gold questions fail - **Poor stratification.** One overall accuracy number across every policy category hides a person who is excellent on the common categories and unreliable on a rare one. Seed gold per category, and report per category. - **Staleness after a guideline edit.** A gold item's recorded verdict is a verdict *under a particular guideline version*. Edit the policy and some of those answers are now wrong, and scoring against them retires exactly the annotators who followed the new rule. Re-derive the affected gold items before the edited guideline goes live. - **Memorisation.** A small gold pool recycled often becomes recognisable by content alone. The set has to be large enough, and refreshed, that seeing one twice is rare. - **Sampling only the easy cases.** If gold items are drawn only from posts the pool decided unanimously, the set measures the work nobody was going to get wrong. Deliberately seed adjudicated hard cases too, at a known mix. ## What it costs The gold rate is pure overhead on annotation spend: those verdicts produce no new labels, because the answer is already held. At a rate of about 5 percent, an operation that hand-labels roughly 18,000 posts a week spends about 900 of those slots re-deciding items it has already decided. That is the price of knowing whether the other 17,100 are worth anything, and it is normally the cheapest quality control in the operation, because the alternative is discovering a bad cohort only when the model trained on their verdicts reaches production. ## The limit worth stating out loud A gold score bounds an annotator's accuracy **on the kinds of items the gold set contains**. Someone who is right everywhere except one narrow slice can pass a headline gold score comfortably, which is why per-category reporting and a periodic re-review of already-accepted verdicts both stay in the operation even when the headline number looks fine.
- Can gold questions catch an annotator who labels correctly everywhere except one narrow policy class?Only if the gold set covers that class. A single headline accuracy averages the narrow slice away, so someone accurate on the common categories passes while systematically mislabelling one. The controls are gold stratified per category with per-category reporting, plus a periodic stratified re-review of verdicts that were already accepted, which is also how an annotator who is wrong on purpose in one narrow place is found.
- What has to happen to the gold set when the written guideline is edited?Its recorded answers are verdicts under the previous version, so every gold item touching an edited clause is re-derived through the adjudication path before scoring resumes. Skipping that step scores annotators against the old policy and penalises exactly the ones who adopted the new one, which reads as a sudden pool-wide accuracy drop with no real cause.
- What rate of gold items is right?High enough that a rolling per-annotator accuracy is meaningful within a review period, low enough that the overhead is tolerable, and raised for new annotators and for categories that were recently edited. It is set per category rather than globally, because a rare category needs a higher rate to produce any usable score at all.
saying these in an interview costs you the question
- Assumes high agreement between annotators proves the labels match the guideline
- Marks gold items visibly, so they are worked more carefully than real posts
- Scores each annotator on one overall accuracy across every policy category
- Keeps the same gold answers unchanged after the guideline is edited
- Recycles a small gold pool until annotators recognise the items
- Draws gold only from unanimous posts, measuring the easy half