skip to content

How does graded relevance change a search-ranking task compared with binary relevant-or-not labels?

level: middleimportance: nice to knowfreq 34%

answer

  1. a set to retrieve, or an order to produce
  2. grades attach to the pair, not the page
  3. the interior boundaries are the fuzzy ones
  4. collapsing to binary throws information away

basics

~20 s

Graded labels make the target an ordering among relevant items rather than a set of hits. With perfect, good, marginal and bad grades, ranking a marginal result above a perfect one is an error binary labels cannot express.

solid answer

~50 s

Suppose site-search results are judged perfect, good, marginal or bad instead of just relevant or not. Three things change in the task definition. First, the target is no longer a set to retrieve but an order to produce: among items that are all "relevant" under a binary scheme, the grades say which must come first. Second, the labelled unit is the query-document pair, not the document — the same page is perfect for one query and bad for another, so relevance is contextual by construction. Third, the annotation guideline becomes part of the specification: the boundary between good and marginal is a product decision, and judges will disagree there more than they do on relevant versus irrelevant. If you later collapse grades to a binary target to train a simpler model, be explicit that you have thrown away the distinction the grades existed to capture.

go deeper

for a junior

Know that relevance can be judged on a scale rather than yes/no, and that a grade describes how well a result answers a specific query, not how good the page is in general.

for a middle

Explain that grades turn the target from a set into an ordering, that the labelled unit is the query-document pair, and what is lost when a four-level ladder is collapsed to a binary target.

for a senior

Show judgement about the labelling operation itself: which grade boundaries judges can actually apply, how to split evaluation data by query, and when a clean binary scheme beats a noisy graded one.

for a principal

Own the decision of how many grades the organisation should maintain, given judging cost, guideline enforcement and the product promise, and be able to defend collapsing the ladder as a deliberate simplification.

## Two different specifications of the same task A binary relevance scheme says each result either belongs in the answer set or does not. A graded scheme assigns a level — a common four-level ladder is **perfect / good / marginal / bad** — and that small change rewrites the task. ### What the model is asked to produce Under binary labels, the task reads as *retrieve the relevant ones*. Any order among the relevant items is equally acceptable, because the labels contain no information distinguishing them. Under graded labels the task reads as *order the results so that higher grades appear before lower ones*. An ordering that places a marginal result above a perfect one is now a defect the labels can describe. This is why graded relevance is the natural home of ranking framing: the label set itself encodes an ordering, so "correct" is a property of the sequence. ### The labelled unit is the pair, not the item A subtle framing point that interviewers probe: a grade attaches to a **(query, document) pair**. A returns-policy page is *perfect* for "how do I return an item" and *bad* for "track my order". Nothing about the document alone determines its grade. Consequences: - Features must describe the match, not just the document — otherwise the model can only learn a global quality prior for pages. - Splitting data for evaluation should generally be done **by query**, so that judgements for the same query do not straddle training and evaluation. - The number of labels you need scales with queries times results judged per query, not with documents. ### The guideline becomes the spec With two labels, judges argue over one boundary. With four, they argue over three, and the interior boundaries are the fuzzy ones. Practical effects: - **Agreement falls as grades multiply.** Distinguishing bad from everything else is easy; distinguishing good from marginal is a judgement call, and disagreement there is not noise to be averaged away — it means the guideline is under-specified. - **Grades are ordinal, not numeric.** Perfect is better than good, but it is not "twice as good". Any mapping of grades to numbers is a modelling choice you are making, not something the labels handed you. - **The grade distribution is usually skewed.** Most judged results for a head query are decent; the perfect grade is rare. That rarity is what makes the top of the list hard. ### Collapsing grades Teams often collapse the ladder — for example treating perfect and good as positive and the rest as negative — so they can train a simpler binary model. That is legitimate, but it is a deliberate loss: - Every ordering that keeps positives above negatives becomes equally good, so the model has no reason to lift perfect results above merely good ones. - Where you draw the collapse line silently redefines the product. Folding *marginal* into the positive side tells the model that a barely acceptable result belongs at the top. - The evaluation must match the framing you shipped. If the product promise is "the best answer first", a measure that only counts hits cannot see whether you kept that promise. ### When binary is the honest choice Graded relevance is not automatically better. Prefer binary when the downstream use really is set-shaped (does this result belong in a fixed answer box at all?), when you cannot afford to write and enforce a four-level guideline, or when judge agreement on the interior grades is so low that the extra levels are mostly noise. A clean binary label set beats a graded one that three judges fill in three different ways. ### Where graded labels come from Two sources, with different failure modes. **Human judges** following a guideline give you grades that reflect the product's intent, at a cost per judgement and with the agreement problem above. **Behavioural signals** such as clicks or dwell are cheap and plentiful, but they are observations of what users did with the ordering you already showed them, so they carry the incumbent ranking's bias and are not a neutral measurement of relevance. Many teams use judged grades as the specification and behaviour as a volume supplement, being explicit about which is which. ### Interview answer shape Say what changes in the *task*: the target becomes an ordering rather than a set; the unit of labelling is the query-document pair; the guideline that defines the middle grades is now part of the product spec; and collapsing grades is a real information loss you should name out loud rather than perform silently.

  • Two judges disagree constantly between good and marginal. What do you do?
    Treat it as a specification problem first, not a labelling problem. Measure agreement per boundary; if the good/marginal line is the one that fails, the guideline does not define it well enough, so I would add worked examples for that boundary and re-judge a sample. If agreement stays low after that, the honest move is to merge the two grades rather than train on a distinction nobody can apply consistently.
  • Can you use clicks instead of paying judges for grades?
    Clicks are cheap and abundant but they are not the same measurement. Users only interact with the ordering you already showed them, so click data reflects position and the incumbent ranking as much as it reflects relevance. I would treat judged grades as the specification of what good means and behavioural data as supplementary volume, and never quote a click-derived number as if it were an unbiased relevance judgement.

saying these in an interview costs you the question

  • Says grades are just relevance scores with more decimals
  • Treats grades as numeric so perfect equals twice good
  • Labels documents rather than query-document pairs
  • Collapses to binary without naming what was lost
  • Assumes more grades always means better labels

context