skip to content

An article can carry any of 20 topic tags at once - how should you frame that target?

level: middleimportance: must knowfreq 66%

answer

  1. ask whether two labels can co-occur
  2. radio buttons or checkboxes?
  3. probabilities need not sum to one
  4. one yes/no decision per tag
  5. combination-as-class explodes past a million

basics

~20 s

That is a multilabel target: the tags are not mutually exclusive, so an article can carry zero, one or several. Frame it as 20 independent yes/no decisions with a probability per tag, not one distribution over 20 tags summing to one.

solid answer

~50 s

The deciding test is mutual exclusivity: can two tags be true of the same article at the same time? Here they can, so it is multilabel, not multiclass. Practically that changes the shape of the output. A mutually exclusive framing produces probabilities across the tags that sum to one and forces exactly one winner; a multilabel framing produces 20 probabilities that need not sum to anything, each answering `does this tag apply`, and each with its own decision rule and its own base rate. The default approach is one binary decision per tag, sometimes called binary relevance. The tempting alternative - making each observed *combination* of tags a class - explodes: 20 tags give over a million possible combinations, and the model can never predict a combination it has not seen. Independent per-tag decisions ignore correlations between tags, which you can recover later if it matters.

go deeper

for a junior

Be ready to state the mutual-exclusivity test in one sentence and apply it: if two labels can be true at once, the task is multilabel. Know that multilabel output is one probability per label, not one distribution over labels.

for a middle

Explain what the summing-to-one constraint does mechanically - classes compete for probability mass - and why that is wrong for co-occurring tags. Be able to say why enumerating tag combinations as classes does not scale past a handful of labels.

for a senior

Show you would look at how the labels were produced before trusting absent tags as negatives, and that you set decision rules per tag because base rates differ by two orders of magnitude. Be ready with a cheap serving-time fix for hierarchy violations.

for a principal

Own the target-design tradeoff: every tag in the schema is an ongoing annotation, monitoring and consistency cost, so the number of labels is a product decision rather than a modelling one. Be ready to defend shrinking a tag taxonomy that nobody acts on.

## Naming the task correctly Four target shapes get confused with each other: - **Binary** - one label, two outcomes. Fraud or not. - **Multiclass** - one label, more than two outcomes, **mutually exclusive**. Exactly one is true per example: a photo shows a cat *or* a dog *or* a horse. - **Multilabel** - many labels, each independently true or false. An article may be tagged `politics`, `election` and `europe` simultaneously - or nothing at all. - **Multi-output** - several separate targets predicted together, each of which may itself be binary, multiclass or continuous. The test is one sentence: **can two of these labels be true of the same example at the same time?** If yes, it is multilabel. Twenty news topics obviously can co-occur, so the tagging task is multilabel. ## What the framing changes about the output A mutually exclusive framing produces one probability vector *across* the classes that sums to one. Raising the probability of one class necessarily lowers another - they compete. That competition is exactly right when only one answer can be true, and exactly wrong here: evidence for `election` should not have to be taken away from `politics`. A multilabel framing produces one probability *per tag*, each in the range 0 to 1, with no summation constraint. Twenty tags means twenty independent yes/no questions. Consequences follow immediately: - **Output cardinality is free.** Zero tags and eight tags are both representable. You must decide as a product question whether an article is allowed to end up with no tags, and what to do when that happens - typically fall back to the highest-scoring tag, or leave it untagged for a human. - **Each tag has its own base rate.** `politics` may apply to 30% of articles and `obituaries` to 0.4%. A single shared decision rule across all 20 will over-fire on the rare tags or under-fire on the common ones, so the cut point is per tag. - **Each tag has its own difficulty.** Some are lexically obvious, some are judgement calls. Per-tag reporting is the only way to see this; a single aggregate number hides which tags are broken. ## Framings that do not work, and why **Label powerset.** Treat every observed combination of tags as its own class. It is a legitimate trick for a handful of labels because it captures correlations exactly, but with 20 tags there are 2^20 - over a million - possible combinations. Almost all of them never appear in training, and the model can only ever predict a combination it has seen. Most observed combinations will have a handful of examples each. It does not scale. **Forcing one primary tag.** Pick the single most important tag per article and run a mutually exclusive model. This is sometimes imposed by a downstream system that has one slot. The costs are real: annotators disagree about which tag is *primary*, so the labels themselves become noisy and inconsistent; the model learns an arbitrary priority ordering rather than topicality; and an article genuinely about both politics and the economy is recorded as an error whichever tag you predict. If a single slot is truly required, the cleaner design is to model all tags honestly and pick the highest-scoring one at serving time - the choice of primary tag then becomes a display rule you can change without retraining. ## Correlation between tags Independent per-tag decisions assume the tags are conditionally independent given the features, which they are not - `election` almost implies `politics`. Two common repairs: chain the decisions, feeding earlier tags' predictions as inputs to later ones, or post-process with the label hierarchy (if `election` fires, force `politics`). Both add complexity, so start with independent decisions and only add structure if the errors show you inconsistent tag sets. ## The subtlety worth raising: missing versus negative In multilabel data an absent tag is ambiguous. Did the annotator judge that `europe` does not apply, or did they stop after the first two tags they thought of? Independent binary framing treats every absent tag as a hard negative, which teaches the model that under-tagged articles genuinely lack those topics. Where labelling was partial, that pushes recall down systematically. Knowing how the labels were produced is part of framing the target - a strong answer says so. ## Hierarchy If some tags nest inside others, the target has structure the flat framing ignores. You can flatten and rely on the model to learn consistency, enforce parents at serving time, or model each level separately. Whichever you pick, decide it deliberately rather than discovering inconsistent parent/child predictions in production.

  • The downstream system has exactly one tag slot per article - what do you lose by training on a single primary tag?
    Label quality first: annotators disagree on which of several correct tags is primary, so the target carries noise that has nothing to do with topic. You also make co-tagged articles unwinnable - either tag counts as an error. Better to model all 20 honestly and choose the highest-scoring tag at serving time, which keeps the slot rule a display decision you can change without retraining.
  • How would you stop a multilabel model predicting 'election' without also predicting 'politics'?
    Independent per-tag decisions have no mechanism to enforce that, so either add structure or fix it downstream. Chaining feeds the parent tag's prediction into the child's inputs; the simpler route is a serving-time rule that switches on any ancestor tag whenever a child fires. Start with the rule, and only pay for a chained model if the inconsistencies are widespread.
  • Under a multilabel framing an article ends up with no tag above its cut point - what do you do?
    First accept that zero tags is a legitimate output of the framing, unlike the mutually exclusive case which always names a winner. Then choose a product rule: emit the top-scoring tag anyway, route the article to a human, or leave it untagged. Track how often it happens, because a rising untagged rate usually means new topics have appeared that the tag set does not cover.
  • When is it correct to collapse 20 tags into a binary target instead?
    When the action is binary. If the only consumer is a filter for one section of the site, model `belongs to this section or not` and skip the other 19 tags. Modelling distinctions nobody acts on costs data, annotation effort and monitoring surface with no return, and each unused tag is another output that can silently break.

Multiclass is a set of radio buttons - selecting one clears the rest. Multilabel is a row of checkboxes: any number can be ticked, including none.

saying these in an interview costs you the question

  • Assumes every classification target has mutually exclusive classes
  • Uses one distribution summing to one for tags that co-occur
  • Enumerates every observed tag combination as its own class
  • Applies one shared cut point across tags with very different base rates
  • Treats an absent tag as a hard negative without checking labelling practice

context