How does distant supervision turn noisy labelling functions into training labels?
answer
- write the knowledge down instead of applying it
- each rule votes or stays silent
- agreement patterns reveal which rules are good
- the labels come out probabilistic
- the trained model must beat the rules
basics
~20 sDomain experts write cheap rules that each vote a class on a row or abstain. The votes are combined, weighted by each rule's estimated accuracy, into probabilistic labels, and a model trained on them generalises beyond the rules.
solid answer
~50 sDistant or weak supervision replaces hand-labelling with programmatic labelling functions: heuristics, keyword and pattern rules, or lookups against an existing database. Each one votes a class on the rows it fires on and abstains elsewhere, so a row can receive several conflicting votes or none. You then aggregate: majority vote is the naive combiner, and the better one estimates each function's accuracy from how often the functions agree, with no ground truth, and emits a probabilistic label per row. The point of the last step is that you train a normal model on those labels, and it learns features correlated with the rules rather than the rules themselves, so it covers rows no rule fires on. Two disciplines are non-negotiable: a small hand-labelled dev set to measure each function's precision and coverage, and a separate random hand-labelled test set — you can never evaluate against the rules' own labels.
go deeper
Know the shape of the idea: instead of labelling rows one at a time, experts write cheap rules that vote a class or stay silent, and the votes are combined into training labels for a model.
Explain aggregation concretely — abstentions, overlaps and conflicts, majority vote versus accuracy-weighted combination, and probabilistic rather than hard labels. Be able to say why a model is then trained on those labels instead of shipping the rules.
Show the operating discipline: a hand-labelled dev set to measure each rule's precision and coverage, a randomly sampled hand-labelled test set kept away from the rules, pairwise overlap checks for correlated rules, and coverage measured per segment.
Own the build-versus-annotate decision and its long tail. Rules are code that needs an owner and decays, but they are versioned and re-runnable in minutes; weigh that against an annotation contract, and set who is accountable when coverage silently drops on a new document type.
## The idea When labels are expensive but domain knowledge is cheap, write the knowledge down as code instead of applying it one row at a time. Two paralegals classifying lease-contract clauses can label perhaps a few hundred clauses a day; the same two paralegals can write twelve pattern-matching rules in an afternoon that fire across the whole corpus. Each rule is far worse than a careful human — but there are twelve of them, and they cover millions of rows. A **labelling function** takes a row and returns a class or abstains. Typical sources: - pattern and keyword rules ("if the clause mentions a renewal period and a notice window, call it a renewal clause"); - lookups against an existing database or list, which is the classic *distant supervision* setup — labelling text by joining it to a knowledge base that already records the relation you want; - an existing legacy model or third-party service used as one noisy voter; - crude structural signals such as section headings or document position. ## Combining the votes Each row ends up with a vector of votes: some functions fired and agreed, some fired and conflicted, most abstained. Aggregation options, in increasing sophistication: **Majority vote.** Take the most common non-abstaining vote. Cheap and often a fine starting point. Its weakness is that it treats a 60%-accurate rule and a 95%-accurate rule as equals, and it double-counts correlated rules. **An accuracy-weighted label model.** The key insight is that you can estimate how good each function is *without any ground truth*, from the pattern of agreements and disagreements alone: a function that agrees with the consensus on the rows where it fires is probably accurate; one that disagrees systematically is probably poor or inverted. Fitting that structure gives per-function accuracy weights and produces a **probabilistic label** for each row — say 0.82 in favour of one class rather than a hard vote. Rows with no vote at all are simply not part of the weakly labelled training set. ## Why train a model on top instead of shipping the rules This is the step candidates most often miss, and it is the whole point. The rules have limited coverage and are brittle — a clause phrased differently is missed. Training a model on the rule-derived labels lets it pick up features that *correlate* with the rules' evidence, so it fires on rows no rule matched, and it smooths over individual rule errors because the probabilistic labels are noisy in different directions. The end model generally outperforms the rule set it was trained from. Training with the probabilities as soft targets, or weighting each row by label confidence, keeps the uncertain rows from being treated as certain. ## The traps **Correlated labelling functions.** If three of the twelve rules are near-copies of each other, both majority vote and a naive accuracy model treat their shared blind spot as three independent confirmations. Check overlap and conflict rates between pairs, and merge or drop the duplicates. **Coverage bias.** Rules fire where the language is conventional. Whole subpopulations — an unusual contract template, a different jurisdiction's phrasing — may receive no votes at all, so they are absent from training and the model is systematically worse there while your metrics never show it. Measure coverage per segment, not just overall. **Evaluating on the weak labels.** Scoring the model against the labels the rules produced measures agreement with the rules, not correctness, and it rewards a model that has merely memorised them. You need a randomly sampled hand-labelled test set, and it must be drawn independently of where the rules fire. **The development set is not optional.** A small hand-labelled dev set is what lets you measure each function's precision and coverage, decide whether a new rule earns its place, and detect an inverted rule. Weak supervision does not remove human labelling; it shrinks it to hundreds of rows instead of hundreds of thousands. **Rule rot.** The rules are code and they decay as documents, products or vocabularies change. Someone must own them. ## When to choose it Weak supervision suits problems where an expert can articulate rules faster than they can apply them, where the corpus is large, and where a moderate label noise level is tolerable. It is a poor fit when the task is genuinely perceptual or subjective, when no rule can be written without reading the whole document carefully, or when the required precision leaves no room for noise. Its practical advantages beyond cost are that the labelling logic is versioned, reviewable and re-runnable: change one rule and regenerate the entire training set in minutes, which no manual annotation process can match.
- How can you estimate a labelling function's accuracy with no ground truth?From agreement structure. Across rows where several functions fire, a reliable function agrees with the emerging consensus more often than chance, while a poor or inverted one disagrees systematically. Fitting that structure yields per-function accuracy weights and probabilistic labels. It only works if the functions are not all correlated — duplicated rules manufacture false consensus, so pairwise overlap and conflict rates must be checked first.
- Why train a downstream model at all rather than shipping the rule set?Coverage and generalisation. Rules fire only on the phrasings someone anticipated; a model trained on rule-derived labels learns features correlated with that evidence and so classifies rows no rule matched. It also averages over rule errors that point in different directions. That is why the end model usually beats the rules it was trained from — and why you evaluate it on hand labels, not on the rules.
- When would you spend the budget on hand labelling instead of writing rules?When the decision needs the whole document read, when it is subjective or perceptual so no rule can express it, when the required precision leaves no room for noisy labels, or when the corpus is small enough to label outright. Rules also need an owner and decay as the data changes; if nobody will maintain them, a one-off hand-labelled set is the cheaper commitment.
Twelve rough rules are twelve smoke detectors of unknown reliability. None is trustworthy alone, but the pattern of which ones go off together tells you which to believe — and you still want one real fire inspection to check the whole system.
saying these in an interview costs you the question
- Evaluates the model against the rule-generated labels
- Treats duplicated rules as independent evidence
- Skips the small hand-labelled development set
- Ships the rule set and calls it a model
- Assumes every rule is equally accurate
- Ignores segments where no rule ever fires