skip to content

Learning Without Full Labels

Self-training and pseudo-labelling, pretext tasks that invent labels, distant supervision from rules, and active learning. Interviewers probe it because labelling is the real project cost.

on this pageshow

questions

4

Why can unlabelled data improve a classifier in semi-supervised learning?

level: juniorimportance: must knowfreq 55%

answer

  1. unlabelled rows carry no targets
  2. they describe the inputs, not the answer
  3. clumps, and the gaps between clumps
  4. boundary belongs in low density
  5. wrong assumption makes accuracy worse

basics

~20 s

Unlabelled data helps only when its shape reveals where the classes separate. Semi-supervised methods assume points in the same dense cluster share a label, so the boundary belongs in a low-density gap. When that assumption is false, unlabelled data hurts.

solid answer

~50 s

Unlabelled rows carry no target values, so they can only help by showing where the input distribution is dense and where it is empty. Semi-supervised methods lean on one of three related assumptions: the cluster assumption (points sitting in the same high-density region tend to share a label), its low-density separation form (the decision boundary should fall in a sparse gap rather than through a crowd of points), and the manifold assumption (the data lies on a lower-dimensional surface along which the label changes smoothly). When one of those holds, a few hundred labelled rows fix *which* side is which while millions of unlabelled rows sharpen *where* the boundary sits. When none holds — classes genuinely overlap, or the true boundary runs straight through a dense region — semi-supervised training can score worse than training on the labelled rows alone, so I always keep the supervised-only baseline to compare against.

go deeper

for a junior

Be ready to say in one breath that unlabelled data describes the inputs, not the answer, and that it helps only when points in the same dense region tend to share a label. Naming the cluster assumption out loud is what the screen is checking for.

for a middle

Explain the mechanism: unlabelled points locate the low-density gap, which reduces the variance of a boundary fitted from few labels. Be able to state the manifold assumption too, and to say why the gain shrinks as the labelled set grows.

for a senior

Show the discipline of treating semi-supervised learning as a hypothesis under test: supervised-only baseline, a randomly sampled human-labelled test set, and a check that the labelled and unlabelled pools come from the same distribution before you trust any gain.

for a principal

Own the decision of whether to invest in a semi-supervised programme at all. Argue it against simply buying more labels, given that the gain is largest when labels are scarcest and evaporates as the labelled set grows, and set the evidence bar the team must clear before shipping one.

## The setting Semi-supervised learning is the regime where you have a small labelled set and a much larger pool of inputs with no targets at all: 8,000 labelled rows and 2,000,000 unlabelled ones is a typical shape. The question every interviewer is really asking is whether those 2,000,000 rows contain any information about the target — because on the face of it they contain none. Nobody wrote a label on them. The resolution is that unlabelled data tells you about `p(x)`, the distribution of the inputs, and says nothing directly about `p(y|x)`, the thing you are trying to learn. Unlabelled data helps exactly when knowing the shape of `p(x)` constrains `p(y|x)`. If those two are unrelated, extra unlabelled rows are decoration. ## The three assumptions that link them **The cluster assumption.** Inputs are not spread uniformly; they clump. The assumption says points inside the same clump usually carry the same label. If it holds, a handful of labels per clump is enough to colour the whole clump, and the unlabelled points define where each clump begins and ends. **Low-density separation.** This is the same idea stated about the boundary instead of the clumps: the decision surface should pass through sparse regions, not through crowds. Given a choice of two boundaries that both classify your labelled rows correctly, prefer the one that slices through empty space. Margin-style semi-supervised methods encode exactly this — they push the boundary away from unlabelled points. **The manifold assumption.** High-dimensional data often lies near a much lower-dimensional surface, and the label varies smoothly as you move along that surface. Two points that are far apart in raw distance but connected by a dense chain of unlabelled points are assumed to share a label. Graph-based semi-supervised methods propagate labels along exactly such chains. All three are geometric bets about the world. None of them is a theorem, and each is falsifiable on your dataset. ## What it looks like when the bet pays off Suppose two classes really do form separated blobs, and your labelled sample is small enough that the supervised boundary is wobbly — a different random draw of 200 labelled rows would give a visibly different line. The unlabelled points do not tell you which blob is which, but they show you exactly where the empty corridor between them is. Nailing the boundary to that corridor removes most of the variance the small labelled sample introduced. The gain is a variance reduction, not new label information, which is why the benefit shrinks as the labelled set grows: with enough labels the supervised boundary is already in the right place. ## What it looks like when the bet fails Three failure shapes recur. 1. **Overlapping classes.** If the two classes genuinely occupy the same dense region — plenty of real problems are like this — there is no low-density corridor to find, and forcing the boundary into a sparse area moves it away from the truth. 2. **A boundary that runs through a crowd.** The dense region may itself straddle the decision surface; the cheapest example is a class that is defined by a subtle feature while the input density is dominated by an unrelated one. 3. **A shifted unlabelled pool.** The unlabelled rows are usually cheap precisely because they were collected differently — a different time window, a different market, a different traffic source. Then the geometry you are fitting is not the geometry your test data has. This is not a theoretical worry. Semi-supervised methods are known to *degrade* accuracy relative to the supervised baseline when their assumption is violated, and a candidate who says "more data always helps" has missed the entire point of the question. ## How to check rather than assume The discipline is simple and interviewers look for it: - Always train the **supervised-only baseline** on the same labelled rows and compare. Semi-supervised is a hypothesis, not a free upgrade. - Evaluate on a **randomly sampled, human-labelled test set** that no semi-supervised step ever touched. - Plot the gain against labelled-set size. A method that helps at 200 labels and does nothing at 20,000 is behaving exactly as theory predicts; a method that helps everywhere is suspicious. - Sanity-check that the labelled and unlabelled pools look alike — if a simple classifier can tell labelled rows from unlabelled ones, you have shift, and the assumption is already in trouble. ## The one-line answer Unlabelled data does not add labels; it adds shape. It pays off when the shape of the inputs and the location of the class boundary are related — and it costs you when they are not.

  • Co-training replaces the cluster assumption with a different one — what does it require?
    Co-training needs two views of each example that are individually sufficient for the task and roughly conditionally independent given the label. Classifying web pages on body text and on inbound anchor text is the classic pairing: each classifier labels the examples it is most confident about and hands them to the other, so the second view supplies genuinely new evidence. If the two views are near-copies of each other, co-training collapses into ordinary self-training and stops helping.
  • How would you prove that the unlabelled data actually helped, rather than assuming it?
    Train the same model family twice on the identical labelled rows — once supervised only, once with the unlabelled pool — and score both on a randomly sampled human-labelled test set that neither run touched. Repeat over several labelled-set sizes and several seeds, because the semi-supervised gain is mostly a variance reduction and it can vanish or invert as labels accumulate.
  • Your unlabelled pool was scraped from a different traffic source than the labelled rows. Why does that matter?
    Every semi-supervised assumption is about the geometry of the input distribution, so it only transfers if the unlabelled pool has the same geometry as the data you will predict on. A shifted pool teaches the boundary to sit in the wrong gap. Check it by training a small classifier to tell labelled from unlabelled rows: if it succeeds easily, the pools differ and the assumption is on shaky ground.

Unlabelled points are like a coastline seen at night: they show you where land is dense and where the open water lies, but not which country you are standing in. Semi-supervised learning bets that the border runs through the water.

saying these in an interview costs you the question

  • Claims more data always helps, labels or not
  • Says unlabelled rows add label information
  • Never runs a supervised-only baseline for comparison
  • Confuses semi-supervised learning with unsupervised clustering
  • Assumes clusters exist without checking the data
  • Ignores that the unlabelled pool may be differently distributed

context

open as a page

In self-training, how do a model's own pseudo-labels reinforce its errors?

level: middleimportance: must knowfreq 58%

basics

~20 s

Self-training adds the model's own confident predictions to its training set as ground truth. Some are wrong, and training on them raises confidence in the same mistakes, so more of them clear the threshold each round.

open as a page

How does uncertainty sampling in active learning choose the next rows to label?

level: middleimportance: should knowfreq 45%

basics

~20 s

Uncertainty sampling scores every unlabelled row by how unsure the current model is — one minus the top predicted probability, the top-two gap, or entropy — and buys human labels for the least certain rows, then retrains and repeats.

open as a page

How does distant supervision turn noisy labelling functions into training labels?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Domain experts write cheap rules that each vote a class on a row or abstain. The votes are combined, weighted by each rule's estimated accuracy, into probabilistic labels, and a model trained on them generalises beyond the rules.

open as a page