How does uncertainty sampling in active learning choose the next rows to label?
answer
- spend the label budget where it matters
- score every unlabelled row for doubt
- top-two gap, or spread of the whole distribution
- a committee can disagree instead
- the resulting pool is no longer random
basics
~20 sUncertainty sampling scores every unlabelled row by how unsure the current model is — one minus the top predicted probability, the top-two gap, or entropy — and buys human labels for the least certain rows, then retrains and repeats.
solid answer
~50 sIn pool-based active learning you train on what you have, score the whole unlabelled pool, and buy labels for the rows the model is least sure about, on the theory that they sit nearest the decision boundary and are worth more than a random row. Three common scores: least confidence (`1 - max(p)`), margin (the smallest gap between the top two class probabilities), and entropy. Margin usually behaves best on multi-class problems because entropy is dominated by the long tail of unlikely classes. Query-by-committee is the ensemble version: fit several models on different resamples or seeds and pick the rows they disagree on most. Two practical constraints follow — you must select in batches with a diversity term, or you buy 200 near-duplicates of the same ambiguous row, and the pool you build is deliberately non-random, so you can never use it to estimate accuracy.
code
python · 24 linesimport math
# predicted class probabilities for three unlabelled rows
pool = {"a": [0.50, 0.30, 0.20],
"b": [0.45, 0.44, 0.11],
"c": [0.90, 0.05, 0.05]}
def least_confidence(p):
return 1 - max(p)
def margin(p): # smaller margin = more uncertain
top = sorted(p, reverse=True)
return top[0] - top[1]
def entropy(p):
return -sum(q * math.log(q) for q in p if q > 0)
for name, p in pool.items():
print(name, round(least_confidence(p), 3),
round(margin(p), 3), round(entropy(p), 3))
print("least-confidence picks:", max(pool, key=lambda k: least_confidence(pool[k])))
print("margin picks:", min(pool, key=lambda k: margin(pool[k])))
print("entropy picks:", max(pool, key=lambda k: entropy(pool[k])))go deeper
Know the loop and the intuition: train, score the unlabelled rows for uncertainty, send the least certain ones to a human, retrain. Be able to name one score, such as one minus the top predicted probability.
Explain how least confidence, margin and entropy differ and when each misleads, and describe query-by-committee as the ensemble-disagreement alternative. Be ready to say why entropy struggles when there are many classes.
Show that you have run one: seed batches randomly, batch to reviewer throughput with a diversity term, guard a randomly sampled test set from the loop, and insist on a random-sampling baseline at equal label count before declaring a win.
Own whether the loop is worth its operating cost at all. Weigh a permanent retrain-and-queue pipeline and a training set curated for one model family against simply buying more random labels, and decide what evidence justifies the commitment.
## The problem it solves When an annotation budget is the binding constraint — one reviewer on a 50-million-listing marketplace catalogue who can label 500 items a day — the question is not *how many* labels you can afford but *which* labels. Randomly sampled rows spend most of the budget on examples the model already gets right. Active learning tries to spend it on examples that change the model. Pool-based active learning is a loop: train on the currently labelled rows, score every row in the unlabelled pool with an acquisition function, send the top-scoring batch for human labelling, add them, retrain, repeat. ## The three uncertainty scores All three read the model's predicted class probabilities for a row and turn them into one number. - **Least confidence:** `1 - max(p)`. How unsure is the model about its own top choice? Simple, but it ignores everything except the top class. - **Margin:** `p_top1 - p_top2`, and you take the rows with the *smallest* margin. This targets rows the model cannot separate between its two leading candidates, which is usually the most useful signal on a multi-class problem. - **Entropy:** `-sum(p_i * log p_i)`, taking the largest values. It accounts for the whole distribution, which is a virtue when classes are few and a liability when they are many, because a long tail of small probabilities can make an otherwise-decided row look uncertain. The three genuinely disagree. Given a three-class problem, a row at (0.45, 0.44, 0.11) has a tiny margin but lower entropy than a row at (0.50, 0.30, 0.20), which has a larger margin. Which one is worth a label depends on what you want: the first is a near-tie between two candidates, the second is diffuse confusion. ## Query-by-committee Instead of trusting one model's probabilities, train a committee — five fits on different bootstrap resamples, different seeds, or different algorithms — and score each row by how much the members disagree, using vote entropy or the divergence of each member's distribution from the committee mean. Five disagreeing fits voting on which 200 rows to send for labelling next is often more robust than one model's confidence, because disagreement across differently-fitted models is harder to fake than the confidence of a single overconfident model. It costs several fits per round. ## The failure modes that separate levels **Outliers and unlabelable rows.** Whatever is most uncertain includes corrupted records, junk listings and genuinely ambiguous items. Pure uncertainty sampling will happily spend the whole budget on them, and the human comes back with shrugs. Density-weighting the score — preferring uncertain rows that also sit in a populated region — is the usual defence. **Batch redundancy.** You cannot retrain after every single label, so you select 200 at a time; but the top 200 by uncertainty are often 200 versions of the same ambiguity. You need an explicit diversity or coverage term, otherwise 200 labels buy you the information of about five. **Cold start.** With almost no labels, the model's uncertainty estimates are noise, and the loop chases nothing. Seed the first batch or two randomly. **The pool is not a random sample.** This is the one interviewers care about most. By construction the labelled set over-represents hard, near-boundary rows, so any accuracy, precision or recall computed on it is biased and usually pessimistic. Hold out a separately drawn, randomly sampled labelled set for evaluation before the loop starts, and never let the acquisition function touch it. **The pool is model-specific.** Uncertainty was defined by the model you were using at the time. Swap the model family later, and you inherit a training set that was curated for a different decision boundary — a real cost that rarely gets mentioned when the loop is proposed. **Human throughput is the real constraint.** A round trip that requires retraining, re-scoring 50 million rows and re-queueing work has an operational cost. In practice rounds are sized to the reviewer's day, not to what the algorithm would prefer. ## When it is worth it Active learning pays when labels are genuinely expensive relative to compute, when the unlabelled pool is large and diverse, and when someone will actually operate the loop. It pays less when labelling is cheap, when the model is already near its ceiling, or when the pool is small enough to label outright. And its gains are usually reported against a random-sampling baseline at equal label count — which is the comparison you should insist on seeing, because a badly configured loop can lose to random sampling.
- How does query-by-committee decide which rows to buy, and when is it worth the extra fits?Fit several models on different resamples, seeds or algorithms, then rank rows by how much the members disagree — vote entropy, or divergence from the committee's mean distribution. It is worth the cost when a single model's probabilities are badly calibrated, since disagreement across independently fitted models is a sturdier signal of ambiguity than one model's confidence. The price is several fits per round.
- Why can't you evaluate the model on the pool that active learning labelled?Because the loop deliberately over-samples hard, near-boundary rows, so that set is not a random draw from the data. Metrics computed on it are biased, usually pessimistic, and shift every round as the acquisition function changes what it buys. Draw a random, human-labelled test set before the loop begins and keep the acquisition function away from it.
- Your reviewer labels 500 items a day. How do you size and select each batch?Size the batch to a reviewer's shift so the queue is never empty and retraining happens between shifts, not between items. Select with uncertainty plus an explicit diversity term so the batch does not contain hundreds of copies of the same ambiguity, and weight toward dense regions so the budget is not spent on junk records. Seed the very first batches randomly, since early uncertainty estimates are noise.
- What baseline must an active-learning result be compared against?Random sampling at the same label count. Plot both learning curves — score against number of labels bought — and the active-learning curve has to sit above the random one to justify the operational complexity. Badly configured loops that chase outliers or buy redundant batches genuinely lose to random sampling, and this comparison is the only thing that catches it.
saying these in an interview costs you the question
- Measures model accuracy on the actively labelled pool
- Selects a large batch with no diversity term
- Starts the loop with almost no labels at all
- Assumes uncertain rows are always informative rows
- Never compares against random sampling at equal cost
- Forgets the curated pool is tied to one model family