How do you make linear-probe comparisons between self-supervised encoders trustworthy?
answer
- encoder is the only variable
- identical tuning budget for every model
- report a label-budget curve
- train nothing: the neighbour vote
- register the protocol before results
basics
~20 sFix everything except the encoder: one layer and pooling convention, one probe family with an identical tuning budget, identical preprocessing and splits, and a reported label-budget curve. Add a parameter-free nearest-neighbour check as a cross-test, and freeze the protocol before results arrive.
solid answer
~50 sLinear evaluation looks objective but has enough free parameters to produce whichever ranking you want, so the protocol is the deliverable. Pin the layer and pooling rule, the feature preprocessing, the probe family, the hyperparameter search budget per encoder, the splits and the seeds - and give every encoder exactly the same allowance. Report a curve over label budgets, typically 1%, 10% and the full set, rather than one number, because encoders reorder as labels get scarce. Then cross-check with a probe-free evaluation: encode the labelled training set, classify each test example by a vote over its 20 nearest neighbours in feature space, and train nothing. Nothing there can be over-tuned, so when the probe ranking and the neighbour ranking disagree, the probe ranking is likely an artefact of tuning. Finally, write the protocol down before you see results, and state that linear rank order need not predict fine-tuned rank order on your task.
go deeper
Recall the core fairness rule: when comparing frozen encoders, everything except the encoder - layer, preprocessing, probe settings, splits - must be held identical.
Explain which knobs move probe accuracy by points, such as layer choice, feature centring, regularisation strength and evaluation-time augmentation, and why each must be fixed across encoders.
Show you run a parameter-free neighbour evaluation alongside the probe and know how to act on a disagreement, plus that you draw label-budget subsets once and share them across models.
Own the protocol as a written, pre-registered artefact with a single owner, and state its limits in the write-up - that linear rank order need not survive fine-tuning and does not transfer across domains.
## Why the protocol, not the number, is the artefact A linear probe reports one scalar per encoder, which invites the belief that the comparison is objective. It is not. Between the frozen encoder and the reported accuracy sit a dozen choices, each worth points, and each one can be tuned - honestly or not - in a direction that favours a preferred model. If you own encoder selection for a team, the defensible deliverable is a written protocol, fixed in advance, that every candidate encoder passes through unchanged. ## The degrees of freedom that move the number - **Which layer and which pooling.** Final block or a middle one; pooled activation or a specific token position. Points, not decimals. - **Which representation.** Some self-supervised methods train an auxiliary projection head on top of the encoder; evaluating before versus after that head gives different numbers. Choose one convention and apply it to all. - **Feature preprocessing.** Centring or standardising cached features changes what a hyperplane can do. Legitimate, but it must be uniform and disclosed. - **Probe regularisation and schedule.** Weight decay, learning rate, epochs. A generous sweep for one encoder and a default for another is the single most common way a comparison is quietly rigged. - **Evaluation-time augmentation.** Averaging features over multiple augmented views inflates accuracy. Fine if applied everywhere, indefensible if applied to one. - **Splits and seeds.** Reporting the best of several seeds for one method and a single run for another manufactures differences. - **Early stopping on the test split.** Choosing the probe's stopping point by test accuracy leaks the test set into the result. The rule that neutralises all of these is simple to state and hard to enforce: *identical everything, identical budget, encoder is the only variable.* ## Report a label-budget curve, not a point Evaluating frozen features at restricted label budgets - a 1% subset, a 10% subset, and the full labelled set - is far more informative than the full-label number alone, and it is standard in self-supervised comparisons. Encoders reorder across the curve: one representation may be marginally better with abundant labels yet clearly better in the scarce regime, which is usually the regime that motivated self-supervision in the first place. The subsets must be drawn once, class-balanced, and shared across encoders, or the comparison collapses into subset luck. ## The probe-free cross-check A nearest-neighbour evaluation trains nothing. You encode the labelled training set once, then classify each test example by a vote among its 20 nearest neighbours in feature space under a fixed similarity, optionally weighting neighbours by similarity. Its virtues are exactly the ones the probe lacks: there is no optimiser, no regularisation strength, no schedule, so there is nothing to tune per encoder and no way to over-fit the evaluation. Its cost is that it answers a different question - it reads local neighbourhood structure rather than global separability - and it inherits sensitivity to the feature scaling and similarity metric you fix in the protocol. Use it as a disagreement detector. When probe ranking and neighbour ranking agree, the ranking is robust to read-out choices and you can act on it. When they disagree, the probe result is likely the product of tuning or of a geometry that a hyperplane flatters, and the honest report says the encoders are not cleanly separated by this evidence. ## Stating the limits alongside the result Two caveats belong in any write-up: 1. **Linear rank order does not have to predict fine-tuned rank order.** Frozen-feature evaluation is a measurement of the representation as it stands; fine-tuning measures how well the initialisation supports being reshaped. Encoders are known to trade places between the two regimes, so a linear-evaluation win is not a promise about the downstream system. 2. **The evaluation dataset is part of the claim.** A ranking obtained on one domain's labels transfers no better than the encoder does. If the target domain differs from the evaluation domain, run the protocol on target-like data or say plainly that you have not. ## The organisational discipline Write the protocol down, name a single owner, and register it before any encoder is run. Publish cached features and split definitions so results are reproducible without re-deriving choices. Require the controls - majority class and a randomly initialised encoder of the same architecture - in the same table, so a reader can see the margin pretraining actually earned. Then treat any deviation from the protocol as a result that has to be re-run, not argued about. The point is not that the protocol is perfect; it is that a fixed imperfect protocol produces a ranking, while a flexible one produces a negotiation.
- Why is a 20-nearest-neighbour evaluation harder to game than a linear probe?Because it trains no parameters. There is no optimiser, no regularisation strength and no stopping rule to sweep per encoder, so the only shared choices are the neighbour count and the similarity, both fixed once in the protocol. That removes the main lever people pull, at the cost of measuring local neighbourhood structure rather than global linear separability.
- The probe ranking and the neighbour ranking disagree. What do you report?That the evidence does not separate the encoders. A disagreement means the ordering depends on which read-out you chose, which is exactly the dependency the comparison was meant to remove. Investigate the likely cause - probe tuning budget, feature scaling, a geometry that suits one read-out - and either fix the protocol and rerun, or decide on other grounds such as cost or latency.
- Why report frozen-feature results at 1% and 10% label budgets rather than only the full set?Because the scarce-label regime is usually the reason self-supervision was chosen, and encoders reorder across the curve - a representation that is marginally ahead with all labels can be clearly ahead or behind with few. A curve also exposes brittleness that a single number hides. The subsets must be drawn once, class-balanced, and shared by every encoder.
- Does winning the linear evaluation mean an encoder will win after fine-tuning on your task?No. Frozen-feature evaluation measures the representation as it stands, while fine-tuning measures how well the initialisation supports being reshaped, and encoders trade places between the two. Treat a linear-evaluation win as a cheap prior that narrows the shortlist, then confirm on the actual target before committing.
saying these in an interview costs you the question
- Sweeps probe hyperparameters for one encoder only
- Reports a single full-label number as the ranking
- Chooses the probed layer after seeing the results
- Early-stops the probe on the test split
- Applies evaluation-time augmentation to one model
- Presents a linear-evaluation win as a fine-tuning guarantee