What does a linear probe on a frozen encoder's features actually measure?
answer
- only the read-out learns
- geometry of a fixed feature space
- linear separability under given labels
- lower bound, not an inventory
- compare against a random encoder
basics
~20 sA linear probe trains only a linear classifier on features a frozen network already produces. Its accuracy measures how linearly separable the target classes are in that fixed feature space; nothing inside the encoder is changed or learned.
solid answer
~50 sA linear probe freezes the encoder, pushes the data through it once, and fits a single linear classifier - logistic regression, in effect - on the resulting feature vectors. No gradient reaches the encoder and its normalisation statistics stay frozen, so the features are bit-for-bit the same before and after probing. The number you get is the accuracy of the best linear decision boundary in that feature space for those labels, which is a statement about the geometry of the representation, not about the encoder's capacity. It is useful because it is cheap - features are extracted once and reused - and comparable, since every encoder is judged through the same deliberately weak read-out. Always report which layer you probed and a control - the same probe on a randomly initialised encoder of the same shape - or you cannot separate learned structure from what architecture plus a linear model gives you for free.
go deeper
Be ready to state the mechanics: features come out of a frozen network, only a linear classifier is trained on them, and its accuracy is the reported number.
Explain why the read-out is kept linear, that the score describes linear separability of the labels in a fixed feature space, and that the probed layer must be named alongside the number.
Show you run the controls - majority class, raw inputs, a randomly initialised encoder - and that you fix one probe hyperparameter budget across everything you compare, never early-stopping on the test split.
Own the framing that probe capacity is the instrument. Argue where probes belong in an evaluation strategy and where they mislead a roadmap, since a probe measures decodability, not what a downstream system will achieve.
## The procedure A linear probe is four steps: 1. Take a trained network and **freeze** it: weights fixed, normalisation statistics fixed, no gradient allowed to flow back into it. 2. Choose a **layer** and a **pooling rule** (for example the pooled activation just before the original output layer, or the averaged token vectors of a chosen block) and run the labelled dataset through once, caching one feature vector per example. 3. Fit a **linear classifier** on those cached vectors: a weight matrix and a bias, trained with cross-entropy. This is ordinary multinomial logistic regression on fixed inputs. 4. Report the probe's accuracy on a held-out split of the same labelled data. Because the features never change, step 2 happens once and steps 3-4 are seconds of work. That cheapness is most of why probes exist: you can probe ten checkpoints, twelve layers and three label budgets in the time one fine-tune takes. ## What the number is a statement about Probe accuracy answers exactly one question: *how well can a hyperplane in this feature space separate these labels?* It is a property of the pair (representation, labelling), not of the encoder alone. Swap the labels and the same features give a different number. The read-out is kept weak on purpose - if you let the read-out be arbitrarily powerful, it would learn the task itself and the score would stop describing the representation at all. The corollary that interviewers press on is that the score is a **lower bound** on what the representation contains. Information can be there and simply not be exposed along a linear direction. So a high score is strong evidence the property is present and easy to reach; a low score is weak evidence that it is absent. ## Layer choice is part of the protocol A probe reads one layer, so "the probe scored X" is meaningless without naming the layer. Probing every block of a frozen twelve-layer encoder for part-of-speech is the classic demonstration: accuracy climbs through the early blocks, peaks in the **middle** of the stack, and falls off at the last block, which has specialised toward the objective the encoder was pretrained on. If you only ever probe the final layer you will systematically understate what the network knows about generic, lower-level properties. ## Controls you must run A bare probe number means little on its own. Three cheap baselines make it interpretable: - **Majority class**, so you know what chance looks like under this label distribution. - **A probe on the raw inputs** (pixels, bag-of-words counts), which tells you what a linear model achieves with no encoder at all. - **A probe on a randomly initialised encoder of the same architecture**, which is the important one. Random projections of a high-dimensional input are often surprisingly linearly separable, and a probe that beats the random-encoder control by two points has demonstrated almost nothing about pretraining. A fourth control used in the probing literature is to refit the same probe on **randomly assigned labels**: if it still fits well, your probe has enough capacity to memorise, and its score on real labels reflects the probe as much as the features. ## Details that decide whether the number is honest - **Probe hyperparameters are part of the measurement.** Regularisation strength, learning rate and number of epochs all move probe accuracy by points. Fix a budget and apply it identically to every representation you compare. - **Preprocessing counts as capacity.** Centring or standardising features before the probe changes what a linear boundary can do. It is legitimate, but it must be applied uniformly and disclosed. - **Never early-stop the probe on the test split.** Hold out a separate probe-validation split. - **Class imbalance** turns raw accuracy into a misleading summary; report per-class or balanced accuracy when the labels are skewed. ## When to reach for a probe Probes are a measuring instrument, not a deployment plan. They are the right tool when you want to compare checkpoints or layers cheaply, sanity-check that a self-supervised run learned anything at all before spending on downstream work, or track whether a property survives at a given depth. They are the wrong tool for predicting the ceiling of a downstream system, because the system will not be limited to a hyperplane.
- Why is the probe deliberately kept linear rather than given a hidden layer?Because the read-out's capacity is the measuring instrument. A weak read-out forces the score to describe the representation: if a hyperplane separates the classes, the encoder did the work. Give the probe hidden layers and it starts solving the task itself, so the score blends representation quality with probe capacity and stops being comparable across encoders.
- You probe all twelve layers of a frozen encoder for part-of-speech and the middle layers beat the last. What do you conclude?That the property is best exposed mid-stack and the final layer has specialised toward the pretraining objective, discarding or reshaping generic structure it no longer needs. It is a normal, well-documented shape, not a bug. The practical consequence is that the layer you probe is a protocol choice you must report, and the last layer is often the wrong place to look for low-level properties.
- What baseline makes a linear-probe number interpretable?The same probe fitted on a randomly initialised encoder of the same architecture, plus the majority-class rate. Random high-dimensional projections are often quite linearly separable, so probe accuracy in isolation can look impressive while the pretraining contributed almost nothing. The gap over the random-encoder control is the part that pretraining earned.
Freezing the encoder and fitting a hyperplane is like testing a filing system by asking whether one straight cut through the drawers separates the folders you want. A clean cut means the filing is good; a messy cut may just mean the right split is not straight.
saying these in an interview costs you the question
- Says the encoder is trained a little during probing
- Treats probe accuracy as the total information in the features
- Reports a probe score without naming the layer probed
- Skips the random-encoder and majority-class baselines
- Tunes probe regularisation per encoder and calls it a fair comparison
- Confuses a linear probe with fine-tuning the last block