skip to content

A linear probe on frozen features scores near chance while a two-layer head succeeds - what does that prove?

level: middleimportance: should knowfreq 38%

answer

  1. same features, different read-out
  2. the bottleneck is the hyperplane
  3. lower bound under a linear constraint
  4. absence of evidence, not evidence of absence
  5. a strong probe needs control comparisons

basics

~20 s

It proves the property is present in the frozen features but not exposed along any single hyperplane. A low linear-probe score bounds linear decodability, not the information itself, so it is never evidence that a feature is absent.

solid answer

~50 s

The features are identical in both runs - only the read-out changed - so the information must already be in the vectors; a hyperplane just cannot carve it out. Linearly inseparable arrangements are ordinary: a class defined by the interaction of two directions, or one that occupies a shell around another, defeats any linear boundary while a hidden layer handles it easily. So probe accuracy is a lower bound on decodability under a linear constraint, and 'the probe scored low' never licenses 'the model does not encode this'. The reverse claim also needs care: with enough capacity a read-out learns the task on its own, which is why a strong probe must be reported alongside controls - the same head on a randomly initialised encoder, and the same head on shuffled labels. What ultimately matters in practice is decodability under the read-out you will actually deploy.

go deeper

for a junior

Recall the key asymmetry: a high linear-probe score shows a property is present and easy to read, but a low score does not show the property is missing.

for a middle

Explain why identical frozen features with a stronger read-out can succeed - interaction, shell and conditional encodings all defeat a hyperplane - and state the lower-bound relationship precisely.

for a senior

Show the discipline that keeps probing honest in practice: fix the read-out family in advance, run random-encoder and shuffled-label controls, and match probe capacity to the head you will actually deploy.

for a principal

Own the epistemics for the team: decide what claims probing results may support in a report, and insist that decodability is never written up as evidence the model uses a feature.

## Why the result is not a contradiction Both runs consume the *same cached feature vectors*. Nothing about the encoder differed. The only thing that changed is the function class fitted on top: a hyperplane in one case, a hyperplane composed with a hidden non-linearity in the other. When the second works and the first does not, the deduction is forced: the information required to make the decision is contained in the vectors, and the linear read-out was the bottleneck. ## What linear inseparability looks like Three familiar geometries defeat a hyperplane while leaving the information fully present: - **Interaction (XOR-like) encoding.** The label depends on the *agreement* of two feature directions, not on either one. Neither direction alone correlates with the label, so every hyperplane is at chance, while one hidden layer of two units resolves it. - **Shell or ring structure.** One class surrounds another in feature space. Distance from a centre separates them perfectly; no linear boundary does. - **Conditional encoding.** The direction that carries the property flips sign depending on some other attribute of the input, so a single global weight vector averages the two cases into nothing. None of these is exotic. Networks have no incentive to make anything linearly readable unless the training objective's own read-out was linear. ## The bound, stated precisely Let *A_lin* be the accuracy of the best linear read-out and *A_f* the best accuracy achievable by any function of the same features. Then *A_lin <= A_f*, always. A probe measures the left-hand side. So: - A **high** probe score is strong evidence: the property is present *and* easy to reach, which is the useful claim for anyone planning to put a light head on frozen features. - A **low** probe score is weak evidence: it constrains the linear case only. Reporting it as "the encoder does not represent X" confuses absence of evidence with evidence of absence. ## The symmetric trap: capacity in the probe Once you allow the read-out a hidden layer, the natural next question is why not go further. The problem is that a sufficiently expressive read-out learns the task itself. Push far enough and you can extract almost any property from almost any representation, including a randomly initialised one - at which point the score describes the probe's training, not the encoder. Two controls keep a stronger probe honest: - **Random-encoder control.** Fit the identical head on a randomly initialised encoder of the same architecture. Only the margin over that control is attributable to what pretraining learned. - **Random-label control.** Fit the identical head on shuffled labels. If it fits them well, the head has the capacity to memorise, and its score on real labels is partly memorisation. The difference between real-label and random-label performance is the standard selectivity check in the probing literature. A disciplined probing report therefore fixes a read-out family, states it, and reports these controls alongside - rather than quietly upgrading the read-out until the result is the one expected. ## Decodability is not use A further limit worth stating out loud in an interview: even a high probe score only shows a property *can* be read out of a layer. It does not show the network's own downstream computation uses it. Probing is a claim about what is recoverable from an activation, not a causal claim about the mechanism. Establishing use requires an intervention - removing or perturbing the direction and observing whether behaviour changes - which is a different tool entirely. ## What to do with the finding If a linear probe fails and a small head succeeds, the practical readings are: 1. Do not delete the encoder or the layer from consideration; it carries the signal. 2. Match the probe to the deployment. If the plan is to serve a light head over frozen features, then a *linear* failure is a real, decision-relevant failure for that plan - it says the plan needs a non-linear head. If the plan is to fine-tune, linear decodability was never the requirement. 3. Report both numbers. "Linear 51%, two-layer head 74%, random-encoder two-layer head 55%" is a complete statement; either number alone is misleading. 4. Consider whether the linear failure is fixable by a different layer, since a property often becomes more linearly exposed a few blocks earlier or later in the stack.

  • If a stronger probe recovers more, why not always use the strongest probe available?
    Because a sufficiently expressive read-out solves the task itself and the score stops describing the representation - you can extract nearly anything from a randomly initialised encoder given enough capacity. Fixing a weak read-out is what makes numbers comparable across encoders. If you do use a stronger probe, report the same head on a random encoder and on shuffled labels so the margin, not the raw score, carries the claim.
  • Does a high linear-probe score show the network actually uses that property downstream?
    No. Probing shows a property is recoverable from an activation, which is a statement about decodability, not about the network's own computation. A direction can be present and ignored by every later layer. Demonstrating use requires intervening - ablating or perturbing that direction and checking whether behaviour changes - which is a different kind of experiment.
  • When is a linear probe's failure still the answer you care about?
    When the deployment plan is a light read-out over frozen features. There the linear constraint is not an artefact of measurement, it is the product constraint, so a linear failure genuinely says the plan will not work as designed. The response is to budget for a non-linear head, probe a different layer, or accept that the features must be reshaped.

A metal detector that stays silent tells you there is no metal near the surface, not that the field is empty. Digging with a spade is a different instrument, and finding a coin does not mean the detector was broken.

saying these in an interview costs you the question

  • Concludes the encoder does not represent the property
  • Says the features must have changed between the two runs
  • Treats probe accuracy as an upper bound on information
  • Escalates probe capacity with no control comparisons
  • Claims a high probe score proves the network uses the property
  • Assumes linear separability is the default for learned features

context