skip to content

Why reuse a trained image classifier's penultimate activations as an embedding rather than its logits?

level: middleimportance: must knowfreq 72%

answer

  1. backbone plus a linear head
  2. the head projects into C dimensions
  3. null space of the head's weights
  4. softmax destroys geometry too
  5. last layer before label commitment

basics

~20 s

The penultimate vector is the general-purpose feature that the final linear layer scores. Logits are that same vector already projected onto the training classes, so they discard every distinction the label set happened to ignore.

solid answer

~40 s

A classifier factorises into a feature extractor and a linear head: the backbone maps an image to a `d`-dimensional penultimate vector `h`, and the head computes `logits = W h + b`. Reusing `h` keeps whatever the backbone learned to encode; the logits are a rank-at-most-`C` linear projection of `h`, so anything living in the null space of `W` is gone. With a 32-category label set and a 2048-wide penultimate layer that is a brutal reduction, and post-softmax probabilities are worse still because softmax is invariant to adding a constant to all logits and squashes the large gaps. For a used-goods marketplace deduplicating 80M listing photos I take the pooled penultimate vector, extract it deterministically (normalisation layers on their inference statistics, stochastic regularisers off), and compare directions there.

go deeper

for a junior

Be able to point at where in a trained classifier the embedding comes from: the vector just before the final linear layer, not the class scores and not the probabilities.

for a middle

Explain the mechanics — the head is a C-by-d linear map, so the logits keep only what survives that projection, and softmax additionally discards the shift direction and saturates the margins.

for a senior

Show you would make extraction reproducible: pinned preprocessing, inference-time normalisation statistics, stochastic regularisers off, and provenance stored with every vector so two encoder generations never mix silently.

for a principal

Own the argument that a frozen penultimate vector is only as general as the objective that produced it, and decide when the organisation should invest in a purpose-trained encoder instead of reusing a classifier.

## The two halves of a classifier Almost every trained image classifier factorises cleanly into two pieces. The **backbone** maps an input image to a vector `h` in `R^d` — in a residual convolutional network this is the global-average-pooled output of the last convolutional block, commonly `d = 2048`. The **head** is a single linear layer: `logits = W h + b`, where `W` is `C x d` and `C` is the number of training classes. A softmax turns the logits into probabilities, and cross-entropy against the label drives training. The *penultimate-layer embedding* is `h`: the last representation before the network commits to the label set. ## Why not the logits `W h + b` is a linear map from `R^d` into `R^C`. When `C < d` that map has a null space of dimension at least `d - C`, and every component of `h` inside that null space maps to exactly the same logits. The head was never asked to preserve it, so it does not. Concretely: a marketplace backbone trained to predict 32 top-level listing categories produces logits that answer one question — *which of 32 categories is this?* Two photographs of two different bicycles, one a genuine re-post of the other, produce nearly identical logits. In `h` they may still differ, because `h` also carries colour, background, framing, wear, and pose that the category loss simply had no opinion about. Probabilities are a further step in the wrong direction. Softmax is invariant to adding the same constant to every logit, so one full direction of the logit space is destroyed outright; it also saturates, compressing every large margin toward the same near-one-hot vector. A vector of probabilities from a confident model is close to a one-hot code and carries almost no geometry. ## Why not an earlier layer The opposite mistake is grabbing something from the middle of the network. Early layers encode edges, colour and texture — generic, transferable, but not semantically organised, so distance in that space measures superficial similarity rather than *is this the same object*. The penultimate layer is the most abstract representation the network builds that has not yet been collapsed by the head's projection, which is why it is the default reuse point. The honest caveat is that `h` is still shaped by the training objective. It received gradients only through `W`, so it is optimised to make the 32 categories linearly separable and nothing else. That is exactly why the reuse can disappoint on a task needing finer distinctions than the labels ever expressed, and it is the root of the low-effective-rank pathology you may later measure in the stored vectors. Under a large domain shift — the backbone saw natural photographs, you are embedding scanned documents — an earlier layer or actually adapting the backbone will beat the frozen penultimate vector. ## Practical extraction Three details decide whether the vectors you store are usable. **Determinism.** The same image must map to the same vector every time, or your near-duplicate threshold means nothing. That requires normalisation layers to use their fixed inference-time running statistics rather than batch statistics, and stochastic regularisers such as dropout to be disabled. It also requires the preprocessing — resize, crop, colour handling — to be pinned and identical between the backfill job and the online path. A silent preprocessing difference between the two is one of the most common causes of a dedup system that works offline and fails in production. **Where exactly you tap.** In a network whose last block ends in a rectified activation, `h` is non-negative. Every vector then lies in one orthant, so cosine similarities are all positive and bunched in a narrow band near the top of the range; thresholds have to be calibrated against that, and centring the set before comparing spreads the range back out. Tapping just before the activation instead gives a signed vector with a wider natural spread. **What you keep alongside it.** Store enough provenance — backbone identity, preprocessing version, extraction date — that you can tell which vectors came from which encoder. Similarity thresholds are not transferable between encoders, so a store that silently mixes two generations of vectors is broken even though nothing errors. ## The unsupervised alternative When there are no labels at all, the analogous vector is the bottleneck code of an autoencoder — train an encoder-decoder to reconstruct the input through a narrow middle layer, then reuse the middle layer. This is common for high-cardinality tabular telemetry, where the code becomes a dense feature block for a downstream model. The tradeoff mirrors the one above: a reconstruction objective preserves whatever explains the most input variance, which is not necessarily what predicts your target. A rare but decisive signal with tiny variance is precisely what such a code will drop, while a supervised penultimate vector keeps whatever its labels rewarded.

  • Would concatenating the pooled outputs of the last two blocks give a better embedding?
    Sometimes, and it is a cheap experiment. Earlier blocks are more generic, so the concatenation is more robust under domain shift. The costs are width and scale: the two parts have different variances, so the higher-variance block dominates any distance unless you normalise each part before concatenating. Validate it on the downstream metric — there is no general answer.
  • You can only get logits from a third-party model. Is that embedding useless?
    Not useless, just coarse. Logits over a rich label set are a usable semantic signature and work acceptably for broad similarity. They cannot support instance-level near-duplicate detection, because two distinct items in the same class collapse together. If you have the choice, take pre-softmax logits over probabilities — softmax discards the shift direction and compresses the margins.
  • Does a frozen penultimate embedding transfer to a domain the backbone never saw?
    Partially, and it degrades with the size of the shift. The backbone's late layers encode concepts specific to its training distribution, so on a distant domain their vectors are less discriminative than the generic textures of earlier layers. Measure it rather than assume it; if the frozen vector underperforms, adapting the backbone on target data is the fix, not a different pooling trick.

The penultimate vector is the full case file; the logits are the one-line verdict. You can appeal a verdict far better with the file than with the sentence.

saying these in an interview costs you the question

  • Calling softmax probabilities an embedding
  • Believing logits are a lossless re-encoding of the penultimate vector
  • Assuming any hidden layer works equally well as an embedding
  • Ignoring that inference-time normalisation statistics change the vector
  • Claiming penultimate features are task-neutral because the head was dropped

context