skip to content

How does a sentence embedding model pool BERT token vectors into one vector?

level: middleimportance: must knowfreq 65%

answer

  1. encoder gives one vector per token
  2. something must collapse the sequence
  3. CLS versus masked mean
  4. pretraining objective never optimized similarity
  5. siamese contrastive training changed the default

basics

~20 s

A transformer encoder emits one vector per token, so a pooling step collapses them into one: take the [CLS] token's vector, or average all real token vectors (mean pooling). Which one works depends on how the model was trained.

solid answer

~60 s

A BERT-style encoder returns a sequence of contextual vectors, one per token, but search and clustering need a single vector per text — so the model applies a pooling operation on top. The two standard choices are **CLS pooling**, taking the vector at the special leading `[CLS]` position, and **mean pooling**, averaging the token vectors while masking out padding. Off-the-shelf BERT does badly at both, because its pretraining objectives (masked-token prediction and next-sentence prediction) never asked it to place similar sentences near each other; people who tried it early found the resulting cosine similarities barely beat averaging static word vectors. What fixed it was Sentence-BERT-style training: run the same encoder over two texts in a siamese setup, pool, and train with a contrastive or triplet objective so paired texts are pulled together and unrelated ones pushed apart. After that training, mean pooling is the usual default. The rule in practice is to use exactly the pooling the model was trained with — the model card states it — and never to swap it because another model used something else.

code

python · 14 lines
python
import torch

def mean_pool(token_embeddings, attention_mask):
    mask = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()
    summed = torch.sum(token_embeddings * mask, dim=1)
    counts = torch.clamp(mask.sum(dim=1), min=1e-9)
    return summed / counts

def cls_pool(token_embeddings):
    return token_embeddings[:, 0]

tokens = torch.randn(2, 5, 8)
mask = torch.tensor([[1, 1, 1, 0, 0], [1, 1, 1, 1, 1]])
print(mean_pool(tokens, mask).shape, cls_pool(tokens).shape)

go deeper

for a junior

Know that an encoder returns one vector per token and that a pooling step turns those into a single sentence vector, and be able to name CLS and mean pooling as the two common options.

for a middle

Explain masked mean pooling concretely, and explain why off-the-shelf BERT vectors are weak for similarity: its pretraining objectives never optimized the space for it. Name siamese contrastive training as the fix.

for a senior

Show that you treat pooling as part of the model contract — matched between queries and documents, taken from the model card, and any change validated against a labelled evaluation set because the failure is silent.

for a principal

Own the framing that an embedding space inherits the objective that trained it. Be ready to reason about when to adopt a published sentence encoder as-is versus investing in objective-specific training, and what that costs in evaluation and re-indexing.

## Where the vectors come from A transformer encoder does not natively output "a sentence vector". Feed it a tokenized text of length n and it returns n contextual vectors, one per token, each a mixture of that token's identity and everything around it. Useful for tagging tokens; useless directly for comparing two documents, because two texts of different lengths give differently shaped outputs and there is nothing to take a distance between. **Pooling** is the operation that collapses that variable-length sequence into one fixed-length vector. It is a small, usually parameterless step bolted onto the top of the encoder, and it is part of the model's definition, not an afterthought. ## The two standard strategies **CLS pooling.** BERT-family tokenizers prepend a special `[CLS]` token to every input. Take the final-layer vector at that position and call it the sentence embedding. The intuition is that attention lets `[CLS]` aggregate the whole sequence, so it acts as a summary slot. **Mean pooling.** Average the final-layer vectors across all positions. Critically, the average must be *masked*: batches are padded to equal length, and including padding positions dilutes the result with meaningless vectors. The correct implementation multiplies by the attention mask, sums, and divides by the number of real tokens. Other variants exist — max pooling over positions, using the last token (common when the backbone is a decoder-only model that only attends leftward), pooling a weighted combination of layers, or a small trained pooling head — but CLS and masked mean cover the overwhelming majority of what you will meet. ## Why vanilla BERT disappoints either way This is the part interviewers actually probe. BERT's pretraining objectives are masked-token prediction and next-sentence prediction. Neither of them shapes the geometry of the output space for similarity: nothing in training ever rewarded the model for putting *this* sentence near *that* semantically equivalent one. `[CLS]` in particular was optimized as a feature for the next-sentence classifier head, so its raw values are tuned for that head, not for cosine distance. The empirical result, well documented when people first tried it, is that raw BERT sentence vectors — CLS or mean — perform poorly on semantic similarity benchmarks, sometimes worse than simply averaging static word vectors. The lesson generalizes beyond BERT: **an embedding space is only as good as the objective that shaped it.** Extracting an internal representation from a model trained for something else gives you a representation optimized for something else. ## What sentence-encoder training changed Sentence-BERT-style training is the fix, and its two ingredients matter: 1. **A siamese (or dual-encoder) setup.** The *same* encoder weights process two texts independently, each pooled to a vector, and the loss is computed on the pair. Independent encoding is what makes the model usable for search — you can embed a corpus offline and compare later, because a document's vector never depends on the query. 2. **A contrastive objective.** Training pulls positive pairs (a question and its answer, a sentence and its paraphrase, a caption and its passage) together and pushes negatives apart, typically with a softmax over in-batch negatives or a triplet margin loss. This directly optimizes the geometry that cosine similarity will later read. Once the encoder has been trained this way *with a specific pooling function in the loop*, that pooling function is baked in. The model learned to put the information where that pooler will look for it. ## Practical rules - **Match the model card.** If the published model uses masked mean pooling, use masked mean pooling. Substituting CLS on a mean-trained model silently degrades quality — no error, just worse retrieval. - **Mask the padding.** An unmasked mean is a real and easy bug: it makes a vector depend on the batch's longest member. - **Be consistent across the index.** Documents and queries must be pooled identically, or the two sets of vectors are effectively in different spaces. - **Follow the model's normalization convention** as documented — many sentence encoders are trained and evaluated on L2-normalized outputs. - **Do not average vectors from different layers, or from different models, hoping for a better representation.** Layer choice is part of the trained configuration; cross-model mixing is meaningless. ## How this shows up as a bug Pooling defects are silent. Vectors still come out the right shape, similarity scores still land in a plausible range, and nothing throws. The only signal is degraded retrieval quality on a labelled evaluation set — which is exactly why any embedding pipeline change should be measured against a held-out set rather than eyeballed on a few queries.

  • What actually goes wrong if you forget the attention mask in mean pooling?
    Padding positions get averaged in as if they were content. Because batches are padded to their longest member, the same text then produces different vectors depending on what else was in its batch — non-deterministic embeddings and quietly degraded similarity. Short texts suffer most, since padding makes up more of their average. Nothing errors; you only see it as poor retrieval on an evaluation set.
  • Why does the siamese setup matter for search specifically, rather than just being a training detail?
    Because both texts are encoded independently, a document's vector never depends on the query. That is what lets you embed a corpus offline, index it, and compare cheaply at query time. A model that scores a pair jointly cannot be precomputed that way — every query would require a fresh pass over every document.
  • If a model's backbone is decoder-only rather than BERT-style, does pooling change?
    Often yes. Decoder-only backbones attend only leftward, so the final token is the only position that has seen the whole input, and last-token pooling is common for them; some are instead trained with a modified attention pattern so mean pooling works. The principle is unchanged: use the pooling the model was trained with, as published, rather than assuming mean.

The encoder is a committee that returns one opinion per member; pooling is the rule that turns the room into a single statement — either you appoint a chair to speak (CLS) or you take the average of everyone present (mean). The rule only works if the committee was trained knowing which one you would use.

saying these in an interview costs you the question

  • Claiming raw BERT's [CLS] vector is a good sentence embedding
  • Averaging over padded positions without the attention mask
  • Swapping CLS for mean pooling because another model uses it
  • Pooling queries and documents differently in the same index
  • Thinking pooling is a post-processing choice rather than trained-in

context