skip to content

Contrastive Self-Supervision

Learning without labels by pulling two augmented views of a sample together and pushing other samples apart, as in SimCLR and MoCo, or with no negatives as in BYOL. Interviewers probe collapse.

on this pageshow

questions

4

What does the InfoNCE contrastive loss optimise, and what does its temperature control?

level: middleimportance: must knowfreq 70%

answer

  1. softmax over similarities, one right answer
  2. the denominator holds every negative
  3. it divides logits before the softmax
  4. small value sharpens onto hardest negatives

basics

~20 s

InfoNCE is a softmax cross-entropy over similarities: an anchor must pick its own positive view out of a pool of negatives. The temperature divides those similarities before the softmax, so a small temperature concentrates the gradient on the hardest negatives.

solid answer

~50 s

For an anchor embedding, InfoNCE forms the similarity to its positive (the same input under a second augmentation) and to every negative, divides them all by a temperature `T`, and applies softmax cross-entropy with the positive as the correct class: `loss = -log( exp(sim_pos/T) / (exp(sim_pos/T) + sum_j exp(sim_j/T)) )`. So it is a (K+1)-way classification whose label is free — nothing but the pairing tells you what is positive. Similarities are cosine on L2-normalised embeddings, which is why `T` is meaningful: it rescales a bounded [-1, 1] quantity into logits. A small `T` sharpens the softmax, so almost all negative gradient comes from the few negatives nearest the anchor — fast separation, but brittle and unforgiving when a 'negative' is really the same thing. A large `T` weights negatives nearly uniformly: stable, less discriminative. Typical values run about 0.05 to 0.5 and are tuned per domain.

code

python · 16 lines
python
import math

pos, negs = 0.8, [0.7, 0.2, -0.1]

for t in (0.5, 0.07):
    logits = [pos / t] + [s / t for s in negs]
    top = max(logits)
    exps = [math.exp(x - top) for x in logits]
    loss = -math.log(exps[0] / sum(exps))
    hardest = exps[1] / sum(exps[1:])
    print("temperature", t,
          "loss", round(loss, 3),
          "hardest negative share", round(hardest, 3))

# temperature 0.5 loss 0.826 hardest negative share 0.637
# temperature 0.07 loss 0.215 hardest negative share 0.999

go deeper

for a junior

Be ready to say in one breath what a positive pair and a negative are, and that the loss asks the anchor to identify its own second view. Knowing that no human labels are involved is the point being checked.

for a middle

Expect to write the loss from memory and explain why the positive appears in the denominator, why embeddings are normalised first, and how dividing by the temperature reweights which negatives supply the gradient.

for a senior

Show you have tuned this. Explain how you picked a temperature for a specific domain, what the failure looked like at each extreme, and why you judged the run by a downstream probe rather than by the loss curve.

for a principal

Own the tradeoff between sharpness and robustness: low temperature buys separation but concentrates all the risk on false negatives you cannot audit. Be able to argue when a corpus is too concept-poor for aggressive contrastive discrimination at all.

## The setup Contrastive self-supervision has no labels. It manufactures a supervision signal from the data itself: take one input, pass it through two random augmentations, and declare those two **views** a *positive pair*. Every other item in the comparison pool is a *negative*. An encoder maps each view to an embedding, the embeddings are L2-normalised, and similarity is cosine — the dot product of two unit vectors, bounded in [-1, 1]. ## The loss Write `s_p` for the anchor's similarity to its positive and `s_1 ... s_K` for its similarities to K negatives. InfoNCE (also called the NT-Xent loss in this setting) is ``` loss = -log( exp(s_p / T) / ( exp(s_p / T) + sum_j exp(s_j / T) ) ) ``` Read it as cross-entropy over a (K+1)-way softmax where the correct class is 'my other view'. Two things follow immediately. First, **the positive term is in the denominator too**. The loss is minimised not by making `s_p` large in absolute terms but by making it large *relative to* the negatives. An encoder that maps everything to one vector gets `s_p = s_j = 1` for all `j` and pays `log(K+1)` — the worst possible score. The negatives are exactly what makes the trivial constant solution unattractive. Second, the objective decomposes into two forces. The numerator pulls the two views of the same input together — **alignment**, an invariance to whatever the augmentations changed. The denominator pushes everything else apart — **uniformity**, a pressure to spread embeddings over the sphere so the encoder does not waste capacity. Good representations need both; either alone degenerates. ## What the temperature actually does `T` divides every similarity before the exponential. Because softmax is scale-sensitive, this is a hardness-weighting knob, not a learning-rate knob. The gradient of the loss with respect to negative `j`'s similarity is proportional to that negative's softmax probability, `p_j = exp(s_j/T) / sum_all exp(s/T)`. With `T` small, a negative that is only slightly closer than the others takes almost all of the probability mass, so almost all of the repulsive gradient lands on it. With `T` large, `p_j` flattens toward uniform and every negative contributes about equally. Concretely: a positive at cosine 0.8 against negatives at 0.7, 0.2 and -0.1. At `T = 0.5` the hardest negative (0.7) carries about 64% of the negative mass. At `T = 0.07` it carries 99.9% — the other two negatives have effectively been deleted from the loss. That has consequences in both directions. Low `T` gives sharp, well-separated clusters quickly, but it makes training sensitive to a handful of examples per step, amplifies label noise you cannot see, and maximally punishes false negatives. High `T` trains gently and tolerantly, but the embedding space stays diffuse and downstream probes underperform. Sweeping `T` from 0.07 to 0.5 while pretraining a wrist-worn accelerometer encoder for human-activity data is a genuinely different experiment from sweeping a learning rate: the low end sharpens the boundary between 'walking' and 'stair-climbing' segments that look nearly identical in raw acceleration, while the high end refuses to commit and leaves them overlapping. There is no domain-independent best value; it interacts with the augmentation strength and with how many negatives you have. ## False negatives InfoNCE assumes every non-positive is a true negative. It is not. Imagine pretraining on a 40-item product catalogue with many photos per item: two augmented crops taken from two *different* photos of the same item are semantically identical, and InfoNCE actively pushes them apart. The rate of such collisions rises with the number of negatives and with how few distinct concepts your corpus contains, and their damage rises as `T` falls, because a false negative is by construction a *hard* negative and low temperature hands it the whole gradient. Mitigations: pretrain on data diverse enough that collisions are rare, raise `T`, or promote confident nearest neighbours to additional positives. ## Practical notes - Normalise before computing similarity. Without unit vectors, `T` is competing with the embeddings' free-floating magnitude and stops meaning anything. - The loss is usually symmetrised: each view of the pair takes a turn as the anchor. - Contrastive training typically applies the loss to the output of a small projection network on top of the encoder, and then discards that projection, transferring the encoder's representation instead — the projection layer absorbs the invariances the loss demands and the layer beneath it retains information the loss would have thrown away. - The loss value is not comparable across different `T` or different negative counts. Compare runs by a downstream probe, never by the raw loss.

  • What goes wrong when two items in the negative pool are semantically the same thing?
    They are false negatives, and InfoNCE pushes them apart anyway — two crops from two different photos of the same catalogue item get treated as a discrimination target. The damage grows with the negative count and gets worse at low temperature, because a false negative is by definition a hard negative and low temperature hands it nearly all the gradient. Fixes: a more diverse corpus, a higher temperature, or promoting confident nearest neighbours to extra positives.
  • Why is the contrastive loss usually applied on top of a projection network that is then thrown away?
    The loss demands invariance to the augmentations, so whatever layer it acts on is squeezed toward discarding colour, scale and crop information. Putting a small projection network there lets it absorb that pressure while the encoder beneath keeps information the objective would otherwise destroy but downstream tasks may need. Transfer therefore uses the encoder output, not the projected one.
  • Does InfoNCE require L2-normalised embeddings?
    In practice yes. Similarity is cosine, which is only defined on directions, and normalisation is what bounds the similarity to [-1, 1] so the temperature has a consistent meaning. Without it the encoder can trivially satisfy the softmax by inflating embedding norms, which changes the effective temperature during training and makes any tuned value meaningless.

It is a multiple-choice quiz in which the anchor must recognise its own second view in a lineup. The temperature sets how much the grader cares about the one impostor who looks most alike.

saying these in an interview costs you the question

  • Calls the temperature a learning rate for the softmax
  • Thinks InfoNCE needs labels to identify negatives
  • Forgets the positive term also sits in the denominator
  • Claims more negatives always help, ignoring false negatives
  • Compares runs by raw loss across different temperatures

context

open as a page

How do BYOL and SimSiam avoid representational collapse without using any negative pairs?

level: middleimportance: should knowfreq 46%

basics

~20 s

They break the symmetry that makes a constant output optimal: one branch carries an extra predictor network, the other is a stop-gradient target. Remove the stop-gradient and training reaches the trivial solution where every input maps to the same vector.

open as a page

In contrastive self-supervision, how does the augmentation set decide what the encoder ignores?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A positive pair is one input under two augmentations, so the encoder learns to discard whatever those augmentations change. Colour jitter buys hue invariance on street photos and ruins a dermatology encoder, where hue is the diagnostic signal.

open as a page

Would you supply contrastive negatives with SimCLR's large batch or MoCo's momentum queue on a small cluster?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Both need many negatives but buy them differently. SimCLR's negatives are the rest of the batch, so a 4096-sample batch and its memory are the price. MoCo decouples negatives from batch size with a momentum-encoder queue, so modest hardware suffices.

open as a page