skip to content

How do BYOL and SimSiam avoid representational collapse without using any negative pairs?

level: middleimportance: should knowfreq 46%

answer

  1. a constant output is a perfect score
  2. the two branches are deliberately asymmetric
  3. one branch receives no gradient
  4. predictor on one side, stop-gradient on the other

basics

~20 s

They break the symmetry that makes a constant output optimal: one branch carries an extra predictor network, the other is a stop-gradient target. Remove the stop-gradient and training reaches the trivial solution where every input maps to the same vector.

solid answer

~50 s

Both take two augmented views. The online branch runs encoder, projector and then a small **predictor**; the target branch runs the other view through a target network and its output is treated as a constant — a stop-gradient. The loss is a negative cosine similarity between the online prediction and the target projection, symmetrised over the two views. Without negatives, mapping everything to one vector would score perfectly, so the question is why gradient descent does not find it. The answer is asymmetry: the predictor makes the online branch *predict* the target rather than equal it, and the stop-gradient means the target branch never receives the gradient that would pull it toward the online output. BYOL also makes the target an exponential moving average of the online weights; SimSiam showed that average is a stabiliser, not the cause. Delete the stop-gradient and every view maps to one constant vector.

go deeper

for a junior

Know that some self-supervised methods use no negatives at all, that the danger is every input mapping to the same vector, and that this failure is called collapse.

for a middle

Be ready to draw the two branches, name the predictor and the stop-gradient, state the loss as a cosine similarity, and explain why removing either asymmetry lets training reach the constant solution.

for a senior

Demonstrate that you monitor for silent collapse in a real run — embedding spread, periodic probes — and that you know the loss curve is uninformative here. Be able to describe how a collapsed run presents.

for a principal

Own the tradeoff between negative-free and contrastive families: one removes the negative-count and false-negative problems, the other has a loss you can partly trust. Argue which risk your team can actually detect and operate.

## The problem these methods create for themselves Contrastive objectives have two forces: pull views of the same input together, push different inputs apart. The repulsive force is the expensive one — it is why you need thousands of negatives, a queue, or a very large batch. So the obvious question is whether you can keep only the attractive force. You cannot, naively. If the loss is only 'make the two views agree', then the function that maps every input to the same constant vector achieves a perfect score. This is **complete collapse**: the encoder has discarded the input entirely, the loss is at its floor, and the representation is worthless. Any negative-free method must explain what stops this. ## The architecture BYOL and SimSiam share a shape: - An input is augmented twice into views `v1` and `v2`. - The **online** branch maps `v1` through an encoder and a projector to `z1`, then through a small **predictor** network to `p1 = q(z1)`. - The **target** branch maps `v2` through a target encoder and projector to `z2`, and `z2` is wrapped in a stop-gradient: it is treated as a fixed constant with respect to backpropagation. - The loss is negative cosine similarity between `p1` and `stopgrad(z2)`, symmetrised by also computing it with the roles of the two views swapped. The difference between the two: in BYOL the target network's weights are an exponential moving average of the online network's weights (a momentum target), while SimSiam shares weights outright between the branches and relies on the stop-gradient alone. SimSiam's contribution was precisely to strip away the moving average and show that training still does not collapse — which relocates the credit from the momentum target onto the stop-gradient and predictor pair. ## Why the asymmetry matters The two ingredients do different jobs. The **stop-gradient** removes one of the two paths by which collapse could be reached. If gradients flowed through both branches, the fastest way to make the two outputs agree is for both to move toward each other and meet at a constant. Freezing one side turns the step into 'match a fixed target' rather than 'meet in the middle'. The **predictor** prevents the remaining path. Without it, the online branch's best move is to reproduce the target exactly, and since the target is itself the network's own output on another view, the degenerate fixed point is still reachable. The predictor means the online branch is optimised to predict the target's output rather than to become it, so the encoder is not directly pushed toward the target's representation. SimSiam offers an interpretation of the pair: the procedure resembles an alternating optimisation, where the stop-gradient defines a target held fixed at each step and the predictor approximates the mapping to it. That is a hypothesis rather than a proof — there is no complete theory of why these methods avoid collapse — but it explains why *both* ingredients are needed and why removing either one is catastrophic. ## The ablation everyone should know Take a working SimSiam-style setup and delete the stop-gradient. Training does not degrade gracefully. Within a small number of steps the loss plunges to its minimum value (-1 for negative cosine similarity of unit vectors) and stays pinned there, and every view of every input maps to essentially the same output vector. The same happens if you remove the predictor. This is the sharpest available demonstration that the loss value carries no information about representation quality in this family: the best-looking loss curve you will ever see is the one produced by total collapse. ## Detecting collapse in a real run Since the loss lies, monitor the representation: - **Per-dimension spread.** L2-normalise the projector outputs and track the standard deviation of each coordinate across a batch. Under healthy training on d dimensions this sits near `1/sqrt(d)`; under collapse it falls toward zero because every sample has the same value in every coordinate. - **A periodic probe.** Freeze the encoder every so often and fit a cheap linear classifier on a small labelled set. This is the only measurement that tracks what you actually want. - **Neighbour sanity checks.** Retrieve the nearest neighbours of a handful of held-out inputs. Under collapse, the neighbour list is arbitrary and the similarity to everything is near 1. ## Practical judgment - The stop-gradient is not optional and is not a stabilisation trick you may drop for a cleaner implementation. - The predictor is deliberately small relative to the encoder; it is a light head, and letting its learning rate lag the rest of the network has been reported to help stability. - BYOL's momentum target genuinely improves stability and results even though it is not the reason collapse is avoided — so 'SimSiam proved the target network is useless' is a misreading. - Negative-free methods remove the negative-count engineering problem and, with it, the false-negative problem, at the cost of a training procedure whose failure mode is silent unless you are monitoring for it.

  • During a run the loss looks excellent — how do you check the representation has not collapsed?
    Never trust the loss in this family, since collapse produces the best possible value. L2-normalise the projector outputs and track the per-coordinate standard deviation across a batch: healthy training sits near one over the square root of the dimension, collapse drives it toward zero. Back that with a periodic linear probe on a small labelled set and a nearest-neighbour spot check.
  • Is the momentum target network required, or is the stop-gradient enough?
    The stop-gradient plus predictor is the essential mechanism — SimSiam demonstrated that a shared-weight setup with no moving average still avoids collapse. The momentum target is not useless, though: it stabilises training and generally improves the resulting representation. Treat it as a stabiliser you would keep, not as the ingredient that prevents collapse.
  • What loss do these methods actually minimise?
    Negative cosine similarity between the online branch's prediction and the stop-gradient target projection, computed on L2-normalised vectors and symmetrised so each view takes a turn on each side. There is no denominator and no partition function, which is exactly what removes the need for negatives and creates the collapse risk in the first place.

Two people are told to agree on an answer. If both can move, they will agree instantly on nothing at all. Freezing one of them, and asking the other to predict rather than parrot, is what forces the answer to carry content.

saying these in an interview costs you the question

  • Says a loss falling to its minimum proves training is working
  • Claims BYOL secretly uses negatives
  • Thinks the target branch is trained by backpropagation
  • Believes the predictor head is an optional detail
  • Cannot say what the trivial solution looks like

context