In a CNN, what is the difference between translation equivariance and translation invariance?
answer
- moves with it, or does not care
- shift the input, watch the feature map
- the body gives one, the head gives the other
- exact only up to stride and borders
basics
~20 sEquivariance means shifting the input shifts the feature map by a corresponding amount. Invariance means the output does not change at all. Stacked convolutions give you equivariance; invariance has to be added by the readout or learned from data.
solid answer
~40 sWrite T for a spatial shift. A layer is equivariant if `f(T x) = T f(x)` — shift the input and the output shifts too — and invariant if `f(T x) = f(x)`. Weight sharing buys the first, not the second: the same kernel is applied at every position, so sliding the input slides the responses. Take a conveyor-belt frame where a scratch moves 40 pixels along the belt. The activation on that scratch moves with it through the convolutional stack, but the pass/fail decision must come out identical — that is invariance, supplied by a readout that aggregates across positions and so discards where the activation was. Equivariance is also only approximate: padding disturbs the borders, and strided downsampling means only shifts that are multiples of the cumulative stride map cleanly.
go deeper
Recall the one-line distinction: equivariant means the feature map moves with the input, invariant means the answer does not change. Be able to say which of the two a convolution itself provides.
Explain why sharing produces equivariance — the same kernel arrives over the pattern at its new position — and name where invariance actually comes from in a classification network. Mention that stride and padding make it only approximate.
Show you have debugged this. Be ready to reason about predictions that wobble under a small shift, to say which architectural choice destroyed or preserved position, and to pick the right property for a dense-prediction head versus a whole-image head.
Own the question of how much symmetry to build in versus teach. Argue the cost of architectural invariance against covering the variation in data, and set the standard for how a team verifies a claimed symmetry rather than assuming it.
## Two properties, stated precisely Let T be an operation that shifts a spatial input by some number of pixels. For a function f — a layer, a stack of layers, or a whole network: - f is **equivariant** to that shift if `f(T x) = T f(x)`. Shift the input, and the output changes in the same, predictable way: it shifts too. Information about *where* is preserved. - f is **invariant** to that shift if `f(T x) = f(x)`. Shift the input, and the output is identical. Information about *where* has been discarded. These are not degrees of the same thing. Equivariance says the output tracks the change; invariance says the output ignores it. A network can be equivariant in its body and invariant at its output, and that is the usual design. ## Why weight sharing produces equivariance A convolution applies the same kernel window at every position. If a pattern that produced a strong response at position p is moved to position p + d, the identical window arrives over it at p + d and produces the identical response there. Nothing about the layer's parameters is position-specific, so responses move exactly as the input does. That is the direct consequence of sharing, and it is what people mean when they say convolution has a built-in translation prior. Notice this gives no invariance whatsoever. A feature map that shifts is a *different* feature map. If you flattened it and fed it into a position-sensitive dense readout, the network's output would change under a shift. ## Where invariance actually comes from Invariance has to be introduced deliberately, and there are essentially three routes: 1. **A readout that discards position.** If the final step aggregates each feature map over all spatial locations into a single number per channel, then a shifted map produces the same aggregate, so the classifier output is unchanged. This is the standard route for a whole-image classification head. 2. **Downsampling that coarsens position.** Reducing spatial resolution makes small shifts stop mattering at the coarse scale, which yields partial, local invariance rather than exact invariance. 3. **Learning it from the data.** If the training set contains the object at many positions with the same label, the network can learn a response that happens to be roughly position-independent, even where the architecture does not enforce it. ## The conveyor-belt case A fixed camera photographs parts on a belt and a model must output pass or fail. A scratch appears 40 pixels further along in one frame than in another. Through the convolutional body, the activation caused by that scratch moves along with it — that is equivariance doing its job, and it is exactly what you want, because the same detector must work at both positions without extra parameters. At the output, though, the two frames must agree: the part is defective either way. That agreement is invariance, and it is supplied by the aggregating readout, not by the convolutions. If you replaced that readout with something position-sensitive, you would get two different scores from two frames that show the same defect, and no amount of extra convolution depth would fix it. ## Equivariance is only approximate in real networks Three things break the clean equality: - **Borders.** Padding at the edges means a pattern near the boundary is not treated identically to the same pattern in the middle, and content shifted off the edge is simply gone. - **Stride and downsampling.** A layer with stride s is only cleanly equivariant to shifts that are multiples of s. Compose several such layers and only shifts that are multiples of the cumulative stride map exactly onto grid positions. - **Aliasing.** Because downsampling samples the response on a coarse grid, a shift of a single pixel can change which peak lands on a sample point, which is why predictions can wobble under a one-pixel shift even though the architecture is nominally shift-friendly. ## When you want to keep equivariance rather than collapse it Dense prediction tasks depend on equivariance surviving all the way to the output. In segmentation the mask must move when the object moves; in detection the box must move; in keypoint or depth prediction the output map is spatial by definition. For those, you deliberately avoid an aggregating readout that would destroy position, and the equivariance of the convolutional body is the property that makes a spatially structured output learnable at all. ## And it is translation only Weight sharing gives a prior about *translation* and nothing else. A convolutional stack is not automatically equivariant or invariant to rotation, scale, reflection, or lighting change. Those need either a different architectural construction or training data that covers the variation. Claiming that a convolution handles rotation for free is a frequent and easily caught mistake. ## Interview framing The crisp answer is: convolution gives equivariance, the readout gives invariance, and knowing which one your task needs at the output is the actual design decision.
- Is a convolutional stack exactly equivariant to any pixel shift you apply?Only in the idealised case: stride one, no boundary effects. A layer with stride s is cleanly equivariant only to shifts that are multiples of s, and a stack is only equivariant to multiples of its cumulative stride. Padding disturbs the borders, and aliasing from downsampling means a one-pixel shift can change which response lands on a sample point, so real networks are approximately rather than exactly shift-equivariant.
- For which tasks do you deliberately want equivariance preserved all the way to the output?Dense prediction. In segmentation the predicted mask must move when the object moves; in detection the box must move; keypoint, flow and depth outputs are spatial by construction. For these you avoid an aggregating readout that would collapse position, because the output is supposed to be a function of where things are, not just whether they are present.
- Does weight sharing give you any robustness to rotation or scale?No. Sharing encodes a prior about translation only. A rotated or rescaled version of a pattern presents a different local arrangement of pixels to the same kernel, so the response is not related to the original in any guaranteed way. Robustness to those variations has to come from data covering them, or from an architecture explicitly constructed for that symmetry.
Equivariance is a shadow: move the object and the shadow moves with it. Invariance is a scale: put the object anywhere on the pan and it reads the same weight.
saying these in an interview costs you the question
- Says a convolution is translation invariant on its own
- Claims equivariance means the feature map is unchanged by a shift
- Assumes CNNs handle rotation and scale for free too
- Thinks a flatten-then-dense head is position independent
- Treats equivariance as always desirable to remove, even for segmentation