What does a sequence model lose about token order when recurrence is removed?
answer
- order lived in the operations, not the data
- shuffle the words and compare
- permute inputs, outputs permute too
- equivariant, not invariant, until pooling
- position must arrive as an input
basics
~20 sOrder itself. A recurrent net reads tokens one at a time, so 'dog bites man' and 'man bites dog' end in different states. A layer that mixes all positions at once sees only a bag of tokens unless position is supplied as input.
solid answer
~50 sRecurrence encodes position in the *process*: the state after three tokens is built by folding in token one, then two, then three, so a different ordering produces a different state without anyone designing that in. A layer that computes each position's output from a weighted mixture of all positions has no ordering in the operation at all, which makes it permutation-equivariant — permute the inputs and you get exactly the same output values back, permuted the same way. Nothing distinguishes the third token from the tenth, and relative distance disappears too. So order has to be re-injected as data, by folding a position-dependent signal into each token's representation before the mixing layer sees it. That is the first of the two things recurrence gave for free, and paying it back explicitly is the price of parallel training.
go deeper
Be ready to say that a recurrent net gets order from reading tokens one at a time, and that a layer seeing all tokens at once needs position handed to it. The 'dog bites man' example is enough to make the point.
Explain the symmetry properly: the operation contains no index, so permuting inputs permutes outputs identically, and pooling turns that into full invariance. Name relative distance as a second casualty, not just absolute position.
Show how you would detect a broken or ignored position signal in a real model — a shuffle test at evaluation, comparing pooled scores — and be clear that no amount of training can break a symmetry built into the function class.
Frame it as a deliberate exchange: an implicit property was swapped for an explicit input in return for parallelism over positions. Be ready to argue what else you would accept trading from implicit to explicit for the same kind of win.
## Order as a by-product of the computation A recurrent layer consumes a sequence one position at a time: `h_1` from `x_1`, `h_2` from `h_1` and `x_2`, and so on. Because the update is applied in sequence, position is baked into the arithmetic. Feed in the tokens of *dog bites man*, then feed in the same three tokens as *man bites dog*, and the two runs pass through different intermediate states and end at different final states. Nobody added a 'position feature'. The order of operations *is* the position information. This is why a recurrent model can learn that a word two tokens back matters, or that a sentence's first token is a question word. Distance is implicit as well: the state at position t has been transformed t-k times since it absorbed the token at position k, so 'how long ago' is encoded in how much processing has intervened. ## What a position-agnostic layer does instead Now consider a layer that produces the output at each position from a weighted combination of *all* positions' representations, where the weights depend on the content of those representations and not on where they sit. Written as a function over a set of vectors, the operation contains no index arithmetic at all. Such a layer is **permutation-equivariant**. Formally: if P is a permutation of the positions, then layer(P(x)) = P(layer(x)). Apply any reshuffling to the inputs and the outputs are the very same vectors, reordered to match. The layer has literally no way to tell an input from its shuffle. Two consequences follow immediately. 1. **Absolute position is gone.** There is no notion of 'the first token' or 'the third token'. Anything the task needs about position — sentence-initial capitalisation, the fact that a subject precedes a verb — is unavailable. 2. **Relative distance is gone.** The layer cannot express 'two tokens to the left'. Every position is equidistant from every other in the operation itself. If the model then pools across positions (sums or averages the outputs), permutation-*equivariance* becomes permutation-*invariance*: shuffled inputs produce a bit-for-bit identical result. That is exactly a bag-of-tokens model, and it cannot separate *dog bites man* from *man bites dog* at all. ## Giving the order back The standard remedy is to make position part of the data rather than part of the computation: attach a position-dependent signal to each token's representation before the mixing layer runs. Once each vector carries a marker of where it came from, the layer's weighted combination can condition on it, and the permutation symmetry is broken — the inputs are no longer interchangeable, because their representations now differ by position even when their tokens do not. How that signal is best constructed is a design question in its own right; the point for this topic is that it must exist, and that it is an explicit cost you took on the moment you removed the recurrence. ## Why the trade was still worth it Order came free with recurrence, but it came bundled with a strictly sequential unroll that cannot be spread over hardware. Adding an explicit position signal is cheap — a per-position vector, a handful of extra numbers per token, no serial dependency introduced. Trading an implicit property for an explicit input in exchange for making the whole sequence computable at once is a very good deal, and understanding *what* was traded is the substance of this question. ## The common misconceptions The first is to say 'the model will learn order from the data'. It cannot: permutation equivariance is a property of the function class, not of the fit. No setting of the weights breaks a symmetry that is built into the operation. The second is to assume token embeddings already encode position — they encode *identity*; the same word maps to the same vector wherever it appears, which is precisely why a separate position signal is needed. The third is to say a position-blind layer's outputs are unchanged by shuffling — they are the same *values*, but at permuted positions; the outputs become genuinely identical only after a pooling step that discards position. ## How to answer under pressure One sentence for the mechanism (recurrence encodes order in the order of operations), one sentence for the loss (a position-agnostic layer is permutation-equivariant, so absolute position and relative distance both vanish), one for the fix (position must be supplied as an input signal), and one for the framing (this was one of two free properties given up in exchange for parallel training).
- How would you demonstrate this loss empirically on a trained model?Shuffle the tokens within each input at evaluation time and compare. A recurrent model's outputs move sharply, because a different order drives it through different states. A position-agnostic layer with no position signal returns the same output values in permuted slots, and any pooled score is unchanged. If your model's accuracy barely moves under shuffling, its position signal is not reaching the layer that needs it.
- Is a position-blind mixing layer permutation-invariant or permutation-equivariant?Equivariant. Permuting the inputs permutes the outputs identically without changing their values — the information is still there, just relabelled. Invariance appears only after an order-destroying reduction such as summing or averaging across positions, which is what turns the model into a true bag of tokens.
- Does supplying position as an input restore everything recurrence gave for free?No. It restores order, but not the second free property: a recurrent cell's per-step cost is bounded by the hidden size regardless of how far back it looks, and that bound does not come back. It also shifts work onto learning — the model must learn to use the position signal, whereas recurrence made order unavoidable.
Reading a sentence aloud word by word, you cannot help but register the order. Handed the same words as a pile of fridge magnets, you would have to number them yourself.
saying these in an interview costs you the question
- Says the model will just learn word order from data
- Confuses permutation equivariance with permutation invariance
- Claims token embeddings already carry position
- Thinks shuffling leaves a position-blind layer's outputs untouched
- Treats losing order as a modelling quirk rather than a symmetry