Why does a transformer need positional encoding, given that attention sees every token?
answer
- attention is order-blind
- a set, not a sequence
- permutation-equivariant outputs
- "dog bites man" vs "man bites dog"
- inject at input or in the scores
basics
~20 sSelf-attention treats its input as a set — reorder the tokens and each one gets the same representation back, just moved. Order has to be injected separately, or "the auditor approved the invoice" and "the invoice approved the auditor" are indistinguishable.
solid answer
~50 sSelf-attention computes, for every token, a weighted average of value vectors where the weights come from query-key dot products. Nothing in that computation refers to an index, so the layer is **permutation-equivariant**: shuffle the input and the outputs are the same vectors in shuffled order. Since meaning in language depends on order, position must be supplied. There are two places to supply it. The first is the input: add an absolute position vector (a fixed sinusoidal formula or a learned table) to each token embedding, once, before layer one. The second is the attention computation itself: bias or transform the query-key interaction so the score depends on the distance between two tokens — rotary encoding (RoPE) is the dominant form of this, applied at every layer. Modern decoder-only LLMs overwhelmingly use the second family, because relative structure is what language actually needs.
go deeper
Be ready to say plainly that self-attention has no built-in notion of order, so position must be added — and to give a two-sentence example where reordering changes meaning but not the computation.
Explain permutation equivariance concretely and name both injection points: absolute vectors added to embeddings versus a relative signal applied to queries and keys inside every attention layer.
Expect to justify why the relative family displaced absolute in production decoder-only models, and to note that the causal mask already leaks a weak position signal that NoPE work exploits.
Own the design call: what your task's structure actually needs — global coordinates or local offsets — and what a positional choice commits you to when sequence lengths later grow well past what you trained on.
## The problem: attention consumes a set, not a sequence A self-attention layer produces, for each position, a weighted sum of value vectors. The weights come from comparing that position's query vector against every key vector. Written out, that computation contains no reference to *where* a token sits — only to *what* its vectors contain. The consequence is precise: self-attention is **permutation-equivariant**. Permute the input tokens and you get exactly the same set of output vectors, permuted the same way. Nothing is lost and nothing is gained; the layer literally cannot tell the two orderings apart. The feed-forward sublayer does not rescue this: it applies the same transformation to each position independently, so it never mixes across positions. Layer normalization is also per-position. Attention is the only mixing operation in the stack, and it is order-blind by construction. So "the auditor approved the invoice" and "the invoice approved the auditor" yield identical bags of representations. Order information must be *injected*, because it cannot be *inferred* from the architecture. ## Two places to inject position **At the input, as absolute position.** Build one vector per position — either from a fixed formula (the original sinusoidal encoding) or from a learned lookup table with one row per index — and add it to the token embedding before the first layer. The token vector now carries "I am the word *invoice*" plus "I am at index 4", and attention can, in principle, read distance out of the difference. This happens once, at the bottom of the stack. **Inside attention, as relative position.** Instead of tagging the inputs, change how a query and a key interact so that the resulting score already depends on the gap between the two tokens. Learned relative-position biases add a per-offset scalar to the attention logits. ALiBi subtracts a penalty proportional to distance. RoPE rotates the query and key vectors by an angle proportional to their absolute positions, so that their dot product ends up depending only on the offset between them. This family is applied at every layer, to queries and keys, not to the token embeddings. ## Why the relative family won Three reasons. First, most of what order contributes to meaning is relative: which word modifies which, how far back the subject was, whether a bracket closed. Second, absolute schemes run out — a learned table has no row for a position it never trained on. Third, a relative signal reapplied at every layer keeps position available deep in the stack, rather than relying on it surviving dozens of residual updates from the embedding. As of mid-2026 essentially every open and frontier decoder-only LLM uses rotary encoding as its baseline; ALiBi belongs to an earlier generation of models. ## The causal-mask nuance worth knowing In a decoder-only model, the causal mask lets position *n* attend only to positions 0..*n*, so each position sees a different-sized prefix. That asymmetry is itself a weak position signal, and research on **NoPE** (no positional encoding) showed that decoder-only transformers can learn order-sensitive behaviour without any explicit encoding at all. This is why "no encoding" is a live architectural option in 2026 rather than an obvious mistake. A bidirectional encoder has no such asymmetry and genuinely cannot work without an encoding. ## When absolute position genuinely matters Relative is the right default for text, but not universally. A symbolic-music sequence model where each token is a note event inside a bar cares about absolute grid position — "beat one of bar seventeen" is a meaningful coordinate, not just an offset from the previous note. Fixed-length structured inputs, where slot *k* always means the same field, are similar. If your task has a genuine global coordinate system, an absolute scheme is not a legacy choice. ## What an interviewer listens for A clean answer says: attention is order-blind because it is a weighted set operation; position is injected either into the inputs (absolute) or into the query-key interaction (relative); modern LLMs use rotary encoding, which is the relative kind; and the causal mask supplies a weak implicit signal on top. Hand-waving that "the model reads the tokens in order" is the answer that fails.
- Doesn't the causal mask already tell a decoder-only model where each token sits?Partly. Under a causal mask, position n attends over a prefix of size n+1, so the number of visible tokens varies with position and gives an implicit signal. Work on NoPE showed decoder-only stacks can exploit this and learn order-sensitive behaviour with no explicit encoding. It is a weak, indirect signal, which is why most models still add one — but it is why NoPE is a serious option rather than a bug.
- Where in the network does the position signal actually enter, for absolute versus rotary schemes?An absolute encoding is added to the token embedding once, before the first layer, and must survive every residual update after that. A rotary or relative scheme is applied inside the attention computation, to queries and keys, in every layer — so position is re-injected dozens of times and never has to be preserved through the residual stream.
- Are there sequence tasks where absolute position is more useful than relative offset?Yes. A symbolic-music model whose tokens are note events inside a metrical grid cares that something falls on beat one of bar seventeen — an absolute coordinate with meaning, not just a gap from the previous event. Fixed-schema structured inputs behave the same way. For natural language, relative structure dominates, which is why relative and rotary schemes became the default there.
Attention is like reading a pile of index cards spread on a table: you can see all of them at once, but nothing tells you which was on top. Positional encoding is the page number written on each card.
saying these in an interview costs you the question
- Says a transformer processes tokens sequentially like an RNN
- Claims the token embedding already encodes where the word appears
- Thinks the feed-forward sublayer mixes information across positions
- Confuses positional encoding with the causal attention mask
- Says position only matters at the input, never deeper in the stack