skip to content

Transformer Architecture

The architecture every modern LLM is built on: self-attention and multi-head attention, positional encodings, feed-forward sublayers, normalization, and why decoder-only stacks won for generation. Interviewers ask it to hear whether you can explain attention without hand-waving.

on this pageshow

explore

questions

20

Why does a transformer need positional encoding, given that attention sees every token?

level: juniorimportance: must knowfreq 74%

answer

  1. attention is order-blind
  2. a set, not a sequence
  3. permutation-equivariant outputs
  4. "dog bites man" vs "man bites dog"
  5. inject at input or in the scores

basics

~20 s

Self-attention treats its input as a set — reorder the tokens and each one gets the same representation back, just moved. Order has to be injected separately, or "the auditor approved the invoice" and "the invoice approved the auditor" are indistinguishable.

solid answer

~50 s

Self-attention computes, for every token, a weighted average of value vectors where the weights come from query-key dot products. Nothing in that computation refers to an index, so the layer is **permutation-equivariant**: shuffle the input and the outputs are the same vectors in shuffled order. Since meaning in language depends on order, position must be supplied. There are two places to supply it. The first is the input: add an absolute position vector (a fixed sinusoidal formula or a learned table) to each token embedding, once, before layer one. The second is the attention computation itself: bias or transform the query-key interaction so the score depends on the distance between two tokens — rotary encoding (RoPE) is the dominant form of this, applied at every layer. Modern decoder-only LLMs overwhelmingly use the second family, because relative structure is what language actually needs.

go deeper

for a junior

Be ready to say plainly that self-attention has no built-in notion of order, so position must be added — and to give a two-sentence example where reordering changes meaning but not the computation.

for a middle

Explain permutation equivariance concretely and name both injection points: absolute vectors added to embeddings versus a relative signal applied to queries and keys inside every attention layer.

for a senior

Expect to justify why the relative family displaced absolute in production decoder-only models, and to note that the causal mask already leaks a weak position signal that NoPE work exploits.

for a principal

Own the design call: what your task's structure actually needs — global coordinates or local offsets — and what a positional choice commits you to when sequence lengths later grow well past what you trained on.

## The problem: attention consumes a set, not a sequence A self-attention layer produces, for each position, a weighted sum of value vectors. The weights come from comparing that position's query vector against every key vector. Written out, that computation contains no reference to *where* a token sits — only to *what* its vectors contain. The consequence is precise: self-attention is **permutation-equivariant**. Permute the input tokens and you get exactly the same set of output vectors, permuted the same way. Nothing is lost and nothing is gained; the layer literally cannot tell the two orderings apart. The feed-forward sublayer does not rescue this: it applies the same transformation to each position independently, so it never mixes across positions. Layer normalization is also per-position. Attention is the only mixing operation in the stack, and it is order-blind by construction. So "the auditor approved the invoice" and "the invoice approved the auditor" yield identical bags of representations. Order information must be *injected*, because it cannot be *inferred* from the architecture. ## Two places to inject position **At the input, as absolute position.** Build one vector per position — either from a fixed formula (the original sinusoidal encoding) or from a learned lookup table with one row per index — and add it to the token embedding before the first layer. The token vector now carries "I am the word *invoice*" plus "I am at index 4", and attention can, in principle, read distance out of the difference. This happens once, at the bottom of the stack. **Inside attention, as relative position.** Instead of tagging the inputs, change how a query and a key interact so that the resulting score already depends on the gap between the two tokens. Learned relative-position biases add a per-offset scalar to the attention logits. ALiBi subtracts a penalty proportional to distance. RoPE rotates the query and key vectors by an angle proportional to their absolute positions, so that their dot product ends up depending only on the offset between them. This family is applied at every layer, to queries and keys, not to the token embeddings. ## Why the relative family won Three reasons. First, most of what order contributes to meaning is relative: which word modifies which, how far back the subject was, whether a bracket closed. Second, absolute schemes run out — a learned table has no row for a position it never trained on. Third, a relative signal reapplied at every layer keeps position available deep in the stack, rather than relying on it surviving dozens of residual updates from the embedding. As of mid-2026 essentially every open and frontier decoder-only LLM uses rotary encoding as its baseline; ALiBi belongs to an earlier generation of models. ## The causal-mask nuance worth knowing In a decoder-only model, the causal mask lets position *n* attend only to positions 0..*n*, so each position sees a different-sized prefix. That asymmetry is itself a weak position signal, and research on **NoPE** (no positional encoding) showed that decoder-only transformers can learn order-sensitive behaviour without any explicit encoding at all. This is why "no encoding" is a live architectural option in 2026 rather than an obvious mistake. A bidirectional encoder has no such asymmetry and genuinely cannot work without an encoding. ## When absolute position genuinely matters Relative is the right default for text, but not universally. A symbolic-music sequence model where each token is a note event inside a bar cares about absolute grid position — "beat one of bar seventeen" is a meaningful coordinate, not just an offset from the previous note. Fixed-length structured inputs, where slot *k* always means the same field, are similar. If your task has a genuine global coordinate system, an absolute scheme is not a legacy choice. ## What an interviewer listens for A clean answer says: attention is order-blind because it is a weighted set operation; position is injected either into the inputs (absolute) or into the query-key interaction (relative); modern LLMs use rotary encoding, which is the relative kind; and the causal mask supplies a weak implicit signal on top. Hand-waving that "the model reads the tokens in order" is the answer that fails.

  • Doesn't the causal mask already tell a decoder-only model where each token sits?
    Partly. Under a causal mask, position n attends over a prefix of size n+1, so the number of visible tokens varies with position and gives an implicit signal. Work on NoPE showed decoder-only stacks can exploit this and learn order-sensitive behaviour with no explicit encoding. It is a weak, indirect signal, which is why most models still add one — but it is why NoPE is a serious option rather than a bug.
  • Where in the network does the position signal actually enter, for absolute versus rotary schemes?
    An absolute encoding is added to the token embedding once, before the first layer, and must survive every residual update after that. A rotary or relative scheme is applied inside the attention computation, to queries and keys, in every layer — so position is re-injected dozens of times and never has to be preserved through the residual stream.
  • Are there sequence tasks where absolute position is more useful than relative offset?
    Yes. A symbolic-music model whose tokens are note events inside a metrical grid cares that something falls on beat one of bar seventeen — an absolute coordinate with meaning, not just a gap from the previous event. Fixed-schema structured inputs behave the same way. For natural language, relative structure dominates, which is why relative and rotary schemes became the default there.

Attention is like reading a pile of index cards spread on a table: you can see all of them at once, but nothing tells you which was on top. Positional encoding is the page number written on each card.

saying these in an interview costs you the question

  • Says a transformer processes tokens sequentially like an RNN
  • Claims the token embedding already encodes where the word appears
  • Thinks the feed-forward sublayer mixes information across positions
  • Confuses positional encoding with the causal attention mask
  • Says position only matters at the input, never deeper in the stack

context

open as a page

In transformer self-attention, what roles do the query, key and value vectors play?

level: juniorimportance: must knowfreq 86%

basics

~20 s

Each token is projected into three vectors: a query saying what it is looking for, a key advertising what it offers, and a value carrying the content it contributes. Query-key dot products score matches, and those scores weight a sum of values.

open as a page

Does FlashAttention change attention's output, and where does its speedup come from?

level: middleimportance: must knowfreq 58%

basics

~20 s

No — FlashAttention computes exactly the same attention, bit-for-bit equivalent up to floating-point reordering. It is faster because it tiles the computation in fast on-chip memory and never writes the full score matrix out to the GPU's main memory, so it moves far less data.

open as a page

In transformer LLMs, how do GQA and MQA differ from multi-head attention?

level: middleimportance: must knowfreq 72%

basics

~20 s

Multi-head attention gives every query head its own key and value projections. Multi-query attention makes all query heads share a single key/value head. Grouped-query attention sits between them: query heads are split into groups, and each group shares one key/value head.

open as a page

An MoE model has 670B total and 37B active parameters — which number sizes GPU memory?

level: middleimportance: must knowfreq 66%

basics

~20 s

Total parameters size memory. Every expert's weights must be resident because any token may route to any expert. Active parameters size compute — the arithmetic and latency per token. A sparse model is therefore memory-hungry like a huge model but computes like a small one.

open as a page

In a transformer, what does a Mixture-of-Experts layer replace, and how does routing work?

level: middleimportance: must knowfreq 75%

basics

~20 s

A Mixture-of-Experts layer swaps a transformer block's feed-forward sublayer for many parallel feed-forward experts plus a small router. The router scores experts for each token and activates only the top few, so most expert weights compute nothing for that token.

open as a page

How does RoPE make attention depend on relative offset using absolute positions?

level: middleimportance: must knowfreq 68%

basics

~20 s

RoPE rotates each consecutive pair of dimensions in a query and a key by an angle proportional to that token's absolute position. Rotations compose, so the resulting dot product depends only on the difference between the two positions.

open as a page

What does the causal mask do in a decoder-only transformer, and what breaks without it?

level: middleimportance: must knowfreq 72%

basics

~20 s

The causal mask sets attention scores for future positions to negative infinity before the softmax, so each position can attend only to itself and earlier tokens. Without it, training leaks the answer: a position can see the token it is meant to predict.

open as a page

Why does scaled dot-product attention divide the QK scores by sqrt(d_k)?

level: middleimportance: must knowfreq 68%

basics

~20 s

Dot products of high-dimensional vectors grow in magnitude with the dimension, so raw scores get large as d_k grows. Large scores push the softmax toward a near one-hot distribution where gradients almost vanish. Dividing by sqrt(d_k) holds the scores near unit scale.

open as a page

Why does quadrupling a transformer's context length raise attention cost roughly 16x?

level: seniorimportance: must knowfreq 64%

basics

~20 s

Self-attention scores every token against every other token, so an n-token sequence produces an n-by-n matrix. Quadrupling n multiplies the number of entries by sixteen, in both the compute to fill the matrix and the memory to hold it during training.

open as a page

How do sinusoidal and learned absolute position embeddings differ beyond the trained length?

level: middleimportance: should knowfreq 48%

basics

~20 s

A learned table stores one trained vector per position, so index 5,000 in a model trained to 2,048 does not exist at all. A sinusoidal encoding is a formula, defined at any index — but the model still never trained on those patterns.

open as a page

Why does a transformer layer use many attention heads instead of one wide head?

level: middleimportance: should knowfreq 58%

basics

~20 s

One softmax distribution can concentrate on only one thing at a time, so a single head must choose which relation to encode. Splitting the same width into several heads lets a layer attend to several relations in parallel, each in its own subspace, then combine the results.

open as a page

What does Multi-head Latent Attention compress, and how is that different from quantizing the cache?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Multi-head Latent Attention projects each token's keys and values down into one small learned latent vector and retains only that. Per-head keys and values are reconstructed on the fly by up-projection. The compression lives in the trained weights, not in the number format.

open as a page

Why does a trainable sparse-attention selector beat a fixed strided pattern at long context?

level: seniorimportance: should knowfreq 33%

basics

~20 s

A fixed pattern decides which earlier positions each query may see using position alone, so it discards relevant tokens that fall outside the pattern. A trainable selector scores earlier blocks by content and picks per query, and because it is trained jointly the model adapts to the sparsity.

open as a page

Three of 64 experts take most tokens in your MoE — what breaks and how do you fix it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

That is routing collapse. Most of the model's capacity goes unused while the hot experts saturate; in a deployment that shards experts across devices, the devices holding them become stragglers and throughput drops to their speed. Fixes push load back toward the idle experts.

open as a page

Why do modern MoE models pair many fine-grained experts with an always-on shared expert?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Splitting the feed-forward budget into many small experts multiplies the routing combinations available at the same active-parameter cost, letting each expert specialize narrowly. An always-on shared expert absorbs the general knowledge every token needs, so the routed experts do not each relearn it.

open as a page

How do NTK-aware scaling, YaRN and LongRoPE rescale RoPE frequencies differently?

level: seniorimportance: should knowfreq 42%

basics

~20 s

All three keep far-out positions inside the angle range a model trained on. NTK-aware scaling raises the rotary base so slow dimensions absorb most of the compression; YaRN treats dimensions differently by wavelength band and adds an attention-temperature correction; LongRoPE searches a per-dimension factor instead of deriving one.

open as a page

When would you interleave linear-recurrent layers with full attention in a long-context model?

level: principalimportance: should knowfreq 26%

basics

~20 s

When the workload needs a very long input but rarely needs exact recall of an arbitrary earlier token. Recurrent layers keep a fixed-size state, so cost per token stops growing with length; the periodic full-attention layers are kept precisely to preserve the exact lookup that a fixed state loses.

open as a page

When is a sparse Mixture-of-Experts the wrong choice for a team self-hosting a model?

level: principalimportance: should knowfreq 34%

basics

~20 s

When your binding constraint is memory rather than compute. Sparsity trades accelerator memory and operational complexity for cheap FLOPs per token, so a low-traffic deployment, a memory-tight footprint, or a weak interconnect all pay the full bill for a discount they never collect.

open as a page

Why would a long-context model leave some head dimensions unrotated by RoPE?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Rotating a dimension makes its contribution to the attention score oscillate and fall away with distance. Leaving part of the representation unrotated gives the model a position-agnostic subspace that can match purely on content, at any distance.

open as a page