skip to content

How do sinusoidal and learned absolute position embeddings differ beyond the trained length?

level: middleimportance: should knowfreq 48%

answer

  1. one is a table, one is a formula
  2. table has no row past max length
  3. defined everywhere is not trained everywhere
  4. both are absolute, so offsets are inferred
  5. fluent output, degraded accuracy

basics

~20 s

A learned table stores one trained vector per position, so index 5,000 in a model trained to 2,048 does not exist at all. A sinusoidal encoding is a formula, defined at any index — but the model still never trained on those patterns.

solid answer

~60 s

Both are **absolute** schemes: a vector representing "I am at index m" is added to the token embedding before the first layer. The learned variant is a lookup table with one trainable row per position, capped at the training length. Feed a model trained on 2,048-token sequences an 8,000-token machine-log file and position 5,000 simply has no row — you must truncate, window, or splice in a garbage vector, and quality falls off a cliff at the boundary. The sinusoidal variant computes each dimension as a sine or cosine of the index divided by a geometrically spaced wavelength, so it is *defined* everywhere. That was advertised as extrapolation, and it mostly did not deliver: the model trained only on the angle combinations that occur below its training length, so beyond it the encodings are in-distribution as numbers but out-of-distribution as inputs, and accuracy degrades even though generation continues. Neither scheme makes attention depend directly on the gap between tokens — that limitation is what drove the field toward relative and rotary encodings.

code

python · 13 lines
python
import numpy as np

def sinusoidal(n_pos, d_model):
    pos = np.arange(n_pos)[:, None]
    i = np.arange(d_model // 2)[None, :]
    angle = pos / np.power(10000.0, (2 * i) / d_model)
    pe = np.zeros((n_pos, d_model))
    pe[:, 0::2] = np.sin(angle)
    pe[:, 1::2] = np.cos(angle)
    return pe

print(sinusoidal(8192, 64).shape)
print(sinusoidal(8192, 64)[5000][:4])

go deeper

for a junior

Know that both schemes tag each token with its index before layer one, and that a learned table has a hard maximum length while the sinusoidal formula does not.

for a middle

Explain the geometric spread of sinusoidal wavelengths, and state precisely what happens at position 5,000 in a model whose table stops at 2,047 — no row, so truncate or window.

for a senior

Be able to argue why sinusoidal extrapolation underdelivered in practice: the encodings are defined but the weights never trained on that region, so fluency survives and accuracy does not.

for a principal

Own the consequence for roadmap: an absolute scheme caps a model's usable length at a value you fix at training time, which is a commitment you cannot cheaply revisit later.

## What "absolute" means here An absolute positional encoding produces one vector per index — position 0 gets one vector, position 1 another — and adds it to the token embedding at the bottom of the stack. From there on, position is just part of the residual stream, and every attention layer has to *infer* the distance between two tokens from the two absolute tags it can see. That indirection is the structural weakness of the whole family, independent of how the vectors are produced. ## The sinusoidal construction The original transformer encoding fills each pair of dimensions with a sine and a cosine of the position divided by a wavelength, with wavelengths spaced geometrically from short to very long across the dimensions. Early dimensions oscillate quickly and distinguish neighbouring positions; late dimensions oscillate so slowly that they act as a coarse "early or late in the sequence" indicator. Two properties follow. It is **parameter-free** — nothing is learned, so it costs no capacity and no training signal per position. And it is **total**: the formula evaluates at index 5,000 as happily as at index 5. That second property is why the original paper speculated it might let a model extrapolate to longer sequences than it saw in training. ## The learned table The alternative, adopted by several influential encoder models, is simply an embedding matrix with one row per position and a fixed maximum length, trained by gradient descent like any other embedding. It is flexible — the model can shape each position's vector however the data rewards — and it costs a modest number of parameters. Its failure mode is not gradual, it is structural. Take a model trained on 2,048-token sequences and hand it an 8,000-token machine-log file. Position 5,000 has no row in the table. There is nothing to look up. Your options are to truncate the input, to slide a window over it, or to fabricate a vector by repeating, tiling or interpolating rows — and every one of those changes the input distribution abruptly at the boundary, which is exactly where quality is observed to collapse. A learned absolute table hard-caps the model's length; it does not degrade past it. ## Why sinusoidal's promised extrapolation underdelivered Being *defined* at a position is not the same as being *trained* at it. During training on sequences up to length L, the slow dimensions never complete even a fraction of a cycle beyond L, so the specific combinations of angles that occur at position 5,000 were never seen by any weight in the network. The encodings are valid numbers in a region of input space the model has no experience of. Empirically, models with sinusoidal encodings pushed past their training length stay fluent — the tokens keep coming and the grammar stays intact — while accuracy on anything requiring precise reference degrades. "It kept generating" is not "it stayed correct", and that distinction is the one to say out loud in an interview. ## The deeper limitation both share When positions 5 and 8 are tagged absolutely, an attention head that wants to know "these are three apart" must recover that from two independent vectors, and it must do so identically for the pair (500, 503). Nothing in the scheme guarantees that. The relative families were introduced precisely to make that guarantee structural: learned relative biases index the attention logit by offset, ALiBi subtracts a penalty proportional to distance, and rotary encoding rotates queries and keys so their dot product depends only on the gap. Rotary won and is the mid-2026 baseline; ALiBi is historical. ## When absolute schemes are still reasonable Fixed-length inputs with a genuine global coordinate — a symbolic-music model where token position maps to a beat inside a bar, or a structured record where slot *k* always means the same field — are legitimately served by absolute encodings, learned ones included. The failure discussed above only bites when you need lengths you did not train on. If your maximum length is a hard product constraint, a learned table is not a defect. ## What interviewers probe Expect the follow-up "sinusoidal was supposed to extrapolate — did it?" The honest answer is: it is defined everywhere, it does not crash, and it still degrades, because the model trained on a bounded region of the encoding space. Candidates who claim sinusoidal encodings give free length generalization are repeating a 2017 hope rather than the observed result.

  • Sinusoidal encodings are defined at every index — so why don't models with them extrapolate cleanly?
    Because the model, not the formula, is the limit. Training on sequences up to length L only ever exposes the network to the angle combinations that occur below L; the slow dimensions never reach the values they take at far positions. Past L the inputs are valid numbers from an unseen region, so output stays fluent while accuracy on precise reference degrades.
  • If a learned table has no row for position 5,000, what do practitioners actually do?
    Truncate the input, slide a fixed-size window over it with overlap, or retrain with a larger table. Fabricating rows by tiling or interpolating existing ones creates a distribution break exactly at the boundary and is not reliable. The structural fix is to move off absolute encodings entirely, which is what rotary schemes did.
  • Why does an absolute scheme make relative distance harder for a head to use?
    The head only sees two independent position tags and must derive the gap from their combination, learning that separately for every region of the sequence. Nothing guarantees the pair (5, 8) is treated like the pair (500, 503). Relative and rotary schemes make that equivalence structural rather than something the model must learn and then generalize.

saying these in an interview costs you the question

  • Claims sinusoidal encodings give free extrapolation to any length
  • Says a learned table degrades gradually past its maximum length
  • Thinks sinusoidal encodings make attention depend on relative distance
  • Confuses the encoding being defined at a position with the model being trained there
  • Assumes fluent output past the training length means the model is still accurate

context