In transformer self-attention, what roles do the query, key and value vectors play?
answer
- Three views of the same token
- One searches, one advertises, one carries
- Dot product scores, softmax normalizes
- Weighted sum of value vectors
- Separate Q and K make it asymmetric
basics
~20 sEach token is projected into three vectors: a query saying what it is looking for, a key advertising what it offers, and a value carrying the content it contributes. Query-key dot products score matches, and those scores weight a sum of values.
solid answer
~50 sSelf-attention treats every token as both a searcher and a searchable item. Three learned linear projections turn each token's representation into a **query**, a **key** and a **value**. To compute one token's output, its query is dotted against every key in the sequence, producing a raw score per token; those scores are scaled and pushed through a softmax so they become non-negative weights summing to one; the output is the weighted sum of the corresponding value vectors. The retrieval analogy is exact enough to be useful: the query is the search, the keys are the index, the values are the payload — except the match is soft, so a token blends a bit of everything rather than picking one row. Queries and keys are separate projections on purpose, which makes attention asymmetric: a pronoun can look for a noun without every noun looking equally hard for that pronoun.
code
python · 14 linesimport numpy as np
def head(x, w_q, w_k, w_v):
q, k, v = x @ w_q, x @ w_k, x @ w_v
scores = q @ k.T / np.sqrt(k.shape[-1])
w = np.exp(scores - scores.max(-1, keepdims=True))
w /= w.sum(-1, keepdims=True)
return w @ v, w
rng = np.random.default_rng(0)
x = rng.normal(size=(5, 8)) # 5 tokens, model dim 8
wq, wk, wv = (rng.normal(size=(8, 4)) for _ in range(3))
out, weights = head(x, wq, wk, wv)
print(out.shape, weights.shape, weights[0].sum())go deeper
Be able to name the three projections and say plainly what each does: query searches, key advertises, value contributes. Then state the pipeline in order — dot product, scale, softmax, weighted sum of values.
Explain why queries and keys are separate matrices and what asymmetry that buys, and why the value projection is a third vector rather than reusing the key. Be precise that the weights are computed per input, not learned.
Show you can reason about shapes and cost: which matrices are parameters, which are activations, what stays fixed as sequence length grows. Be ready to say what you would inspect if a head's outputs looked degenerate.
Own the framing that attention is a learned, content-addressed routing layer, and that every architectural variant in the field is a tradeoff against the cost of that routing. Be able to argue where the mechanism's real capacity limits sit rather than reciting the formula.
## What problem attention is solving A language model has to build a representation of each token that depends on the other tokens around it. "Bank" means something different next to "river" than next to "savings". Earlier architectures propagated that information sequentially, one step at a time, so information from a distant token had to survive many hops. Self-attention lets every token read from every other token in one step, with learned, content-dependent weights over what to read. That is the whole idea; queries, keys and values are the machinery that implements it. ## Three projections of the same input The layer receives a matrix X of shape (n, d_model): n tokens, each a d_model-dimensional vector. Three learned weight matrices produce three new views of the same tokens: - Q = X W_Q — the **query**: what this token is looking for. - K = X W_K — the **key**: what this token offers to anyone looking. - V = X W_V — the **value**: the content this token contributes when it is attended to. Each projection maps d_model down to a smaller head dimension d_k (values sometimes use d_v, usually equal). Crucially, the same input row produces all three vectors — this is *self*-attention, as opposed to cross-attention, where the queries come from one sequence and the keys and values from another. Note what the projections do not depend on: the sequence length. W_Q, W_K and W_V are fixed-size matrices, so a head has the same parameter count whether it processes 10 tokens or 100,000. Only the activations grow with n. ## Scoring and mixing The computation is: Attention(Q, K, V) = softmax(Q K^T / sqrt(d_k)) V Read it left to right. Q K^T is an n-by-n matrix of dot products: entry (i, j) is how well token i's query matches token j's key. Dividing by sqrt(d_k) keeps those raw scores in a range where the softmax stays soft. The softmax is applied along each row, converting the scores in row i into a probability distribution over all positions — non-negative, summing to one. Multiplying that distribution by V produces the output for token i: a convex combination of every token's value vector, weighted by how relevant that token was judged to be. The softmax matters more than it looks. It forces competition: attention mass is a fixed budget of 1.0 per token, so paying more attention to one position necessarily means paying less to others. That is what makes a head able to *select* rather than merely average. ## Why queries and keys are separate A common instinct is that one projection would do — just dot the token representations against each other. That would make the score matrix symmetric, meaning token i attends to j exactly as much as j attends to i. Language is not symmetric. In "the trophy didn't fit in the suitcase because it was too big", the pronoun "it" needs to look hard at "trophy", but "trophy" has no particular reason to look at "it". Separate W_Q and W_K let a token advertise something different from what it searches for, which is exactly what asymmetric relations like reference, agreement and dependency require. The value projection is separate for a related reason: what makes a token *findable* need not be what makes it *useful* once found. A key might encode "I am a singular noun"; the value it contributes might encode the noun's semantics. Collapsing keys and values would force one vector to do both jobs. ## What comes out Each head outputs an (n, d_v) matrix — one vector per token, same sequence length as it started with. Attention never changes how many tokens there are; it changes what each token's vector contains. Those outputs are combined and added back into the token's running representation, so attention layers refine representations rather than replacing them. ## Common misreadings Attention weights are not parameters. They are computed fresh for every input from the content of that input; the learned parameters are only the projection matrices. Two different sentences of the same length produce completely different weight matrices from the same weights. Also, attention weights are not an explanation. A high weight from "it" to "trophy" is suggestive, but a single layer's weights are one term in a deep residual computation, and reading them as the model's reasoning is a well-known overclaim.
- What would break if you dropped the value projection and attended directly over the raw token vectors?The head would lose the ability to separate "what makes me findable" from "what I contribute". A token's retrieval signature and its payload would be forced into one vector, so any change that made a token easier to match would also change what downstream layers received from it. It also removes a learned degree of freedom: the value projection is where a head decides which subspace of the input it actually reads out.
- How does a head's parameter count change as the sequence gets longer?It does not. The learned parameters are the projection matrices, whose shapes depend on the model dimension and head dimension only. Sequence length affects activations and compute — the score matrix and the number of rows flowing through — not the weights. That is why the same checkpoint can run at 4K or 400K tokens without changing shape.
- How is cross-attention different from self-attention in terms of where Q, K and V come from?In self-attention all three come from the same sequence. In cross-attention the queries come from the sequence being generated while the keys and values come from a separate encoded sequence, so the decoder reads from that other sequence rather than from itself. The arithmetic — scaled dot product, softmax, weighted sum of values — is identical.
Think of a soft database lookup: the query is your search string, every row publishes a key as its index entry, and instead of returning one matching row you get a blend of all rows' payloads weighted by how well each key matched.
saying these in an interview costs you the question
- Calling attention weights learned parameters rather than per-input activations
- Saying queries and keys are the same thing with different names
- Claiming attention outputs one selected token instead of a weighted blend
- Believing attention changes the number of tokens in the sequence
- Treating attention weights as a faithful explanation of the model's reasoning