skip to content

In transformer LLMs, how do GQA and MQA differ from multi-head attention?

level: middleimportance: must knowfreq 72%

answer

  1. Query heads stay put; something else shrinks
  2. Two named variants, one shared axis
  3. Groups of query heads share one KV head
  4. MHA and MQA are the endpoints
  5. Memory dial, not a FLOP dial

basics

~20 s

Multi-head attention gives every query head its own key and value projections. Multi-query attention makes all query heads share a single key/value head. Grouped-query attention sits between them: query heads are split into groups, and each group shares one key/value head.

solid answer

~50 s

All three keep the same number of **query** heads; they differ only in how many distinct **key/value** heads exist. In MHA, heads and KV heads are one-to-one. In MQA, every query head reads one shared K/V head, which shrinks what has to be stored and re-read for every generated token, but concentrates all retrieval into a single subspace and costs quality. GQA generalises both: with 64 query heads and 8 KV groups, eight query heads share each K/V head, so the stored keys and values shrink eightfold while each group still keeps its own view of the sequence. GQA with one group is MQA; with as many groups as heads it is MHA. It is cheap to adopt because an existing MHA checkpoint can be converted by mean-pooling the K/V projections within each group and briefly continuing training, rather than pretraining from scratch.

code

python · 6 lines
python
import numpy as np

n_q_heads, n_kv_heads, seq, d = 8, 2, 4, 16
k = np.random.randn(n_kv_heads, seq, d)
k_broadcast = np.repeat(k, n_q_heads // n_kv_heads, axis=0)
print(k.shape, k_broadcast.shape)

go deeper

for a junior

Know the three names and the one axis they sit on: MHA gives every head its own keys and values, MQA shares one set across all heads, GQA shares one set per group of heads. Say plainly that this is about shrinking what the model keeps around while generating.

for a middle

Be ready to explain that query head count is unchanged and only key/value head count varies, that GQA with one group is MQA and with G equal to H is MHA, and that the shared head is broadcast across its group before scores are computed. Name the quality-versus-state direction.

for a senior

Expect to justify a group count for a real serving target and to explain why the benefit shows up as throughput and batch size rather than fewer operations. Knowing that an MHA checkpoint can be uptrained into GQA by mean-pooling projections is the detail that separates reading about it from shipping it.

for a principal

Own the framing that head sharing is a coarse, bounded lever: it divides the retained state by a constant but does not change how it grows with sequence length. Be able to say when that ceiling forces a move to latent compression, learned sparsity or hybrid layers instead, and what each of those costs in training risk.

## The quantity that actually differs Attention computes, per head, a set of query vectors that are compared against key vectors and used to weight value vectors. Nothing about that computation changes across MHA, MQA and GQA. What changes is **how many distinct key/value projections the layer has**, and therefore how many distinct key and value tensors must be kept around while the model generates text token by token. During autoregressive generation, the keys and values for every token already produced are retained so the next token does not have to recompute them. That retained state is proportional to the number of key/value heads. Query heads, by contrast, are recomputed for the single new token each step and cost nothing to retain. So the number of query heads is a quality dial; the number of key/value heads is a memory dial. Head sharing exists to separate the two. ## The three points on the spectrum **Multi-head attention (MHA)** is the original design: H query heads, H key heads, H value heads, one-to-one. Every head has its own subspace for both asking and answering. Maximum expressiveness, maximum retained key/value state. **Multi-query attention (MQA)** keeps H query heads but has exactly one key head and one value head, shared by all of them. The retained state collapses by a factor of H. The cost is real: every query head now has to phrase its lookup against the same key subspace, and models trained this way tend to lose a measurable amount of quality and can be less stable to train. **Grouped-query attention (GQA)** interpolates. Query heads are partitioned into G groups; each group has its own key head and value head. G = H recovers MHA, G = 1 recovers MQA. Typical production settings put G somewhere between 4 and 8 for models with 32 to 128 query heads. Empirically this recovers most of the quality gap to MHA while retaining most of MQA's memory saving, which is why GQA became the conservative default rather than an exotic option. ## Why the memory dial matters more than the FLOP dial A common misreading is that head sharing is a compute optimisation. It is not, primarily. Sharing K/V heads does not reduce the number of query-key comparisons — every query head still attends over the whole sequence, so the score computation is unchanged. What shrinks is the amount of state that must be held and re-read from memory on every single decoding step. Generation is dominated by moving that state, not by the arithmetic on it, so reducing it translates into higher throughput and larger batch sizes rather than fewer operations. In implementation terms the shared key/value head is simply broadcast (repeated) across the query heads in its group before the score computation. ## Conversion rather than retraining A practically important property of GQA is that it can be retrofitted. Given a trained MHA checkpoint, you can mean-pool the key projection matrices of the heads in each intended group, do the same for the value projections, and then continue training on a small fraction of the original pretraining budget. The resulting model behaves close to a natively-trained GQA model. This is why GQA spread across model families quickly: labs did not have to bet a full pretraining run on it. ## Where the spectrum ends Head sharing is a coarse instrument — it reduces the number of key/value subspaces but each surviving one is still full-width and still grows with sequence length. That ceiling is what pushed the field past GQA. Two later directions attack the same problem differently: compressing keys and values into a learned low-rank latent so the stored object is smaller than any head, and making attention itself selective so a query only consults part of the sequence. GQA remains the safe baseline, and as of mid-2026 it is still what most models ship when they are not chasing extreme context lengths; the alternatives are what appear in models built specifically for very long inputs. ## Answering it well State the invariant first — query head count is unchanged, key/value head count is the variable — then place the three names on that one axis, then name the tradeoff direction: fewer KV heads means less retained state and faster decoding, more KV heads means more expressive retrieval. Mentioning uptraining from an MHA checkpoint signals you know why the industry actually adopted it.

  • If GQA does not reduce the number of query-key comparisons, where does the speedup during generation actually come from?
    From memory traffic. Each decoding step re-reads all retained keys and values for the sequence so far, and that read dominates the step because the arithmetic per byte is tiny. Cutting the number of key/value heads cuts what must be read, which raises throughput and lets you hold more concurrent sequences. The score arithmetic itself is essentially unchanged.
  • How would you pick the number of groups for a new model, and what fails at the extremes?
    Treat it as a quality-versus-state tradeoff measured empirically: sweep a few group counts at small scale and look at both validation loss and the retained-state size at your target context. One group (MQA) usually shows a real quality drop and can be less stable to train; groups equal to head count gives no saving at all. Most models land at four to eight groups.
  • Can you convert an existing multi-head checkpoint to GQA without pretraining again?
    Yes. Mean-pool the key projections of the heads that will form each group, do the same for the value projections, then continue training on a small fraction of the original pretraining compute. The converted model recovers close to native-GQA quality, which is why this became the standard adoption path rather than a fresh pretraining run.

saying these in an interview costs you the question

  • Says GQA reduces the number of query heads
  • Claims head sharing changes the attention formula itself
  • Thinks the win is fewer FLOPs rather than less retained state
  • Assumes MQA is free with no quality cost
  • Confuses grouping query heads with routing tokens to experts

context