skip to content

How do Bahdanau's additive attention score and Luong's multiplicative score differ?

level: middleimportance: should knowfreq 55%

answer

  1. one is a network, one is a product
  2. count the learned parameters in each
  3. the plain dot needs matching state sizes
  4. general inserts one learned matrix
  5. additive: project, add, tanh, read out

basics

~20 s

Bahdanau scores a decoder-encoder pair with a small one-hidden-layer network (tanh, then a learned vector), tolerating different sizes. Luong multiplies the two states directly, optionally through a learned matrix: cheaper, but the plain dot needs matching sizes.

solid answer

~50 s

Both produce one scalar per source position; they differ in the function. Bahdanau's **additive** score is `e_j = v^T tanh(W s + U h_j)`: the decoder state and the encoder state are each projected into a shared space, added, squashed by tanh, and read out by a learned vector. It carries three parameter blocks and handles mismatched state sizes for free. Luong's **multiplicative** family is a product — `dot` is `s^T h_j`, parameter-free but requiring equal dimensions, and `general` is `s^T W h_j`, one learned matrix that fixes a size mismatch and learns which directions should match. The practical difference is shape: multiplicative scoring is a single dense matrix product over all source positions, while additive scoring needs a per-pair projection and nonlinearity. They also wire differently — Bahdanau feeds the context into the recurrent step, Luong into the output layer.

go deeper

for a junior

Be ready to say that both produce one score per source position and that the difference is the function: a small network with a tanh in one case, a direct product of the two states in the other.

for a middle

Write both scores out and name the parameters: two projections plus a read-out vector for additive, nothing for dot, one matrix for general. Explain which variants tolerate mismatched encoder and decoder widths.

for a senior

Expect to justify a choice under real constraints — matrix-product scoring for throughput, general when widths differ or the identity metric is wrong — and to know that the two papers also wire the context into different places in the decoder.

for a principal

Own the tradeoff between expressiveness and hardware shape: a nonlinear per-pair score buys flexibility that a bilinear form often recovers with far better utilisation, and that argument, not parameter count, is what settles the design.

## Same slot, two functions Attention needs a function `score(s, h)` that turns a decoder state `s` and an encoder state `h_j` into one number. Everything downstream — softmax over source positions, weighted average — is identical. The two named families differ only in that function, and the naming comes from the two papers that introduced attention for translation. ## Bahdanau: additive (also called concat) scoring ``` e_j = v^T tanh(W s + U h_j) ``` Read it as a one-hidden-layer network over the pair. `W` maps the decoder state (size `d_dec`) into an attention space of size `d_a`; `U` maps the encoder state (size `d_enc`) into the same space; the two projections are **added**, hence *additive*; `tanh` makes the combination nonlinear; and `v`, a vector of size `d_a`, collapses it to a scalar. Parameter count: `d_a * d_dec` for `W`, `d_a * d_enc` for `U`, `d_a` for `v`. Two consequences follow immediately. First, `d_dec` and `d_enc` may be anything — the projections reconcile them, which mattered in the original design where the encoder was bidirectional and each `h_j` was the concatenation of a forward and a backward state, so twice the decoder's width. Second, the score is not linear in the pair; the tanh lets it express relevance patterns a plain inner product cannot. The cost is the shape of the computation. `W s` is computed once per output step, but `U h_j` is per source position (and can be precomputed once for the whole source, a common optimisation), and the addition, tanh and read-out are per pair — an elementwise nonlinearity over an `S x d_a` buffer at every output step. ## Luong: multiplicative scoring Three variants were proposed; two are the multiplicative ones. ``` dot: e_j = s^T h_j general: e_j = s^T W h_j ``` **dot** has no parameters at all. Relevance is cosine-like alignment in the existing state space, scaled by the two magnitudes. Its constraint is rigid: `d_dec` must equal `d_enc`, and the two spaces must already be comparable — the encoder and decoder must have learned to place related content in the same directions, which they can, since both are trained jointly. **general** inserts one matrix of size `d_dec x d_enc`. That restores flexibility on both counts: sizes may differ, and the model learns a bilinear form saying which decoder directions should match which encoder directions rather than assuming the identity. The reason multiplicative scoring is preferred on real hardware is that the whole set of `S` scores is one matrix-vector (or, batched over output steps, one matrix-matrix) product. There is no elementwise nonlinearity and no `S x d_a` intermediate. One caveat worth naming: the magnitude of an inner product grows with the dimension of the states, so with wide states raw dot scores can become large, and a softmax over large-magnitude scores is very peaked. The `general` variant's matrix can absorb some of that; designers otherwise control it explicitly. (Luong's third variant, *concat*, is the additive form; the family split is really additive versus multiplicative, not one paper versus the other.) ## The wiring difference candidates forget The scoring function is not the only difference between the two designs. - **Bahdanau** computes the alignment from the decoder state *before* the current recurrent update — the previous state `s_{t-1}` — and feeds the resulting context vector into that recurrent step as extra input. Context influences the state. - **Luong** computes the alignment from the decoder's *current* top-layer state `s_t`, then combines context and state (concatenate, then a linear map and tanh) to form the vector that feeds the output distribution. Context influences the prediction; an *input-feeding* variant also carries that combined vector into the next step's input. Both are trained end to end and neither ordering is universally better; interviewers ask about it to see whether you have read the mechanism rather than a one-line summary. ## Choosing - Same-width encoder and decoder, want the cheapest thing that works: `dot`. - Widths differ, or you want a learned notion of which directions match: `general`. - Small model, plenty of compute per score, or you want the extra expressiveness of a nonlinear score: additive. The difference in parameter count is small next to the rest of the model, so the decision is usually driven by shape compatibility and by how much per-pair work you are willing to do.

  • Your encoder is bidirectional so its states are twice the decoder's width — which scoring functions still work?
    Additive works unchanged: its two projection matrices map the differently sized states into a shared attention space. Luong's `general` also works, since its matrix is `d_dec x d_enc`. The plain `dot` does not — it needs equal dimensions, so you would have to project one side first, which is effectively `general` with the matrix named differently.
  • Where does the context vector enter the decoder in each design?
    In Bahdanau's, the context is computed from the previous decoder state and fed into the recurrent step as additional input, so it shapes the new state. In Luong's, it is computed from the current top-layer state and merged with it to form the vector that produces the output distribution; the input-feeding variant additionally carries that merged vector into the next step.
  • Why is multiplicative scoring usually faster in practice even though the parameter counts are comparable?
    All `S` multiplicative scores are one dense matrix product, which maps directly onto hardware built for exactly that. Additive scoring needs a per-pair sum, an elementwise tanh over an `S x d_a` buffer, and a read-out, so it materialises a much larger intermediate and spends time in memory-bound elementwise work rather than in the matrix unit.

saying these in an interview costs you the question

  • Says additive attention has no learned parameters
  • Claims the plain dot score works with any two state sizes
  • Describes additive scoring as simply adding the two state vectors
  • Thinks the two differ only in which paper named them
  • Says the general variant normalises the alignment weights

context