skip to content

Matrix Multiplication

How (m x n) and (n x p) shapes must line up, why AB rarely equals BA, the (AB)^T = B^T A^T rule, and reading a matrix-vector product as a linear map. Shape questions are a standard screen.

on this pageshow

questions

5

For the matrix product AB, what shapes must A and B have, and what shape is the result?

level: juniorimportance: must knowfreq 86%

answer

  1. not entrywise like addition
  2. one dimension is shared by both
  3. rows on the left, columns on the right
  4. the shared dimension disappears in the result

basics

~20 s

Matrix multiplication needs matching inner dimensions: if A is m x n and B is n x p, then AB is m x p. Entry (i,j) is the dot product of row i of A with column j of B.

solid answer

~50 s

The product `AB` is defined only when the number of columns of `A` equals the number of rows of `B`. Writing the shapes side by side, `(m x n)(n x p)`, the two inner numbers must agree; they cancel, and the result takes the outer numbers, `m x p`. The entry in row `i`, column `j` of `AB` is the inner product of row `i` of `A` with column `j` of `B`: sum over `k` of `A[i,k] * B[k,j]`, so the shared dimension `n` is exactly the length of the vectors being dotted. That is why multiplying a `3 x 2` by another `3 x 2` fails — 2 does not equal 3 — even though the two operands have identical shapes and could be added entrywise. Matching shapes is the rule for addition, not for multiplication.

go deeper

for a junior

Be able to state the rule instantly and apply it: inner dimensions must match, the outer ones survive. Practise writing the shape chain before touching any arithmetic, and know that addition and multiplication have different shape rules.

for a middle

Explain where the rule comes from — each output entry is a dot product between a row and a column, so the shared dimension is the length of those vectors. Be ready to work an entry out by hand and to compare row-vector and column-vector orientations.

for a senior

Show that you debug shape errors structurally rather than by trial and error: read the shape chain end to end, identify which operand is oriented wrongly, and decide whether the fix is a transpose or a reordering. Mention the mnp cost when it matters.

for a principal

Frame shape discipline as an interface question. Fixing a convention for how data is laid out — observations as rows or as columns — removes a whole class of silent errors across a team, and you should be able to argue for one convention and enforce it consistently.

## The definition Matrix multiplication is defined entry by entry, and every shape rule follows from that definition rather than from convention. Let `A` have `m` rows and `n` columns, and let `B` have `n` rows and `p` columns. The product `C = AB` has `m` rows and `p` columns, and its entries are ``` C[i,j] = A[i,1]*B[1,j] + A[i,2]*B[2,j] + ... + A[i,n]*B[n,j] ``` which is the sum over `k` from 1 to `n` of `A[i,k] * B[k,j]`. In words: **entry (i,j) of the product is the dot product of row i of the left matrix with column j of the right matrix.** That sentence contains the whole shape rule. A dot product is only defined between two vectors of the same length. Row `i` of `A` has length `n` (one number per column of `A`); column `j` of `B` has length `n` (one number per row of `B`). So the number of columns of `A` must equal the number of rows of `B`. Nothing else has to match. ## Reading shapes off the page A reliable habit is to write the shapes next to each other with the operands in order: ``` (m x n)(n x p) -> m x p ``` The two inner numbers must be equal; they are consumed by the summation and vanish. The two outer numbers survive and become the shape of the answer. Some worked cases: - `(4 x 7)(7 x 3)` is fine and gives `4 x 3`. - `(3 x 2)(3 x 2)` is **not** defined: inner numbers 2 and 3 disagree. - `(1 x 4)(4 x 1)` gives `1 x 1` — a single number, the dot product of the two vectors. - `(4 x 1)(1 x 4)` gives `4 x 4` — an outer product, a whole matrix from the same two vectors. The last two lines are worth staring at. The same pair of length-4 vectors produces either a scalar or a 4-by-4 matrix depending purely on the order and orientation, which is the first concrete sign that order matters in matrix multiplication. ## Why identical shapes are not the rule A very common beginner error is to assume `A` and `B` must have the same shape, by analogy with addition. Matrix addition **is** entrywise: `A + B` requires identical shapes and adds corresponding cells. Matrix multiplication is not entrywise at all — it is a structured collection of dot products, and it exists to represent the *composition of linear maps*. If `B` maps vectors from a `p`-dimensional space into an `n`-dimensional space and `A` maps that `n`-dimensional space into an `m`-dimensional one, then `AB` maps the `p`-dimensional space straight into the `m`-dimensional one. The shape rule is the algebraic shadow of "the output space of the second map must be the input space of the first". There *is* an entrywise product of two equally shaped matrices (often called the elementwise or Hadamard product), but it is a different operation with different rules, and writing `AB` never means it. ## Counting the work Each of the `m * p` entries of the result costs `n` multiplications and `n - 1` additions, so the standard algorithm performs about `m * n * p` multiply-add operations. That count is worth remembering: it is the basis for reasoning about which order to evaluate a longer chain of products, and it explains why the shared inner dimension, not the visible output size, usually drives the cost. ## Special shapes - **Square matrices** (`n x n`) can always be multiplied by each other, in either order, and both products are `n x n`. - **Matrix times column vector**: `(m x n)(n x 1)` gives an `m x 1` column — the most common case in practice. - **Row vector times matrix**: `(1 x m)(m x n)` gives a `1 x n` row. - Vectors are just matrices with a 1 in one slot, so they obey exactly the same rule; most "dimension mismatch" confusion comes from forgetting whether a vector is a row or a column. ## Checking your own work Before computing anything by hand, write the shape chain. If you can carry `(m x n)(n x p)(p x q)` through and the inner numbers pair up all the way along, the expression is well formed and the answer is `m x q`. If any adjacent pair disagrees, the expression is meaningless and no amount of arithmetic will rescue it — one of the operands needs to be transposed or the order needs to change.

  • Why does multiplying a 3 x 2 matrix by another 3 x 2 matrix fail, when adding them is fine?
    Addition is entrywise and needs identical shapes. Multiplication needs the columns of the left operand to match the rows of the right: here that is 2 against 3, so the row-by-column dot products have no common length. Transposing one operand fixes it — `(3 x 2)(2 x 3)` gives `3 x 3`, and `(2 x 3)(3 x 2)` gives `2 x 2`.
  • If A is 3 x 2 and B is 2 x 3, are both AB and BA defined, and are they the same size?
    Both are defined, but they are different objects. `AB` is `(3 x 2)(2 x 3)` and comes out `3 x 3`; `BA` is `(2 x 3)(3 x 2)` and comes out `2 x 2`. They cannot even be compared entry by entry, which is the crudest demonstration that matrix multiplication does not commute.
  • Roughly how many multiply-add operations does the standard algorithm need for an (m x n) by (n x p) product?
    About `m * n * p`. There are `m * p` output entries, and each is a dot product of length `n`, costing `n` multiplications and `n - 1` additions. So the shared inner dimension multiplies the cost even though it never appears in the output shape.

Think of each matrix as a pipe with an input width and an output width. You can only join two pipes when the first pipe's output width equals the second's input width; the joint itself is invisible from outside.

saying these in an interview costs you the question

  • Claims both matrices must have the same shape, as in addition
  • Multiplies corresponding entries instead of rows against columns
  • Says the result takes the shape of the larger operand
  • Keeps the shared inner dimension in the output shape
  • Dots columns of the left matrix with rows of the right

context

open as a page

Why can the matrix products AB and BA differ, even when both are defined?

level: middleimportance: must knowfreq 71%

basics

~20 s

A matrix product is a composition of linear maps, and composition depends on order. In AB the right factor acts first, so rotating then stretching is a different transformation from stretching then rotating. Only special pairs commute.

open as a page

What does the matrix-vector product Ax compute in terms of the columns of A?

level: middleimportance: should knowfreq 56%

basics

~20 s

Ax is a linear combination of the columns of A, weighted by the entries of x. Applying A to the basis vector with a single 1 in slot j returns column j, so the columns are where the axes go.

open as a page

Why does the transpose of a matrix product equal B^T A^T rather than A^T B^T?

level: middleimportance: should knowfreq 47%

basics

~20 s

Transposing swaps each matrix's row and column counts, so the factors must reverse for the shapes to line up. Entry (i,j) of (AB)^T is entry (j,i) of AB, which is row i of B^T dotted with column j of A^T.

open as a page

How does associativity of matrix multiplication let you cut the cost of a three-matrix chain?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Associativity means (AB)C and A(BC) give identical results, so the grouping is yours to choose. An (m x n) by (n x p) product costs about mnp multiply-adds, so pick the grouping whose intermediate is smallest.

open as a page