For the matrix product AB, what shapes must A and B have, and what shape is the result?
answer
- not entrywise like addition
- one dimension is shared by both
- rows on the left, columns on the right
- the shared dimension disappears in the result
basics
~20 sMatrix multiplication needs matching inner dimensions: if A is m x n and B is n x p, then AB is m x p. Entry (i,j) is the dot product of row i of A with column j of B.
solid answer
~50 sThe product `AB` is defined only when the number of columns of `A` equals the number of rows of `B`. Writing the shapes side by side, `(m x n)(n x p)`, the two inner numbers must agree; they cancel, and the result takes the outer numbers, `m x p`. The entry in row `i`, column `j` of `AB` is the inner product of row `i` of `A` with column `j` of `B`: sum over `k` of `A[i,k] * B[k,j]`, so the shared dimension `n` is exactly the length of the vectors being dotted. That is why multiplying a `3 x 2` by another `3 x 2` fails — 2 does not equal 3 — even though the two operands have identical shapes and could be added entrywise. Matching shapes is the rule for addition, not for multiplication.
go deeper
Be able to state the rule instantly and apply it: inner dimensions must match, the outer ones survive. Practise writing the shape chain before touching any arithmetic, and know that addition and multiplication have different shape rules.
Explain where the rule comes from — each output entry is a dot product between a row and a column, so the shared dimension is the length of those vectors. Be ready to work an entry out by hand and to compare row-vector and column-vector orientations.
Show that you debug shape errors structurally rather than by trial and error: read the shape chain end to end, identify which operand is oriented wrongly, and decide whether the fix is a transpose or a reordering. Mention the mnp cost when it matters.
Frame shape discipline as an interface question. Fixing a convention for how data is laid out — observations as rows or as columns — removes a whole class of silent errors across a team, and you should be able to argue for one convention and enforce it consistently.
## The definition Matrix multiplication is defined entry by entry, and every shape rule follows from that definition rather than from convention. Let `A` have `m` rows and `n` columns, and let `B` have `n` rows and `p` columns. The product `C = AB` has `m` rows and `p` columns, and its entries are ``` C[i,j] = A[i,1]*B[1,j] + A[i,2]*B[2,j] + ... + A[i,n]*B[n,j] ``` which is the sum over `k` from 1 to `n` of `A[i,k] * B[k,j]`. In words: **entry (i,j) of the product is the dot product of row i of the left matrix with column j of the right matrix.** That sentence contains the whole shape rule. A dot product is only defined between two vectors of the same length. Row `i` of `A` has length `n` (one number per column of `A`); column `j` of `B` has length `n` (one number per row of `B`). So the number of columns of `A` must equal the number of rows of `B`. Nothing else has to match. ## Reading shapes off the page A reliable habit is to write the shapes next to each other with the operands in order: ``` (m x n)(n x p) -> m x p ``` The two inner numbers must be equal; they are consumed by the summation and vanish. The two outer numbers survive and become the shape of the answer. Some worked cases: - `(4 x 7)(7 x 3)` is fine and gives `4 x 3`. - `(3 x 2)(3 x 2)` is **not** defined: inner numbers 2 and 3 disagree. - `(1 x 4)(4 x 1)` gives `1 x 1` — a single number, the dot product of the two vectors. - `(4 x 1)(1 x 4)` gives `4 x 4` — an outer product, a whole matrix from the same two vectors. The last two lines are worth staring at. The same pair of length-4 vectors produces either a scalar or a 4-by-4 matrix depending purely on the order and orientation, which is the first concrete sign that order matters in matrix multiplication. ## Why identical shapes are not the rule A very common beginner error is to assume `A` and `B` must have the same shape, by analogy with addition. Matrix addition **is** entrywise: `A + B` requires identical shapes and adds corresponding cells. Matrix multiplication is not entrywise at all — it is a structured collection of dot products, and it exists to represent the *composition of linear maps*. If `B` maps vectors from a `p`-dimensional space into an `n`-dimensional space and `A` maps that `n`-dimensional space into an `m`-dimensional one, then `AB` maps the `p`-dimensional space straight into the `m`-dimensional one. The shape rule is the algebraic shadow of "the output space of the second map must be the input space of the first". There *is* an entrywise product of two equally shaped matrices (often called the elementwise or Hadamard product), but it is a different operation with different rules, and writing `AB` never means it. ## Counting the work Each of the `m * p` entries of the result costs `n` multiplications and `n - 1` additions, so the standard algorithm performs about `m * n * p` multiply-add operations. That count is worth remembering: it is the basis for reasoning about which order to evaluate a longer chain of products, and it explains why the shared inner dimension, not the visible output size, usually drives the cost. ## Special shapes - **Square matrices** (`n x n`) can always be multiplied by each other, in either order, and both products are `n x n`. - **Matrix times column vector**: `(m x n)(n x 1)` gives an `m x 1` column — the most common case in practice. - **Row vector times matrix**: `(1 x m)(m x n)` gives a `1 x n` row. - Vectors are just matrices with a 1 in one slot, so they obey exactly the same rule; most "dimension mismatch" confusion comes from forgetting whether a vector is a row or a column. ## Checking your own work Before computing anything by hand, write the shape chain. If you can carry `(m x n)(n x p)(p x q)` through and the inner numbers pair up all the way along, the expression is well formed and the answer is `m x q`. If any adjacent pair disagrees, the expression is meaningless and no amount of arithmetic will rescue it — one of the operands needs to be transposed or the order needs to change.
- Why does multiplying a 3 x 2 matrix by another 3 x 2 matrix fail, when adding them is fine?Addition is entrywise and needs identical shapes. Multiplication needs the columns of the left operand to match the rows of the right: here that is 2 against 3, so the row-by-column dot products have no common length. Transposing one operand fixes it — `(3 x 2)(2 x 3)` gives `3 x 3`, and `(2 x 3)(3 x 2)` gives `2 x 2`.
- If A is 3 x 2 and B is 2 x 3, are both AB and BA defined, and are they the same size?Both are defined, but they are different objects. `AB` is `(3 x 2)(2 x 3)` and comes out `3 x 3`; `BA` is `(2 x 3)(3 x 2)` and comes out `2 x 2`. They cannot even be compared entry by entry, which is the crudest demonstration that matrix multiplication does not commute.
- Roughly how many multiply-add operations does the standard algorithm need for an (m x n) by (n x p) product?About `m * n * p`. There are `m * p` output entries, and each is a dot product of length `n`, costing `n` multiplications and `n - 1` additions. So the shared inner dimension multiplies the cost even though it never appears in the output shape.
Think of each matrix as a pipe with an input width and an output width. You can only join two pipes when the first pipe's output width equals the second's input width; the joint itself is invisible from outside.
saying these in an interview costs you the question
- Claims both matrices must have the same shape, as in addition
- Multiplies corresponding entries instead of rows against columns
- Says the result takes the shape of the larger operand
- Keeps the shared inner dimension in the output shape
- Dots columns of the left matrix with rows of the right