skip to content

In a 12288-512-128-10 MLP fed a batch of 32 flattened aerial tiles, what shape is each activation?

level: juniorimportance: must knowfreq 78%

answer

  1. one axis never changes
  2. rows are examples, columns are features
  3. (batch, in) times (in, out)
  4. the product consumes only the feature axis

basics

~10 s

Input (32, 12288), then (32, 512), then (32, 128), and finally (32, 10). The batch size stays on the leading axis at every layer; only the trailing feature width changes, once per weight matrix.

solid answer

~40 s

Each 64x64x3 tile flattens to 64*64*3 = 12288 numbers, so the batch enters as (32, 12288). With weights stored as (inputs, outputs), the first layer computes `X W1 + b1`: (32, 12288) x (12288, 512) gives (32, 512), and the (512,) bias is added to every row. The nonlinearity is elementwise and changes nothing. The next two products give (32, 512) x (512, 128) -> (32, 128) and (32, 128) x (128, 10) -> (32, 10), one score per class per example. The batch axis is never consumed by a product, which is why the same weights accept a batch of 1 or 512. A transpose is needed only if a weight is stored as (outputs, inputs), in which case you compute `X W^T`.

go deeper

for a junior

Be ready to state every tensor shape in a small MLP given only the layer widths and the batch size, and to say which axis holds the examples.

for a middle

Explain why a product consumes only the feature axis, when a stored weight forces a transpose, and why elementwise activations leave the shape untouched.

for a senior

Show how you make shapes checkable in real code: named or asserted shapes at each module boundary, so a mis-sized layer fails where the mistake is rather than at the head.

for a principal

Own the convention itself. Pick and document one house layout - examples on the leading axis, features trailing - so models, loaders, serving and export code agree without per-project renegotiation.

## The layout convention comes first A fully-connected layer is a matrix product plus a bias. Before you can name a single shape you have to fix how a batch is laid out, and the near-universal convention in deep learning is **rows are examples, columns are features**: `B` examples with `F` features each form a `(B, F)` array. Under that convention a layer with `F` inputs and `H` outputs stores a weight matrix of shape `(F, H)` and a bias of shape `(H,)`, and computes `Y = X W + b`. The shape rule follows: `(B, F) x (F, H) -> (B, H)`. The two inner `F`s meet and cancel; the outer dimensions survive. `B` is never touched. That single sentence is the whole leaf. ## Walking the stack An aerial survey tile of 64 x 64 pixels with 3 colour channels flattens to 64 * 64 * 3 = 12288 numbers. Thirty-two of them stacked give the input batch. Layer by layer: - input: `(32, 12288)` - layer 1: weight `(12288, 512)`, bias `(512,)` -> pre-activation `(32, 512)` - nonlinearity: elementwise -> still `(32, 512)` - layer 2: weight `(512, 128)`, bias `(128,)` -> `(32, 128)`, nonlinearity -> `(32, 128)` - head: weight `(128, 10)`, bias `(10,)` -> `(32, 10)` The output `(32, 10)` is ten class scores for each of the thirty-two examples. Every hidden width appears exactly twice in the parameter list: once as the output width of the layer that produces it and once as the input width of the layer that consumes it. Reading a stack that way is the fastest manual check there is - if a width appears with two different values on the two sides of a boundary, that boundary is the bug. ## Where transposes come from Textbook notation writes a layer as `y = W x + b`, with `x` a column vector of length `F` and `W` of shape `(H, F)`. Batched code puts examples in rows instead, because a row-major batch is what a data loader naturally produces and what keeps each example contiguous. The two forms differ by a transpose: with `W` stored as `(F, H)` you write `X W`; with `W` stored as `(H, F)` you write `X W^T`, which is the same numbers arranged the same way. Either storage order is fine. What breaks is mixing them without noticing - multiplying `(32, 12288)` by a `(512, 12288)` matrix is simply not defined, and multiplying `(12288, 32)` by `(12288, 512)` is not defined either, though for a different reason. ## Bias and elementwise functions The bias is a length-`H` vector, one learned offset per output unit, added identically to all `B` rows. Activation functions, dropout scaling and normalisation all act elementwise or per-feature; none of them change the rank or the extent of the tensor. So in a plain MLP the only places a shape changes at all are the weight products. ## Why the batch size is not a model hyperparameter Because the batch axis never enters a product, nothing in the parameter set depends on it. A layer trained with batches of 256 evaluates one example at a time without a single change, and the parameter count is identical either way. This is also why a shape written as `(None, 512)` or `(*, 512)` in a summary is not a mystery: the leading axis is free. ## Flatten order matters, but only for consistency Going from 64 x 64 x 3 to a flat 12288 requires choosing an order - channels fastest, or rows fastest. Any fixed order works, because the first layer simply learns one weight per input position and does not know what those positions mean. What is fatal is using one order at training time and a different one at inference: the learned weights then line up with the wrong pixels, no shape error is raised because the length is still 12288, and accuracy collapses to chance. ## Where a mismatch actually surfaces Shapes are checked pairwise, at each product, in order. If you widen the second hidden layer from 128 to 256 and forget the head, layer 2 still runs happily - `(32, 512) x (512, 256)` is fine - and the failure is raised at the head, `(32, 256) x (128, 10)`, one layer past the mistake. The error message names the innocent layer. The habit that fixes this is to write the intended shape at each boundary, or assert it, so the check happens where the intent lives rather than wherever two mismatched widths finally collide.

  • Why does none of the weight shapes mention the batch size of 32?
    Because the parameters map features to features, and the batch axis rides along untouched through every product. Each row is transformed independently, so the same weights serve a batch of 1 or 512. Batch size is a throughput and optimisation choice, not part of the model definition.
  • Where in this stack would a transpose actually be required?
    Only at a layer whose weight is stored as (outputs, inputs) rather than (inputs, outputs). Then the batched form is `X W^T`: (32, 12288) x (12288, 512). Transposing the input instead - `W X^T` - also computes valid numbers, but returns (512, 32), putting the batch on the trailing axis and breaking the convention for everything downstream.
  • Does it matter in what order a 64x64x3 tile is flattened into 12288 values?
    Not intrinsically - the first layer learns one weight per position and is indifferent to which position is which. It matters absolutely that the order is identical at training and inference. A changed order keeps the length at 12288, so no shape error fires, and the model silently reads scrambled pixels.

A weight matrix is a fixed adapter: it turns a row of 12288 numbers into a row of 512. The batch dimension is just how many rows you push through the adapter at once.

saying these in an interview costs you the question

  • Says the weight matrix has a batch dimension
  • Thinks the activation function changes the tensor shape
  • Multiplies (32, 12288) by a (512, 12288) matrix without transposing
  • Claims batch size is fixed when the model is defined
  • Assumes flatten order can differ between training and inference

context