In semantic segmentation, how does a CNN classifier become a dense per-pixel predictor?
answer
- the head is the obstacle
- flattening throws away where
- dense layer equals a convolution
- C score maps, softmax per pixel
- coarse grid, then a decoder
basics
~20 sDrop the flatten-and-dense head and make every layer convolutional. The dense head is rewritten as an equivalent convolution, the last layer emits one score map per class, and a decoder upsamples those maps back to the input resolution.
solid answer
~50 sA classifier ends by flattening its final feature map into a vector and running dense layers, which discards spatial layout and pins the network to one input size. A dense layer over a `7x7x512` feature map is arithmetically identical to a `7x7` convolution with one output channel per unit, and any dense layers after it are `1x1` convolutions, so the head can be rewritten as convolutions without changing what it computes. The network is then fully convolutional: it accepts any image size and returns a coarse grid of class scores instead of a single vector, so an arbitrarily large scene is labelled densely in one forward pass. The final layer has `C` channels, one per class; a softmax runs across channels at each spatial location and training uses cross-entropy averaged over pixels. Because the encoder has downsampled, often 32x, a decoder upsamples the score grid back to `HxW`.
go deeper
Be ready to state the output shape without hesitating: one score map per class, at input resolution, with a softmax over classes at each pixel and cross-entropy averaged over pixels.
Explain the convolutionalisation arithmetic out loud: which kernel size replaces the dense layer over a given feature map, why the layers after it become 1x1 convolutions, and what the coarse output grid size is for a stride-32 encoder.
Show you have run this in production: how you pick crop and inference sizes, how you handle ignore labels in the loss, and why sliding-window classification is not a serious alternative on real image sizes.
Own the framing decision. Argue when per-pixel labelling is the right product output at all versus boxes or region summaries, and what dense output costs in annotation budget, storage and inference time across a fleet.
## The task Semantic segmentation assigns a class label to **every pixel** of an image. For an `H x W` input and `C` classes, the model produces `H x W x C` scores, and the prediction at each pixel is the argmax over the `C` channels. Note what this does *not* do: it labels pixels, not objects. Two adjacent cells of the same class fuse into one connected region, because nothing in the output distinguishes one instance from another. ## Why a classifier's head is the obstacle A typical image classifier is an encoder of convolution and downsampling stages followed by a head: flatten the final feature map into a long vector, then one or more dense (fully connected) layers, then a softmax over `C` classes. Two things break dense prediction here. 1. **Flattening destroys spatial layout.** Once the `7x7x512` map is a 25088-long vector, the network can no longer say *where* anything is; it only says *what* the whole image contains. 2. **A dense layer fixes the input size.** Its weight matrix has a row per input element, so the feature map must always have exactly `7x7x512` elements, which forces one fixed input resolution. ## The convolutionalisation trick A dense layer applied to a `7x7x512` feature map computes, for each of its `N` units, a weighted sum over all `7*7*512` values. That is exactly what a convolution with a `7x7x512` kernel and `N` output channels computes at a single spatial position. So you can **reshape the dense weights into convolution kernels** and get a layer that computes the identical function. Any dense layer stacked after it consumes a `1x1xN` map and is therefore a `1x1` convolution with the same weights. The result is a network with no dense layers at all: a **fully convolutional network**. Nothing about its behaviour on a `224x224` input has changed. What has changed is what happens on a larger input. A convolution slides; it does not care how big the plane is. Feed a `500x700` image and the head, now convolutional, evaluates at every position of the final feature grid, returning a small `H/32 x W/32 x C` map of class scores instead of one vector. Each score in that map is what the original classifier would have said about the receptive field centred there. You get dense scores over a whole large scene from a **single forward pass**, instead of cropping thousands of overlapping windows and classifying each one. ## Output shape and loss The last convolution has `C` output channels, one logit map per class. The softmax is taken **across the channel axis at each spatial location independently**, so every pixel gets its own probability distribution over classes. The loss is cross-entropy computed per pixel and then averaged (or summed) over all labelled pixels of the image. Pixels marked as ignore or unlabelled in the ground truth are excluded from the average rather than assigned a class. This is the main mental shift from classification: one image is not one training example with one label, it is tens of thousands of per-pixel examples that happen to share a forward pass. ## Getting back to full resolution The coarse score grid is the remaining problem. If the encoder downsamples by 32, the prediction grid is 32 times smaller than the image in each dimension, so upsampling it directly gives blocky masks whose boundaries are accurate only to about half a stride. A **decoder** fixes this: it repeatedly upsamples, either by learned transposed convolution or by fixed interpolation followed by a convolution, until the output is back at `H x W`. The overall shape is an encoder-decoder, wide and shallow at the ends, narrow and deep in the middle. A decoder that only upsamples still produces soft, rounded boundaries, because the detail was destroyed on the way down and interpolation cannot invent it. That is what skip connections from the encoder into the decoder are for, and it is why the plain encoder-decoder is a starting point rather than a finished design. ## Common mistakes - Keeping global average pooling before the head. It collapses the spatial axes, which is exactly what you must not do. - Applying softmax over the spatial axes instead of the channel axis. Class probabilities must sum to one per pixel, not per class map. - Training at one crop size and assuming the network is therefore locked to it. A fully convolutional network runs at any size that survives the downsampling arithmetic; only the statistics of the receptive field context change.
- Why can a fully convolutional network accept an input size it never saw in training?Every layer is a sliding local operation whose parameters are tied across positions, so none of them has a weight per input coordinate. A larger image simply yields a larger feature grid. Only the arithmetic constraint matters: the size must survive the downsampling and upsampling factors cleanly, or shapes stop matching.
- How does the loss treat a class that covers most of the image?Plain per-pixel cross-entropy averaged over pixels weights every pixel equally, so a class covering 40 percent of the pixels contributes 40 percent of the gradient. Rare thin classes contribute almost nothing, which is why segmentation training often reweights pixels by class or samples crops that contain the rare classes.
- What is the cheapest way to double the output resolution of a stride-32 network?Reduce the stride of the last downsampling stage so the encoder ends at stride 16 instead of 32. It halves the amount the decoder must recover, but the final stages then run on four times as many spatial positions, so memory and compute rise sharply, and each output position sees less surrounding context.
A classifier is a stamp that prints one verdict for whatever it is pressed onto. Convolutionalising it turns the stamp into a roller: the same verdict machinery, rolled across the whole image, leaving a verdict at every position.
saying these in an interview costs you the question
- Says segmentation outputs one label per image region
- Thinks a dense layer cannot be written as a convolution
- Applies softmax across pixels instead of across classes
- Believes upsampling alone recovers destroyed boundary detail
- Confuses per-pixel class labels with per-object instances