A 300x300 image through five stride-2 stages breaks a decoder's skip concatenation — why?
answer
- odd sizes do not halve cleanly
- the floor discards half a row
- concatenation needs identical height and width
- five halvings need a factor of 32
- pad to a multiple, crop at the end
basics
~20 s300 is not divisible by 32. The floor in the output-shape formula drops a row at every odd-sized stage, so the encoder produces 37 where the decoder's doubling produces 36, and concatenation needs identical height and width.
solid answer
~50 sWalk the sizes: 300 -> 150 -> 75 -> 37 -> 18 -> 9. The third stage receives an odd 75, and `floor((75 - 2)/2) + 1 = 37`, not 37.5 — half a row is discarded and never recorded. The decoder doubles from 9 as 18 -> 36 -> 72, so at the stage-three junction it holds 36 against a 37-wide skip, and concatenation needs the spatial axes to match exactly. The root cause is that five exact halvings need an input divisible by `2^5 = 32`, and 300 is not. The clean fixes are to constrain inputs to multiples of 32, or to pad the input up to 320 and crop the output back. Interpolating the skip to fit compiles but shifts features by a fraction of a pixel, which hurts dense predictions near boundaries.
code
python · 21 linesimport math
def out_size(n, k, s, p=0, d=1):
k_eff = k + (k - 1) * (d - 1)
return math.floor((n + 2 * p - k_eff) / s) + 1
n = 300
enc = [n]
for _ in range(5): # five 2x2 stride-2 downsamples
n = out_size(n, k=2, s=2)
enc.append(n)
print(enc) # [300, 150, 75, 37, 18, 9]
n = enc[-1]
dec = [n]
for _ in range(5): # the decoder doubles each stage
n *= 2
dec.append(n)
print(dec) # [9, 18, 36, 72, 144, 288]
print(enc[3], dec[2]) # 37 vs 36 -> the stage-3 skip cannot concatenatego deeper
Remember that concatenating two feature maps requires identical height and width, and that the error message names the two sizes that disagree — read them before guessing.
Walk the encoder sizes stage by stage with the floor formula and show exactly where an odd size loses a row that the decoder's doubling cannot restore.
Show the production instinct: reproduce with the failing resolution, fix by constraining or padding the input, and add a shape assertion that runs across several sizes.
Decide and document the input-size contract for the serving stack, so arbitrary user image sizes never reach a model that silently assumes divisibility by 2^depth.
## What actually happened Encoder-decoder architectures with skip connections — U-Net being the canonical example — pair each encoder stage with a decoder stage at the same resolution and concatenate the two tensors along the channel axis. Concatenation along channels requires the **spatial axes to agree exactly**. Every other axis may differ; height and width may not. Run the shape formula down the encoder. With a 2x2 stride-2 downsample and no padding, `out = floor((n - 2)/2) + 1`: ``` 300 -> 150 -> 75 -> 37 -> 18 -> 9 ``` The interesting stage is the third. It receives 75, an odd number. The ideal halving is 37.5, and there is no such thing as half a pixel, so the floor keeps 37 and the final row and column are simply not covered by any window. The half-row is discarded and nothing records that it existed. The decoder does not know that. Starting from 9 and doubling: ``` 9 -> 18 -> 36 -> 72 -> 144 -> 288 ``` At the junction with the stage-three skip it presents 36 against a 37-wide encoder tensor, and the concatenation fails. The same one-pixel debt propagates upward: the decoder ends at 288 rather than 300. ## Why 300 and not 256 Five exact halvings require the input to be divisible by `2^5 = 32`. 256, 288, 320 and 512 all pass; 300 factors as `4 * 75`, so it survives two halvings and rounds at the third. This is exactly why the bug is so often reported as "it works on our test images and fails in production": development sets are frequently resized to a power of two, and the first arbitrary user image with a size like 300 or 1023 trips it. Note also that the two axes are independent. A 300x320 input rounds on height and not on width, producing an error that names two sizes differing on only one axis — a good diagnostic clue. ## Ranking the fixes **Constrain the input contract.** Require sizes divisible by `2^depth` and validate at the boundary. Simple, explicit, and it makes the failure loud instead of latent. The cost is that callers must comply. **Pad up, crop back.** Pad the input from 300 to 320, run the network, then crop the output back to 300x300. Every halving is exact, the alignment is preserved because padding is applied on known sides, and callers see the size they asked for. This is the standard production answer for arbitrary input sizes. **Pass explicit target sizes to the upsampling stages.** Instead of asking the decoder to double, tell each stage the spatial size of the encoder tensor it must match. This is exact and handles odd sizes correctly, but it makes the decoder depend on the encoder's recorded sizes. **Crop the skip.** The original U-Net used unpadded convolutions and cropped the larger encoder tensor to fit the decoder tensor. It is correct and well established, but it discards a border and shifts the effective field of view, so it must be a deliberate design choice rather than a patch. **Interpolate the skip to fit.** This is the fix most likely to be reached for and the one to argue against. It compiles, it trains, and it silently misaligns features by a fraction of a pixel at every stage where it fires. For dense prediction — segmentation masks, depth, keypoints — that misalignment shows up as soft, wandering boundaries that are very hard to trace back to a resize buried in the decoder. **Round the pooling up instead of down.** Tempting, but it relocates the mismatch rather than removing it: rounding up gives an encoder chain of 300 -> 150 -> 75 -> 38 -> 19 -> 10, and a decoder doubling from 10 gives 20 against a 19-wide skip. The parity problem is structural, not a rounding-direction preference. ## The habits that prevent it Trace the size chain by hand, or in a few lines of arithmetic, before training anything. Assert on the encoder and decoder sizes at every junction rather than discovering the mismatch as a runtime error deep in a training run. Test at several input resolutions — including a deliberately awkward one such as 300 or 401 — not only at the resized training resolution. And write down the divisibility requirement as part of the model's public contract, because it is a real constraint of the architecture and not an implementation detail. One more subtlety worth knowing: even when sizes happen to match, a chain that floors is still throwing away edge pixels at each rounding stage. That is usually harmless for classification but matters when the task is to predict something at the border of the image.
- Why does the same model run fine at 320x320?320 is 32 * 10, so all five halvings are exact: 320 -> 160 -> 80 -> 40 -> 20 -> 10, and the decoder's doubling retraces those sizes precisely. No stage ever floors away a row, so every skip junction matches. The requirement is divisibility by 2^depth, not size or squareness.
- Why is interpolating the skip tensor to match a poor fix?It removes the error message without removing the misalignment. Resampling by one pixel out of 37 shifts features by a fraction of a pixel at that stage and every stage above it, which softens and displaces boundaries in dense predictions. It also hides the real constraint, so the next architecture change reintroduces the bug silently.
- How would you catch this before a training run rather than during one?Compute the encoder and decoder size chains up front and assert equality at each skip junction, then run that check across a set of input sizes including deliberately awkward ones such as 300 and 401. A few lines of arithmetic replace a crash that otherwise surfaces minutes into training.
- Does the rounding also matter when the sizes happen to match?Yes, though more mildly. Whenever the numerator is not divisible by the stride, the final row and column are not covered by any window, so the border is quietly dropped. Classification rarely notices; segmentation, depth and keypoint tasks that must predict at the image edge can.
Integer division loses a remainder at every stage, and the decoder's doubling has no record of the remainder to put back.
saying these in an interview costs you the question
- Blames the data loader rather than the shape arithmetic
- Claims a stride-2 stage always halves the size exactly
- Interpolates the skip tensor and calls the bug fixed
- Only ever tests power-of-two input sizes
- Removes the failing skip connection to make the error go away
- Thinks rounding the pooling up removes the mismatch