In a CNN, what is the real cost of choosing valid padding over same zero padding?
answer
- one mode invents values, the other loses them
- shrinkage compounds with depth
- a border ring gets no output position
- constant zeros are a position cue
- same stops preserving size once stride exceeds one
basics
~20 sValid padding computes outputs only where the kernel fits inside real data, so the map shrinks at every layer and a border ring gets no output position. Same zero padding preserves size but feeds invented zeros into border windows.
solid answer
~50 sValid means no padding: the kernel is only placed where its whole window lies inside real data, so a 3x3 layer trims one pixel from each side and the loss compounds with depth. Stack ten such layers and the outer ten-pixel ring of the input has no corresponding output position anywhere. On a scanned document field that is exactly where the underline beneath a signature lives, so the model is blind to it. Same padding rings the input with zeros so the output keeps the input's spatial size, at the price of border windows that partly average over invented values - and those constant borders leak a usable cue about absolute position, which a model can learn to exploit in ways that do not transfer to differently framed inputs. Neither choice adds parameters. Pick valid when fabricated border values are unacceptable, same when you need the geometry to line up.
go deeper
Be ready to state the mechanical difference: valid means no padding and a smaller output, same means a zero border so a stride-1 layer keeps its input's height and width.
Explain both costs, not just the sizes. Valid removes any output position centred on the border ring; same makes border outputs partly a function of invented zeros.
Demonstrate the diagnosis: an accuracy band along image edges, or a model that breaks on tiles, points at padding. Know why tiled inference favours unpadded convolutions with overlap.
Own the framing contract between training and deployment. If inputs will be tiled, cropped or resized in production, decide early whether the network is allowed to learn anything from its border.
## The two modes **Valid padding** means no padding at all. The kernel is placed only where its entire window sits inside real input, so the output is smaller than the input: a 3x3 kernel at stride 1 loses one pixel from each of the four sides, a 5x5 kernel loses two. **Same zero padding** rings the input with a border of zeros - one ring for 3x3, two for 5x5 - so that a stride-1 layer emits a map with the same height and width as its input. "Same" is a statement about stride-1 behaviour; with a stride above 1 the output shrinks by the stride regardless of padding, which is a routine misunderstanding. Neither mode introduces learnable parameters. Padding is data manipulation, not capacity. ## What valid padding costs: a shrinking, blind border The shrinkage compounds. Ten stacked valid 3x3 layers remove ten pixels from every side. Beyond the bookkeeping annoyance, the substantive cost is that the outer ring of the input has no output position that is centred on it. Content there still influences some outputs - it appears in windows centred one pixel inward - but it can never be the subject of a prediction. On a scanned document field this is the failure that bites: the signature's underline sits in the bottom few pixel rows of the crop, and after a valid stack there is simply no output unit whose job is that region. The model behaves as though the field ends slightly higher than it does. Anything where evidence hugs the frame - marks at the edge of a form, structures at the boundary of a tile - is exposed to the same problem. The shrinkage also makes architecture with skip connections awkward: two paths that started the same size no longer match, so a valid-padded design has to crop one side to align them. ## What same padding costs: fabricated context Same padding buys size preservation by inventing values. A window centred on a corner pixel of a 3x3 layer contains five real values and four zeros. Its output is therefore a blend of measurement and fiction, and border statistics differ systematically from interior statistics. There is a subtler consequence. The zero border is a constant, reliable signal that appears only near the edges, and a convolutional stack can learn to use it to infer *absolute* position - how far a unit is from the frame - even though convolution is nominally position-agnostic. Sometimes that is useful. Often it is an invisible shortcut: a model trained on centred, uniformly framed crops can lean on the border cue and degrade when inputs are tiled, cropped differently, or presented at another size. Zeros are also not the only option. Reflecting or replicating the edge values keeps the border statistics closer to the real distribution and avoids injecting a hard discontinuity, at the cost of duplicating information the model may then over-count. ## Choosing Use **same** padding when geometry must be preserved: dense prediction where an output pixel should correspond to an input pixel, architectures with additive skip connections, or deep stacks where valid shrinkage would eat the image. Use **valid** padding when fabricated values are genuinely unacceptable. The classic setting is tiled inference over an image far too large to process at once: you run the network on overlapping tiles and stitch the results, and you want every output to depend only on real data so that seams do not show. A design in that spirit uses unpadded convolutions throughout and supplies the missing context by overlapping the tiles rather than by inventing it - and at the outer boundary of the whole image, where no real context exists, it mirrors the image rather than zeroing it. ## Interview framing The weak answer is "same keeps the size, valid shrinks it" and stops there. That is the definition, not the cost. The strong answer names both prices: valid gives up a border region that ends up unrepresented in the output, and same gives up border authenticity and hands the model a positional shortcut it did not ask for. The strong answer also flags the trap that same padding does not preserve size once the stride is above 1. ## Diagnostic tells If a segmentation model is systematically worse in a band along the image edge, look at padding first. If a model trained on full frames collapses on cropped or tiled inputs, suspect it learned something from the constant border. And if two branches of a network refuse to add, count the valid convolutions on each path.
- Why can zero padding leak absolute position information to a supposedly position-agnostic network?The zero ring is a constant that occurs only near the frame, so a unit's window composition tells it how close to the edge it is. Stacked layers can compose that into a fairly precise distance-from-border signal. It helps when framing is consistent and hurts when the model later sees crops or tiles framed differently.
- When would you deliberately choose valid padding despite the shrinkage?When outputs must depend only on real measurements - most typically tiled inference over an image too large to process whole. Unpadded convolutions guarantee no output was computed from invented values, and the missing context comes from overlapping neighbouring tiles instead. At the true image boundary you supply context by mirroring rather than zeroing.
- Does same padding still preserve the output size when the stride is 2?No. Same padding is defined by stride-1 behaviour. With stride 2 the output is roughly half the input in each dimension whatever the padding; padding only controls how the borders are handled and whether the halving rounds up or down. Expecting size preservation at stride 2 is a common source of shape surprises.
saying these in an interview costs you the question
- Thinks same padding gives border windows real data
- Assumes same padding preserves size at any stride
- Says padding adds learnable parameters
- Treats border loss in a deep valid stack as negligible
- Cannot name any cost of zero padding at all