skip to content

What does raising a convolution's stride from 1 to 2 change, and what does it cost?

level: middleimportance: must knowfreq 64%

answer

  1. it is a step size, not a size
  2. positions dropped, weights untouched
  3. roughly four times cheaper in two dimensions
  4. resolution lost cannot be recovered later
  5. stride above kernel size skips pixels

basics

~20 s

Stride is the step between successive kernel placements. Stride 2 evaluates the filter at every other position, so each spatial dimension comes out roughly halved and compute drops about fourfold. Parameters are unchanged; spatial precision is lost.

solid answer

~50 s

Stride sets how far the kernel moves between evaluations, so stride 2 keeps only every second position in each dimension. The output map is about half as tall and half as wide, which cuts this layer's multiply-adds and the activation memory of everything downstream by roughly four. The weights do not change at all - the same filter is simply applied at fewer places. What you pay is resolution: the layer only ever measures its filter's response on a coarse lattice, so a pattern that sits between two sampled positions is never seen centred, and small shifts of the input can flip the response. On 512x512 chest radiographs a stride-2 stem is the standard cheap first cut, because the earliest maps are the largest and the saving compounds through the whole network. If the stride ever exceeds the kernel size, some input pixels are read by no window at all.

go deeper

for a junior

Be ready to define stride as the step between kernel placements and to say that stride 2 roughly halves each spatial dimension of the output while leaving the weights alone.

for a middle

Explain the arithmetic of the saving: about four times fewer positions, so four times fewer multiply-adds here and a smaller map for everything after. Name what is paid for it - spatial precision.

for a senior

Show diagnostic instinct. Count strided stages when small objects disappear from predictions, and recognise shift-sensitivity as an aliasing symptom of coarse sampling with little window overlap.

for a principal

Own the placement policy across a backbone: how many reductions, at which depths, and what final spatial resolution downstream tasks need. Argue the cost curve, not a convention inherited from a classifier.

## What stride is A convolution slides a kernel over its input. **Stride** is the step size of that slide: with stride 1 the kernel is evaluated at every position; with stride 2 it is evaluated at every second position along each spatial axis; with stride 3, every third. Stride is purely about *where* the filter is placed. It does not change the filter, its size, its depth or its weight count. Because positions are dropped in both height and width, stride 2 produces an output roughly half as tall and half as wide - about a quarter of the positions. ## What you gain **Compute.** The cost of a convolution layer is proportional to the number of output positions it must produce, multiplied by the work per position. Stride 2 divides the position count by about four, so the layer itself is about four times cheaper - and, more importantly, every layer after it now operates on a map a quarter of the size. **Activation memory.** Training stores activations for the backward pass. Cutting the spatial size early is the single most effective way to reduce that footprint, because early feature maps are the biggest thing in the network even though they carry few channels. **Reach per layer.** With a coarser grid, each subsequent layer's window covers more of the original image, so a network reaches a wide view with fewer layers. This is why a stride-2 stem on 512x512 chest radiographs is the near-universal first move: at that point the map is enormous and the channel count is small, so the cut is nearly free in representational terms and enormous in cost terms. ## What you pay **Spatial precision, permanently.** Once positions are dropped, no later layer can recover them. For a classification head that is usually fine. For anything that must localize - segmentation masks, keypoints, small lesions a few pixels across - each stride-2 stage halves how finely you can point at something. **Sensitivity to sub-stride shifts.** Consider a 7x7 kernel at stride 3. Windows still overlap, so every input pixel is read by some window; nothing is skipped outright. But the filter's response is only ever *measured* at every third location. Shift the input by one pixel and no unit sees the pattern in the same alignment it saw before, so the layer's output can change more than the tiny shift warrants. This is aliasing: the response map is being sampled below the rate at which it varies. **Pixels that are never read.** The overlap guarantee fails when the stride exceeds the kernel size. A 2x2 kernel at stride 3 leaves a column and a row of input untouched by every window - those values contribute nothing to the layer's output. That is almost always a bug rather than a design, and it is worth checking whenever both numbers are unusual. ## What stride does *not* change - **Parameter count.** Same kernel, same weights, applied at fewer places. A candidate who says stride 2 halves the parameters has confused compute with capacity. - **Channel count.** The number of filters sets output channels; stride is orthogonal to it. In practice designers often double the channel count at a stride-2 stage, but that is a separate choice made to keep the layer's representational budget roughly flat, not something the stride does by itself. - **The kernel's window size.** Each individual placement still covers the same k x k patch. ## Choosing a stride in practice Stride 1 and stride 2 cover almost every real design. Stride 2 appears at stage boundaries where you deliberately trade resolution for cost. Larger strides show up mainly in stems, where a big kernel at stride 2 or higher digests a very large input in one step. The engineering question is always *where* the reductions sit, not how many exist. Placing them early is cheapest; placing them late means you paid full price for high-resolution computation you then threw away. Placing too many means the final map is so coarse that fine structure has no representation left - the classic failure when a classification-derived backbone is dropped into a dense-prediction task without adjustment. ## Debugging tells If small objects vanish from a model's predictions while large ones are fine, count the strided stages between input and the layer where prediction happens: a lesion eight pixels wide has essentially nothing left after three of them. If accuracy wobbles under one-pixel translations of the input, aggressive striding with little overlap is a prime suspect.

  • Does stride 2 reduce the layer's parameter count?
    No. The same kernel is simply evaluated at fewer positions, so the weight count is identical. What drops is the number of multiply-adds and the size of the output map, which then reduces activation memory and compute for every later layer. Confusing the two is the most common error on this question.
  • Why is a stride-2 layer usually placed at the very start of a network rather than deep inside it?
    Cost per layer scales with spatial size times input channels times output channels. Early on the spatial map is huge and the channel counts are tiny, so halving each dimension there is cheap in representation and saves the most work - and the saving compounds, because every subsequent layer inherits the smaller map.
  • When can a strided convolution leave input pixels that no window ever reads?
    When the stride is larger than the kernel's spatial size. A 2x2 kernel at stride 3 leaves a gap between consecutive windows, so those pixels contribute to no output at all. As long as the stride is at most the kernel size, consecutive windows touch or overlap and every pixel is covered.
  • Why might a heavily strided network's predictions change under a one-pixel shift of the input?
    Striding samples the filter's response map coarsely. A pattern that lands between two sampled positions is never measured in its best alignment, so a shift that moves it onto a sampled position can change the response sharply. This aliasing is worse when overlap between successive windows is small.

saying these in an interview costs you the question

  • Says stride 2 halves the number of learnable parameters
  • Treats stride purely as a speed knob with no representational cost
  • Confuses stride with kernel size or dilation rate
  • Claims later layers can recover the dropped positions
  • Thinks stride changes the number of output channels

context