What does AlexNet's 11x11 stride-4 first layer discard that a 3x3 stem keeps?
answer
- what layer one can never give back
- stride sets the sampling grid
- one wide linear filter, one nonlinearity
- cheap stem, coarse features
basics
~20 sFine spatial detail. A stride-4 first layer samples the image on a coarse grid, aliasing away structure finer than four pixels, and one wide linear filter plus one nonlinearity is a weaker map than a stride-1 3x3 stack.
solid answer
~50 sTwo things are lost, and they are separate. The stride: sampling every fourth position fixes the finest detail the network will ever localize, and no later layer can recover structure that was never sampled — high-frequency texture and small objects are the first casualties. The kernel: an 11x11 filter is a single linear map over a wide window followed by one nonlinearity, so whatever it detects must be expressible that way, whereas a stack of 3x3 stride-1 layers builds comparable reach through several nonlinear stages and keeps full resolution while doing it. What the aggressive stem buys is compute: cost scales with height times width, and the stem is where feature maps are largest, so a stride of 4 cuts the whole downstream network's work by about sixteen times. VGG spends that budget instead, and its early full-resolution layers hold almost no weights yet dominate its multiply-adds.
go deeper
Recall what stride 4 means concretely: the layer is applied at every fourth position, so its output map is a quarter of the input size in each dimension right after the first layer.
Explain that one 11x11 filter is a single linear map plus one nonlinearity over a wide window, while a stride-1 3x3 stack builds comparable reach through several nonlinear stages and keeps full resolution while doing it.
Show the trade in both directions: the aggressive stem is a large compute saving exactly where feature maps are biggest, and its price is spatial detail that nothing downstream can reconstruct. Name the tasks that cannot pay that price.
Own the stem as a budget decision. Decide how much resolution to spend before any features exist, tie it to the deployment latency target and to what the downstream task needs, and be able to defend a deliberately coarse stem.
## Two knobs, often conflated AlexNet's first layer uses 11x11 kernels at stride 4; VGG's first layer uses 3x3 kernels at stride 1. Candidates usually answer as if that is one decision, but kernel size and stride do different damage and buy different things, and a good answer separates them. ## What the stride costs Stride 4 means the layer evaluates its filter only at every fourth position, so the output map is a quarter of the input's size in each dimension right after layer one. The sampling grid is now coarse, and anything that varies faster than that grid can resolve is gone — folded, in the aliasing sense, into whatever the filter happened to average over. That loss is permanent: later layers see only the subsampled map, and no amount of depth reconstructs structure that was never measured. Tasks that depend on fine detail feel it first — small objects, thin structures, fine texture, precise localization. For coarse whole-image classification the loss is often affordable, which is precisely why the design survived as long as it did. ## What the kernel size costs An 11x11 filter over 3 input channels computes one weighted sum over a wide window and passes it through one nonlinearity. Whatever pattern it responds to must be expressible as a single linear template plus a pointwise nonlinearity. Stacked small kernels reach comparable window size through several linear-plus-nonlinear stages, which is a strictly richer composition per unit of reach. This is the same factorization argument that makes two 3x3 layers preferable to one 5x5 layer, applied at the stem. ## What the aggressive stem buys Compute. A convolution layer costs about `k*k*C_in*C_out*H*W` multiply-adds, and `H*W` is largest at the input. Striding by 4 at the stem shrinks `H*W` by roughly sixteen for that layer and for everything after it. The saving is enormous and it is why an early network could be trained at all on the hardware of its day. VGG's decision to convolve at full resolution for the first block is the opposite bet: those layers hold a trivial number of weights — a 3x3 layer at 64 to 64 channels is under forty thousand — and yet, running on a 224x224 map, they consume a large share of the network's arithmetic. ## Why the stem is not where the parameters are A common wrong answer blames the 11x11 kernels for AlexNet's size. Count it: 11 times 11 times 3 input channels times the filter count is a few tens of thousands of weights. Like VGG, AlexNet keeps the overwhelming majority of its parameters in the fully connected layers at the end. Large kernels are expensive in arithmetic at high resolution, not in storage at the stem. ## The judgment to demonstrate The stem is a budget decision, and framing it that way is what separates a senior answer from a recital. You are choosing how much resolution to spend before the network has learned anything, and the right answer depends on the downstream task and the latency budget rather than on a rule. A coarse, cheap stem is right when the task is whole-image and the compute budget is tight. It is the wrong call when the task needs fine localization, because the first layer sets a ceiling on precision that nothing later can lift. In between, the usual compromise is to reach the same output stride over several layers rather than in one jump, so that some nonlinear processing happens before the detail is thrown away. ## The trap in the factorization argument One last subtlety. Replacing an 11x11 stride-4 layer with three stride-1 3x3 layers does reduce weights, but it multiplies the arithmetic, because those layers now run at full resolution. The cheap-small-kernels argument compares kernels at equal spatial size; it says nothing about downsampling. If you want the stem's compute saving, you have to keep the downsampling — you can only spread it out.
- Why would anyone accept that loss and downsample so aggressively at the stem?Because the stem is where the feature maps are largest and convolution is most expensive — cost scales with height times width. Striding by 4 shrinks that by about sixteen for the first layer and everything after it, which is what made early large networks trainable at all. It is a compute-versus-detail trade, and for coarse whole-image classification the detail is often affordable to lose.
- Would replacing the 11x11 stride-4 stem with three 3x3 layers make the network cheaper?Not at stride 1. The factorization argument compares kernels at equal spatial resolution, whereas the stem's saving comes from the stride. Three stride-1 3x3 layers on a full-resolution image cost far more multiply-adds than one strided 11x11 layer, even though they hold fewer weights. To keep the saving you must keep the downsampling and merely spread it over more layers.
- Does the 11x11 stem explain AlexNet's large parameter count?No. That layer holds only a few tens of thousands of weights — kernel area times three input channels times the filter count. As in VGG, the parameters sit overwhelmingly in the fully connected layers at the end. Large kernels are costly in arithmetic at high resolution, not in storage.
- Which downstream tasks suffer most from an aggressive stem?Anything that needs precise spatial answers: detecting small objects, segmenting thin structures, reading fine texture, or localizing a boundary to within a few pixels. The stem sets a floor on the spatial precision the rest of the network can express, and depth cannot lift it. Coarse whole-image labels are the case that tolerates it best.
Sampling a signal at every fourth point: whatever wiggles between the samples is gone before any later stage gets a chance to look for it.
saying these in an interview costs you the question
- Says a larger kernel is simply better because it sees more
- Claims later layers can recover detail the stride discarded
- Blames the 11x11 kernels for the model's parameter count
- Treats kernel size and stride as the same design knob
- Assumes small kernels are always cheaper regardless of stride