Why does a convolutional layer share one kernel across all positions instead of using per-position weights?
answer
- one detector, reused everywhere
- two restrictions, not one
- local window plus identical weights
- 27 numbers against 150,528
basics
~20 sA convolution reuses one small kernel at every position, so it learns a feature detector once instead of relearning it at each pixel. That cuts the parameter count enormously and lets the same feature be found anywhere in the input.
solid answer
~50 sA dense layer gives every input element its own weight into every unit. On a 224x224x3 image that is 150,528 inputs, so a single dense unit already carries 150,528 weights, and the edge detector it learns for the top-left corner is useless in the bottom-right. A convolutional layer imposes two separate restrictions instead. First, **sparse local connectivity**: each output reads only a small window across the input channels, everything else is structurally zero. Second, **weight sharing**: every position in one feature map uses the same window weights, so a 3x3 kernel over 3 input channels is 27 weights plus a bias, reused everywhere. The payoff is not only memory. Sharing encodes a prior — image statistics are roughly the same everywhere — which shrinks the hypothesis space, improves sample efficiency, and makes the layer's parameter count independent of input height and width.
go deeper
Be ready to state the two structural differences from a dense layer in one breath: each output sees only a small local window, and the same window weights are reused at every position. Have the rough scale of the parameter saving on hand.
Explain that sharing makes the parameter count independent of input height and width, that each output channel owns its own kernel, and that sharing is spatial only rather than across channels. Name locally connected layers as the unshared middle ground.
Show that you treat sharing as an assumption about the data, not a default. An interviewer wants to hear when the stationarity it asserts holds for your inputs, and what you would measure or inspect before trusting it on an unfamiliar modality.
Own the framing that architecture choice is prior selection under a data budget. Be able to argue when a strong built-in bias is the right trade against a weaker one that needs far more data, and how you would decide that for a team rather than per model.
## What a dense layer assumes A fully connected layer treats its input as an unordered vector: every input element gets its own private weight into every output unit. Feed it a 224x224x3 colour image and the input is 150,528 numbers, so one output unit holds 150,528 weights plus a bias. Nothing inside that layer knows that one pixel sits next to another, or that the same vertical edge could appear in the top-left corner and again in the bottom-right. If the layer learns an edge detector at one location, it must learn a completely separate copy for every other location, from separate evidence. ## The two restrictions a convolution adds A convolutional layer replaces the dense connection pattern with two restrictions that are logically independent of each other. **1. Sparse local connectivity.** Each output unit reads only a small window of the input — a k-by-k spatial patch spanning all input channels — and the weight on everything outside that window is exactly zero by construction, not because training drove it there. This encodes the belief that whatever matters for a low-level feature is nearby: edges, corners, colour transitions and textures are local phenomena. **2. Weight sharing.** Every output position within the same feature map uses the *same* window weights. The layer stores one kernel and slides it. This encodes a second, stronger belief: the statistics of the input are approximately the same everywhere along the shared axes, so an edge is an edge wherever it appears and there is no reason to learn a location-specific detector. Concretely, a 3x3 kernel over 3 input channels holds 27 weights and one bias, and those same 27 numbers produce every position of the output map. Set that against the 150,528 weights one dense unit needed for the same image and the scale of the difference is obvious. ## What sharing actually buys you - **Parameters and memory.** Fewer weights to store, transmit and regularise. - **Statistical efficiency.** The constraint is a prior. It removes from consideration every function that would treat the top-left of an image differently from the bottom-right, so the network has far fewer ways to fit noise, which matters most when labelled data is scarce. - **Detect anywhere.** A feature learned from examples where the object was centred still fires when the object is at the edge, because the same weights are applied there too. - **Input-size flexibility.** Because the parameter count does not depend on spatial height and width, the same layer accepts a larger or smaller image without any change to its weights. ## What you pay Sharing is an assumption, and assumptions can be wrong. It asserts that a pattern means the same thing at every position along the shared axes. When the position itself carries meaning, that assertion is a false constraint that the network cannot train its way out of, because the parameters simply do not exist to express a position-dependent response. That is a judgment call about the data, not about the architecture. ## Two things sharing is *not* A very common muddle is to think a convolution shares weights across everything. It does not. Sharing is **spatial only**: - Each input channel gets its own slice of the kernel — a 3x3 kernel over 3 input channels really has 3 distinct 3x3 grids of weights, not one grid applied three times. - Each output channel has its own separate kernel; different output channels are different learned detectors. - Sharing across *positions within one layer* is also distinct from tying weights between different layers at different depths, which is a separate technique with different motivations. ## The middle ground You can keep locality and drop sharing. A **locally connected** layer wires each output to a small local window, like a convolution, but gives every position its own weights. It saves parameters relative to a dense layer while allowing position-specific detectors, which can pay off when inputs are registered so that position is meaningful and consistent — aligned faces or registered medical scans, for instance. It costs you the ability to recognise the feature at a position never seen during training, and the parameter count grows with input size again. ## How to answer this in an interview Say the parameter saving, but do not stop there — every candidate says the parameter saving. The answer that lands separates locality from sharing, names them as two distinct structural priors, and says explicitly that sharing is an inductive bias about the data rather than a compression trick.
- Locality and weight sharing are separate ideas. What does a locally connected layer without sharing give you?It keeps the small local window but gives every position its own weights. You still avoid the full dense parameter count, and you gain the ability to learn position-specific detectors, which helps on registered inputs where a given location always means the same thing. You lose translation equivariance and the ability to recognise a pattern at a position that was never seen during training, and the parameter count grows with input size again.
- Beyond the parameter count, why does sharing help most when labelled data is scarce?Because it is a prior, not just a compression. Sharing removes from the hypothesis space every function that would respond differently at different positions, so there are far fewer ways for the model to fit noise. Each kernel is also supported by evidence from every position of every image rather than from one location alone, so the same amount of data pins down far fewer numbers much more tightly.
- Does a convolutional layer share weights across input channels as well as across positions?No. Sharing is spatial only. A 3x3 kernel over 3 input channels holds three distinct 3x3 grids, one per input channel, which are summed after applying them. Each output channel then owns its own complete kernel and bias. So the layer has many independent detectors; what is reused is each detector across space, not across channels.
Rather than hiring one inspector per spot on a conveyor belt, each trained only on their own spot from their own examples, you train a single inspector and walk them along the whole belt.
saying these in an interview costs you the question
- Calls weight sharing purely a memory optimisation with no effect on generalisation
- Thinks each output position has its own separately learned kernel
- Claims a convolution also shares one weight grid across all input channels
- Confuses spatial sharing with tying weights between different layers at different depths
- Cannot separate sparse local connectivity from weight sharing as two distinct constraints