When is a convolution's shared-weight assumption wrong for the data you are modelling?
answer
- ask what each axis means
- slide the pattern: does meaning survive?
- frequency is not like time
- absolute position can be the answer itself
basics
~10 sSharing asserts that the same local pattern means the same thing everywhere along an axis. It fails where position itself carries meaning: a spectrogram's frequency axis, absolute-coordinate targets, and tabular columns with no order.
solid answer
~50 sSharing weights along an axis claims the data is stationary along it — a pattern at one position means what it means at every other. Test that per axis. On a mel spectrogram it holds along time, since a sound is the same sound later, but not along frequency: moving a formant pattern up in frequency changes which vowel you hear, so a shared kernel encodes a false invariance. Fix it by convolving along time with kernels spanning the full frequency extent, or by feeding a frequency-index channel. Coordinate-dependent targets break it differently: to output where in the frame a gripper should close, a fully shared, spatially aggregated stack has no representation of absolute position, and the fix is two input channels holding each pixel's coordinates. Tabular columns break both assumptions — arbitrary order means no locality and no shift structure.
go deeper
Recall that convolution assumes neighbouring positions are related and that a pattern means the same thing wherever it appears. Be able to name one input type where that is plainly untrue, such as columns of a table.
Explain the assumption axis by axis and give a concrete counterexample, such as the frequency axis of a spectrogram, along with one mechanism that relaxes sharing there — an index channel, or sharing along time only.
Demonstrate diagnosis. Describe the symptom you would see when the prior is wrong, such as coordinate predictions collapsing toward the mean, and the change you would make, and say how you would confirm the fix rather than assume it.
Own the modality decision. Weigh a strong built-in prior that is slightly wrong against a weaker one that needs far more data, and set the expectation that a team justifies the symmetry it builds in before committing a modelling stack to it.
## State the assumption before testing it A convolution makes two claims about the data, and they can fail independently. - **Locality**: whatever matters for a low-level feature lives in a small neighbourhood of indices, so neighbouring positions are genuinely related. - **Stationarity along the shared axis**: the same local pattern carries the same meaning at every position along that axis, so one detector suffices. The useful habit is to take each axis of your input in turn and ask: if I slide a pattern along this axis, does its meaning survive? If yes, share. If no, sharing is a constraint the network cannot train away, because position-dependent parameters simply do not exist in that layer. ## Case 1: the frequency axis of a spectrogram A time-frequency representation looks like an image, which is exactly why people reach for a 2D kernel and share weights along both axes. Along **time** the assumption is sound: a burst of energy occurring later is the same burst. Along **frequency** it is not. In speech, the identity of a vowel is determined by where its formants sit on the frequency axis; move that same pattern of peaks up in frequency and you have a different vowel. Sharing weights along frequency asserts that a pattern centred low means the same thing when centred high, which is precisely the invariance you do not want. It also throws away a strong cue: harmonic structure, noise floor and channel effects are all frequency-dependent, and a shared kernel is forbidden from conditioning on which band it is looking at. Workable responses: - Convolve along **time only**, with kernels that span the full frequency extent, so the layer sees the whole spectral shape at once and shares only over the axis where sharing is justified. - Keep 2D kernels but add an input channel holding the **frequency index** of each row, so the layer can learn a band-dependent response despite shared weights. - Use **locally connected** or band-wise layers along frequency: local windows, but separate parameters per band. ## Case 2: targets that live in absolute coordinates Suppose a fixed overhead camera watches a workspace and the model must output where in the frame a robot gripper should close. The convolutional body is translation equivariant, which means it tells you *what* is where relative to the input, but a globally shared stack with an aggregating readout has genuinely no representation of absolute position — the same appearance at two places produces the same evidence, and position has already been discarded by the time the head runs. The symptom is a model that identifies the right object and then predicts coordinates poorly, often collapsing toward the mean location of the training targets. The standard fix is to make position an input feature: append two extra channels to the input, one holding each pixel's normalised x coordinate and one its y coordinate. The layer's weights are still shared, but the values it reads now differ by position, so a position-dependent response becomes representable. The same trick handles scene priors from a fixed camera — sky at the top, the belt in the middle — that a fully shared stack cannot express. The alternative is to keep a spatially structured output rather than aggregating: predict a heatmap over positions and read the coordinate off it. That keeps equivariance working for you instead of fighting it. ## Case 3: tabular columns A table of features breaks both assumptions at once. Column order is arbitrary — swap two columns and the data means exactly the same thing — so adjacency carries no information and locality is meaningless. There is also no shift structure: sliding a kernel from the age column to the postcode column and applying the same weights is not a symmetry of the problem, it is nonsense. A convolutional stack over such input buys nothing that a dense network does not, while adding an arbitrary constraint. The correct baseline is a dense network, or a model family suited to heterogeneous columns. If a candidate proposes a conv stack for a table because "it has fewer parameters", that is the parameter-count reflex overriding the question of whether the prior is true. ## The general diagnostic For each axis, ask three questions: 1. **Is adjacency meaningful?** Are indices i and i+1 genuinely related, or is the ordering arbitrary? 2. **Is the axis stationary?** Does sliding a pattern along it preserve the label semantics? 3. **Does the target depend on absolute position along it?** If so, either supply position as an input or keep a spatially structured output. An axis that passes 1 and 2 is a good axis to share over. An axis that passes 1 but fails 2 wants locality without sharing, or an explicit index channel. An axis that fails 1 should not be convolved at all. ## Interview framing What distinguishes a strong answer here is that the candidate treats convolution as a hypothesis about the data and says how they would check it, rather than as a default that is always safe on anything shaped like a grid.
- How would you keep convolution on a spectrogram without asserting invariance along frequency?Three options. Convolve along time only, with kernels spanning the full frequency extent, so you share only over the axis where sharing is justified. Or keep 2D kernels and add an input channel holding each row's frequency index, letting shared weights still condition on the band. Or use band-wise locally connected layers: local windows along frequency but separate parameters per band.
- A model must predict absolute pixel coordinates and keeps collapsing toward the mean location. What is going on?A translation-equivariant body followed by a readout that aggregates over space has no representation of absolute position, so identical appearance at two locations yields identical evidence and the best it can do is the average target. Either append coordinate channels to the input so position becomes a readable feature, or drop the aggregating readout and predict a spatial heatmap you read the coordinate off.
- How do you decide quickly whether convolution is worth trying on an unfamiliar modality?Check two things per axis. First, is adjacency real — would permuting indices along that axis destroy meaning? If permuting is harmless, there is no locality to exploit. Second, is the axis stationary — does sliding a pattern along it preserve what the label should be? Only an axis that passes both is a sound axis to share weights over.
Sharing weights along an axis is claiming the ruler's markings are interchangeable. Along a spectrogram's time axis they nearly are; along its frequency axis they are not.
saying these in an interview costs you the question
- Reaches for a conv stack on tabular columns to reduce parameters
- Treats a spectrogram as just an image, so 2D shared kernels are always fine
- Claims a shared, spatially aggregated stack can recover absolute position anyway
- Never asks whether the axis being shared over is stationary
- Believes more depth or data can train away a wrong architectural constraint