skip to content

In a 1D convolution over a 6-channel 50 Hz sensor stream, what does the kernel slide over?

level: juniorimportance: must knowfreq 65%

answer

  1. one axis moves, two are summed
  2. channels are combined, not traversed
  3. a filter is a k-by-C block plus bias
  4. width divided by sampling rate
  5. does the kernel span the event

basics

~20 s

A 1D kernel slides only along time. Each filter carries weights for every input channel at every offset in its width, so one kernel position sums k times 6 numbers into a single output value.

solid answer

~50 s

Only the time axis is convolved. A filter of width `k` on a `C`-channel stream is a `k x C` block of weights plus a bias; at each time position it multiplies that block against the window of the last `k` steps across all six channels and sums everything into one scalar. The channel axis is summed over, never slid over, so accelerometer and gyroscope axes are combined at every layer rather than processed independently. One filter produces one output channel over time; the layer stacks as many filters as you want output channels. The width is only meaningful in physical units: at 50 Hz a width of 9 spans 9/50 = 0.18 s, and a 128-step window is 2.56 s. Size the width against how long the motion you care about actually lasts, not against a number that looked reasonable for images.

go deeper

for a junior

Be ready to say which axis a 1D convolution moves along and that each filter reads every input channel at once. Convert the kernel width into seconds before you talk about whether it is big enough.

for a middle

Explain that a width-k filter on C channels holds k times C weights plus a bias, and that the channel axis is summed rather than traversed. Say how the sampling rate changes your choice of width.

for a senior

Show that you size kernels from the signal itself - how long the event you care about lasts - and that you re-derive widths whenever the stream is resampled or a sensor is added to the array.

for a principal

Own the input representation as a product decision: sampling rate, window length and channel set trade latency and battery against accuracy, and they silently fix every kernel width downstream.

## The shape of the problem Human-activity recognition from a wrist device gives you a stream shaped `channels x time`: three accelerometer axes and three gyroscope axes sampled at 50 Hz, cut into windows of 128 steps. That is 6 channels by 128 time steps, or 2.56 seconds of movement per example. A 1D convolution is the operator that walks a small learned kernel along the time axis of that array. ## What one filter is A single filter of width `k` on a `C`-channel input is **not** a vector of `k` numbers. It is a `k x C` array of weights plus one bias - here `9 x 6 = 54` weights plus a bias for a width-9 filter. At output position `t` the filter takes the window of `k` consecutive time steps, multiplies element-wise across **all six channels at once**, sums every product, adds the bias, and emits one scalar. So the operator reduces over two axes and slides over one: - **slid over:** time - **summed over:** the kernel offsets within the window, and the channel axis A layer with `F` filters emits `F` output channels over time, each one a different learned mixture of the six inputs. Those `F` channels are then the input channels of the next layer, which sums over them in exactly the same way. Depth therefore mixes channels at every level; nothing keeps the accelerometer stream separate from the gyroscope stream unless you deliberately build separate branches. ## Why people get this wrong The two standard misreadings are worth naming. **"Each channel gets its own 1D filter."** That is a different operator - a depthwise convolution - and it is a design choice, not the default. The plain convolution deliberately mixes channels, which is what lets it learn that a particular pattern of forward acceleration co-occurring with a wrist rotation means one activity rather than another. If you convolved each channel independently, no layer would ever see a cross-sensor pattern. **"The kernel slides across channels too."** That is a 2D convolution applied to the `6 x 128` array as if it were an image, and it quietly asserts that channel 2 sits between channel 1 and channel 3 in a meaningful way. Sensor axes have no such ordering; reordering the leads on the input would change the model. The 1D convolution avoids the question entirely by treating the channel axis as an unordered set that is fully connected inside every filter. ## Kernel width is a duration The number that matters is not `k` but `k / sampling_rate`. - 50 Hz wrist sensor, `k = 9` -> 0.18 s. Enough to see one arm swing beginning, nowhere near a whole gait cycle. - 500 Hz 12-lead ECG, `k = 5` -> 10 ms. A QRS complex - the sharp ventricular depolarisation spike - runs roughly 80-100 ms in a normal beat, about 40-50 samples at that rate. A width-5 kernel sees a fragment of one edge of it. That second case is the standard interview trap. The channels here are leads, not colours, and the sampling rate is ten times higher, so a kernel width copied from an image network covers a physiologically meaningless slice of signal. You have three honest fixes: widen the kernel, stack more layers so the composed span grows with depth, or downsample the stream first so each step is worth more time. The corollary is that kernel widths do not survive a change of sampling rate. Resample the same signal from 50 Hz to 100 Hz and every layer suddenly covers half as much time; a model that was reading a whole gesture is now reading half of one. Either rescale the widths or resample back before inference. ## What to say in an interview State the axis (time), state what is summed (kernel offsets and channels), give the weight count of one filter (`k x C` plus bias), and immediately convert the width into seconds for the sensor at hand. Candidates who go straight to the physical duration signal that they have actually shipped a model on a real stream rather than copied a block diagram.

  • How would you pick the kernel width for arrhythmia detection on a 12-lead ECG sampled at 500 Hz?
    Start from the waveform, not from habit. A QRS complex lasts roughly 80-100 ms, which is about 40-50 samples at 500 Hz, so the model needs a span of at least that before it can recognise one beat and considerably more before it can compare beats. Get there with a wider first kernel, with depth, or by decimating the signal - but state the target span in milliseconds first and derive the width from it.
  • Your sensor stream is resampled from 50 Hz to 100 Hz. What has to change in the network?
    Every kernel now covers half as much real time, so a stack that spanned 1.2 s spans 0.6 s. Either double the widths or dilations to hold the span constant, or downsample back to the rate the architecture was designed for. Nothing errors out - the shapes still work - which is exactly why this regression is easy to ship.
  • Why not feed the 6 x 128 window to a 2D convolution instead?
    A 2D kernel also slides along the channel axis, which asserts that neighbouring channels are neighbours in some meaningful sense. Sensor axes and ECG leads have no such ordering, so the model would become sensitive to how you happened to order the columns. The 1D convolution treats channels as an unordered set and connects all of them inside every filter.

Think of the six channels as six parallel tracks on a tape and the kernel as a playback head that reads all six tracks at once. It travels only forwards along the tape; it never moves sideways between tracks.

saying these in an interview costs you the question

  • Says the kernel slides across channels as well as time
  • Thinks each input channel gets its own separate 1D filter
  • Quotes a kernel width without converting it to seconds
  • Copies an image-network width onto a 500 Hz signal
  • Assumes a filter of width 9 holds only 9 weights

context