skip to content

How does splitting a convolution into g groups change its cost and its channel connectivity?

level: middleimportance: should knowfreq 50%

answer

  1. channels get partitioned, not shortened
  2. each filter sees a slice, not everything
  3. parameters divide by the group count
  4. who talks to whom becomes block-diagonal
  5. the limit case has one channel per filter

basics

~20 s

Grouping partitions input and output channels into g disjoint sets, each output group seeing only its own input group. Weights and multiply-accumulates both drop by a factor of g, but no channel combines across group boundaries.

solid answer

~50 s

A grouped convolution with `g` groups splits the `C_in` input channels into `g` slices of `C_in/g` and the `C_out` output channels into `g` slices of `C_out/g`; group `i`'s filters convolve only slice `i`, so both counts must be divisible by `g`. Each filter shrinks to `k*k*C_in/g` weights, and with `C_out` filters the layer costs `k*k*C_in*C_out/g` weights and the same fraction of the arithmetic. The price is connectivity: an output channel is blind to the other groups, and stacking grouped layers with nothing between them leaves the channels in permanent silos. The usual fixes are a dense 1x1 layer between them, or a parameter-free channel shuffle that permutes channels across groups. At the extreme `g = C_in` every filter owns one input channel, which is a depthwise convolution; ResNeXt sits in between, using 32 groups to spend a fixed budget on many narrow parallel paths instead of one wide one.

go deeper

for a junior

Know the picture: input and output channels are cut into g matching slices, each slice convolved on its own, and the weight count drops by a factor of g.

for a middle

Be able to derive kkC_in*C_out/g from the filter shape, state the divisibility constraint on both channel counts, and say that g equal to C_in gives you a depthwise convolution.

for a senior

Demonstrate that you have hit the silo problem in practice: diagnose stacked grouped layers with no mixing, and choose between a dense 1x1 and a channel shuffle knowing what each costs.

for a principal

Argue cardinality as a scaling axis against width and depth at a fixed budget, and be honest that the arithmetic saving from small groups may not survive contact with the target hardware.

## Definition A standard convolution is densely connected across channels: every one of the `C_out` filters spans all `C_in` input channels, so the weight tensor is `C_out x C_in x k x k`. A grouped convolution adds one integer hyperparameter, the group count `g`, and partitions both axes: - the input channels split into `g` contiguous slices of `C_in/g` channels each; - the output channels split into `g` slices of `C_out/g` channels each; - output slice `i` is produced only from input slice `i`. Both `C_in` and `C_out` must be divisible by `g`. Setting `g = 1` recovers the ordinary dense convolution. ## The arithmetic Each filter now spans `C_in/g` channels instead of `C_in`, so it holds `k * k * C_in / g` weights. There are still `C_out` filters in total, so ``` params = k * k * (C_in / g) * (C_out / g) * g = k * k * C_in * C_out / g ``` and the multiply-accumulates, `H_out * W_out * params`, fall by the same factor `g`. Concretely, a 3x3 layer with 256 input and 256 output channels holds `9 * 256 * 256 = 589,824` weights densely, `73,728` with eight groups, and `2,304` when `g = 256`. Note what does *not* change: the output shape, the spatial receptive field, and the number of activations produced. Grouping is purely a sparsity pattern imposed on the channel connectivity. ## The connectivity cost, and the silo problem The factor-`g` saving is not free capacity. Inside one grouped layer, an output channel is a function of only `C_in/g` input channels. If you stack two grouped layers with the same grouping and put nothing between them, the composition is still block-diagonal: features that entered group 3 can never influence group 5, no matter how deep the stack. The network degenerates into `g` independent narrow networks sharing an input and an output — an information silo. Two standard remedies: 1. **A dense 1x1 (pointwise) convolution between grouped layers.** It is ungrouped, so it mixes every channel with every other, at a cost of `C_in * C_out` weights. 2. **A channel shuffle.** Reorder the channel axis so that each group's outputs are redistributed across all groups before the next grouped layer, as ShuffleNet does. This costs no parameters and almost no arithmetic — it is a permutation — and restores cross-group flow. It is the cheap answer when even the dense 1x1 is too expensive, because that 1x1 is often the dominant cost in a compact block. A good interview answer names the silo problem unprompted; a weak one claims the grouping "doesn't really matter because gradients mix everything anyway", which is false — the block-diagonal structure applies to the backward pass exactly as it does to the forward pass. ## Cardinality: groups as a design axis ResNeXt reframed the group count as a third scaling axis it calls *cardinality*, alongside depth and width. Its block is a 1x1 reduction, a 3x3 grouped convolution with 32 groups, and a 1x1 expansion, wrapped in a residual addition. Because grouping frees up a factor of `g`, the 3x3 stage can be made proportionally wider at the same parameter and arithmetic budget: instead of one wide dense path you get 32 narrow parallel paths whose outputs are then recombined by the following dense 1x1. The empirical claim ResNeXt makes is that at a *matched* budget, buying cardinality is a better trade than buying more width or more depth. Interviewers like this because it forces you to say what is held fixed — the budget — rather than the vague "more groups is cheaper". ## The two endpoints - `g = 1`: the dense convolution. Maximum cross-channel expressivity, maximum cost. - `g = C_in` (with `C_out` a multiple of `C_in`): every filter owns exactly one input channel. This is a **depthwise** convolution — the extreme of grouping, where channel mixing is entirely delegated to whatever layer follows. Seeing depthwise as "grouped convolution taken to its limit" is the cleanest mental model, and it explains why a depthwise layer is nearly always followed by a pointwise 1x1: something has to reintroduce the cross-channel mixing that grouping removed. ## Practical notes Grouping has a hardware history — it was first used to split a large model's channels across two devices — and it retains a hardware flavour today: the FLOP reduction from grouping is real, but very small groups can execute inefficiently, because the layer becomes many small independent convolutions with little data reuse. Choosing `g` in practice means balancing the parameter saving, the loss of cross-channel expressivity, whether a mixing layer follows, and how the grouped kernel actually runs on the target device. Divisibility also constrains architecture search: widths must stay multiples of `g` throughout the stage.

  • Two grouped convolutions are stacked with the same group count and nothing between them. What goes wrong?
    The composition stays block-diagonal, so channels are permanently siloed inside their group and the stack behaves as `g` independent narrow networks. Fix it with a dense 1x1 layer between them, which mixes every channel at a cost of `C_in * C_out` weights, or with a parameter-free channel shuffle that permutes channels across groups before the next grouped layer.
  • What does ResNeXt's cardinality of 32 actually buy compared with just making the block wider?
    At a fixed parameter and arithmetic budget, grouping frees up a factor of 32 that is spent on making the grouped 3x3 stage proportionally wider. You get 32 narrow parallel transformations recombined by the following dense 1x1, rather than one wide dense transformation. The claim is that this trade is a better use of the same budget than extra width or depth.
  • What constraint does the group count place on the layer's channel counts?
    Both `C_in` and `C_out` must be divisible by `g`, since each group takes an equal slice of each axis. That propagates through a stage: every width you choose has to remain a multiple of the group count, which quietly restricts architecture search and any later change to the layer's width.

Groups are like splitting one wide team into g sub-teams that never talk to each other: each is g times cheaper to run, but no member ever hears what the other sub-teams found unless you deliberately hold a joint meeting afterwards.

saying these in an interview costs you the question

  • Says grouping reduces the spatial receptive field
  • Claims gradients mix channels across groups anyway
  • Thinks groups shrink the output feature map
  • Ignores the divisibility constraint on channel counts
  • Presents cardinality without saying the budget is held fixed

context