Why is a 3D convolution over a 16-frame video clip usually factorized into (2+1)D?
answer
- one more axis, three times the weights
- memory grows with clip length
- split spatial from temporal
- an extra nonlinearity between them
- pooled frame features lose ordering
basics
~20 sA 3D kernel spans time, height and width at once, so weights and compute scale with clip length too. The (2+1)D form splits it into a spatial convolution then a temporal one: fewer weights, an extra nonlinearity, easier optimization.
solid answer
~40 sA 3D convolution slides a `t x h x w` kernel across a clip, summing over input channels as a 2D one does, so a 3x3x3 kernel mapping 64 channels to 64 holds `27*64*64 = 110,592` weights and its compute scales with clip length. The (2+1)D factorization replaces it with a `1 x 3 x 3` spatial convolution followed by a `3 x 1 x 1` temporal one: `36,864 + 12,288 = 49,152` weights at the same widths, under half. Two further gains matter. A nonlinearity now sits between the spatial and temporal stages, and the two easier sub-problems optimize more readily than the joint one - which is why some designs widen the intermediate layer to spend the saved budget instead. You keep motion modelling, which per-frame features averaged over time throw away.
go deeper
Be ready to say that a 3D kernel spans frames as well as height and width, and that this is what lets a network see motion rather than just appearance.
Explain the parameter arithmetic - a 3x3x3 kernel is three times a 3x3 one at the same channel widths - and describe what the (2+1)D split replaces it with.
Show that you weigh a 3D model against a per-frame backbone by asking whether the label depends on motion, and that you know memory, not parameters, is what caps clip length.
Own the build-versus-borrow call: image pretraining is abundant and video pretraining is not, so inflation, factorization and clip budgeting are how you buy video capability without a video-scale data programme.
## What a 3D convolution is Give a network a clip shaped `channels x frames x height x width` - say 3 x 16 x 112 x 112 for a 16-frame action-recognition clip. A 3D convolution slides a kernel of size `t x h x w` over the three spatial-temporal axes at once. As always, the channel axis is summed rather than slid over, so one filter holds `t * h * w * C_in` weights plus a bias and emits one output channel over a 3D volume. The operator sees motion directly: a filter can respond to an edge that moves left between consecutive frames, which no single-frame feature can express. ## What it costs Cost is the problem. - **Parameters.** A 3x3x3 kernel from 64 to 64 channels is `3*3*3*64*64 = 110,592` weights. The 2D equivalent, 3x3, is `9*64*64 = 36,864`. Three times more for one extra axis of size 3. - **Compute.** Multiply-adds scale with output positions, and there are now `T` times as many of them as in a per-frame network. - **Activation memory.** Every stored activation carries a temporal axis, so training memory scales with clip length. This is usually what actually caps you: batch size collapses, and you end up with short clips and small batches. ## The (2+1)D factorization Split the `t x d x d` kernel into two convolutions in sequence: 1. `1 x d x d` - spatial only, applied within each frame independently; 2. `t x 1 x 1` - temporal only, applied across frames at each spatial location. At equal channel widths of 64 and `t = d = 3`: ``` full 3D: 3*3*3*64*64 = 110,592 (2+1)D : 1*3*3*64*64 = 36,864 + 3*1*1*64*64 = 12,288 = 49,152 ``` Under half, and the ratio holds generally: `27*C^2` against `12*C^2` for equal widths, or four-ninths. But parameter savings are only part of the argument, and the more interesting part is the other two: - **An extra nonlinearity.** Placing an activation between the spatial and temporal stages doubles the number of nonlinear transformations for the same nominal depth, which increases the complexity of functions the block can represent. - **Easier optimization.** Decomposing one joint spatio-temporal estimation into two smaller ones empirically yields lower training error at matched capacity. That is why designs in this family often widen the intermediate channel count until parameters match the full 3D block: the goal is the optimization and nonlinearity gain, with the size reduction available as an alternative way to spend the same budget. ## The alternatives you should name **Per-frame 2D backbone plus temporal aggregation.** Run a 2D network on each frame, then average or pool the features over time. Cheap, and it reuses everything known about image backbones, but average pooling is order-invariant: opening a door and closing a door produce the same pooled feature. Fine when the task is appearance-dominated - which scene is this, does this object appear - and wrong when the label depends on motion direction or ordering. **Inflation from 2D.** Rather than train 3D kernels from scratch, take a pretrained 2D kernel, repeat it along the temporal axis `t` times and divide by `t`. On a static clip - every frame identical - the inflated 3D network then produces the same activations as the 2D one, so you inherit the image-pretrained representation as an initialisation and fine-tune from there. This is the standard trick for making 3D models trainable on video datasets that are far smaller than image datasets. ## When full 3D earns its cost Choose full or factorized 3D when the label genuinely lives in the motion: distinguishing actions that share appearance but differ in direction or speed, fine-grained gesture recognition, anything where reversing the clip should change the answer. Choose per-frame features when the label is visible in a single frame. And be honest about the budget - 3D models want more data, more memory and longer training, so on a small dataset a 2D backbone with a light temporal head often wins outright. The clip itself is a design parameter too: how many frames, at what stride, and whether you aggregate several clips per video at inference. Longer clips buy temporal context at a memory cost that is linear in frames, and the choice interacts directly with how far your temporal kernels can reach.
- When would you skip 3D convolutions entirely and use per-frame 2D features with temporal pooling?When the label is visible in a single frame - scene classification, object presence, most appearance-dominated tasks - and when data or latency is tight. Average pooling over frames is order-invariant, so it cannot distinguish an action from its reverse; that is a fatal flaw for motion tasks and irrelevant for appearance ones. Decide by asking whether reversing the clip should change the label.
- How do you initialize 3D kernels from a network pretrained on still images?Inflate them: repeat each 2D kernel t times along the temporal axis and divide the weights by t. On a clip of identical frames the inflated network then reproduces the 2D network's activations exactly, so you start from the image representation rather than from noise. It is the standard way to make 3D video models trainable on datasets far smaller than image corpora.
- What usually limits clip length in practice - parameters or memory?Memory. Parameters do not depend on clip length at all, but every stored activation carries a temporal axis, so training memory scales linearly with frames. That is what forces short clips and small batches. The usual responses are fewer frames at a wider sampling stride, or downsampling the temporal axis early in the network.
saying these in an interview costs you the question
- Thinks a 3D convolution slides over the channel axis
- Assumes per-frame 2D features already capture motion
- Describes (2+1)D as a 3D kernel with smaller spatial size
- Ignores that activation memory grows with clip length
- Claims the only gain from factorizing is parameter count