Why must a causal 1D convolution pad only on the left of the sequence?
answer
- the future must not be visible
- all the padding on one side
- reach equals (k-1) times dilation
- shapes match, so nothing errors
- check every layer, not just the last
basics
~20 sCausality means output at step t may use only inputs up to t. Padding (k-1)*d zeros on the left, then a valid convolution, preserves length and enforces that. Symmetric padding centres the kernel on t, so it reads the future.
solid answer
~50 sA convolution is causal when output `t` depends only on inputs at or before `t`. The standard construction pads the left edge with `(k-1)*d` zeros - kernel width `k`, dilation `d` - then runs a valid convolution, which returns a sequence of the original length whose every position looks strictly backwards. Symmetric or `same` padding splits that pad across both ends, which centres the kernel on `t` and lets it read `(k-1)*d/2` future steps. Nothing crashes; the shapes are identical, so the bug is silent. Every layer in the stack has to be causal - one symmetric layer anywhere makes the whole model acausal, and the leak widens as it propagates upward. This matters whenever the future genuinely is not available at inference: streaming forecasting, live detection, autoregressive generation. If the entire window is recorded before you score it, a non-causal model is a legitimate and usually stronger choice.
go deeper
Be ready to state the rule: output at time t may depend only on inputs at or before t, and that is achieved by putting all the padding on the left before a valid convolution.
Explain the arithmetic - a left pad of (k-1) times the dilation - and why the bug is silent: the output shape is identical either way, so nothing raises an error.
Demonstrate the operational instinct: a leak shows up as excellent offline metrics and collapsing live performance, and you catch it with a perturb-the-future test rather than by reading the code.
Own the framing decision. Causality is a deployment constraint, not a virtue; decide whether the future is available at inference before the architecture is chosen, and make that assumption explicit in the design.
## The definition A temporal convolution is **causal** if the value it emits at time `t` is a function of inputs at times `<= t` only. That property is not a feature of the convolution operator; it is a property of how you pad and align it. ## The construction Take a kernel of width `k` and dilation `d`. The set of input offsets one output position touches is `{0, d, 2d, ..., (k-1)d}`, so the kernel reaches `(k-1)*d` steps away from its anchor. To keep output length equal to input length you must add `(k-1)*d` zeros somewhere. - Put **all** of them on the left, then run a valid (no extra padding) convolution: output `t` sees inputs `t-(k-1)d ... t`. Causal. - Split them evenly across both ends - the usual `same` padding: output `t` sees inputs `t-(k-1)d/2 ... t+(k-1)d/2`. Acausal, by exactly half the reach. With `k = 3, d = 4`, the left pad is `(3-1)*4 = 8` zeros. Note that dilation multiplies the pad; a stack whose dilations double must have pads that double alongside them. Forgetting that is the most common implementation slip, and again it does not change any tensor shape, so nothing complains. An equivalent formulation pads both ends and then discards the last `(k-1)*d` outputs. It is the same operator; the discard is doing the work the asymmetric pad would otherwise do. ## Why the failure is silent Suppose you build a WaveNet-style stack of dilated convolutions over raw 16 kHz audio and one block is padded symmetrically. Training runs. Validation - computed on complete recorded windows, where the future really is present - looks excellent, because the model has learned to peek a few milliseconds ahead and that is genuinely informative. Then you deploy on a live stream, where the future does not exist, and quality collapses. Offline metrics never warned you because train, validation and test all shared the leak. This is the temporal analogue of target leakage: the model is not wrong, the evaluation was. ## How to detect it A direct test beats any code review. Take one input sequence, copy it, and change the copy at every position after some index `t0` - random noise is fine. Push both through the model and compare outputs at positions `<= t0`. If a single one differs, the model is not causal, and the first index that moves tells you how far the leak reaches. Run it as a unit test; it is cheap and it catches the dilation-pad slip that reading the code does not. ## The layers compound Causality is a property of the composition, not of any one layer. If layer 3 peeks one step ahead and layers above it each have their own reach, the top of the stack sees far more than one future step, because layer 4 aggregates layer-3 outputs that were themselves contaminated. So the check is per-layer: every convolution, every residual branch, and any pooling or normalisation that mixes across time. ## When you should not be causal Causality costs accuracy. At the same parameter budget a model that may look both ways has twice the context per layer, and for a whole class of problems the future is available: - classifying a **recorded** activity window - all 2.56 seconds are on disk before you score them; - offline annotation, segmentation or captioning of a stored signal; - any batch job over completed sequences. In those settings a symmetric stack is the right answer and imposing causality is self-harm. Causality is mandatory only when the model must produce output at `t` while `t+1` has not happened yet: live monitoring, streaming detection, real-time control, and autoregressive generation where the model's own previous output is the next input. The senior version of this answer is therefore not "always pad left". It is: decide first whether the deployment has access to the future, then make the padding match that decision, then test the property rather than assume it.
- How would you prove an already-trained temporal model is causal?Perturb the future and watch the past. Duplicate an input sequence, randomise every step after index t0 in the copy, run both, and compare outputs at positions up to t0. Any difference means the model reads ahead, and the earliest changed index tells you how far. It is a cheap unit test and it catches dilation-pad mistakes that code review misses.
- Does only the first layer need causal padding, or all of them?All of them. Causality is a property of the whole composition: a single symmetric layer contaminates its outputs, and every layer above aggregates those contaminated positions, so the leak widens as it propagates. Audit every convolution, every residual branch, and anything else that mixes across the time axis.
- When is a non-causal temporal convolution the better choice?Whenever the full window exists before you score it - classifying a recorded activity segment, annotating a stored signal, any batch job over completed sequences. A symmetric stack sees twice the context per layer at the same cost, so forcing causality there just discards information for no operational benefit.
A causal convolution is live commentary: you may quote anything already said, never the replay. Symmetric padding is a commentator who has secretly seen the tape.
saying these in an interview costs you the question
- Says same padding is fine if the last layer is causal
- Forgets that dilation multiplies the required left pad
- Thinks causality only matters for generation, not detection
- Treats strong validation scores as proof of no leakage
- Imposes causality on offline scoring of complete windows