skip to content

How do you compute a dilated temporal convolution stack's receptive field in time steps?

level: middleimportance: should knowfreq 44%

answer

  1. start at one, add each layer's reach
  2. reach is (k-1) times dilation
  3. doubling buys exponential coverage
  4. two convolutions per block, not one
  5. then divide by the sampling rate

basics

~20 s

Each layer adds (k-1)*d input steps, so the receptive field is 1 plus the sum of (k-1)*d over the layers. Width 3 with dilations 1, 2, 4, 8, 16 covers 63 steps - divide by the sampling rate for seconds.

solid answer

~40 s

Start at one step and add each layer's reach: `rf = 1 + sum over layers of (k-1)*d`. Width-3 convolutions with dilations 1, 2, 4, 8, 16 give `1 + 2*(1+2+4+8+16) = 63` steps. Doubling the dilation makes coverage grow exponentially in depth while parameters grow linearly. Then convert: 63 steps at 50 Hz is 1.26 s, but 63 samples of 16 kHz audio is under 4 ms, which is why an audio stack uses width 2 with dilations doubling to 512 - one cycle covers 1024 samples, 64 ms - and repeats the cycle. Design rule: name the longest dependency the task has, express it in steps, and stack cycles until the field covers it. A field shorter than the dependency is a ceiling no amount of data lifts.

code

python · 14 lines
python
def receptive_field(kernel, dilations, convs_per_level=1):
    rf = 1
    for d in dilations:
        rf += convs_per_level * (kernel - 1) * d
    return rf

cycle = [1, 2, 4, 8, 16]

print(receptive_field(3, cycle))                 # 63 steps
print(receptive_field(3, cycle, 2))              # 125, two convs per level
print(receptive_field(3, cycle + cycle))         # 125, two stacked cycles
print(receptive_field(3, cycle) / 50.0)          # 1.26 seconds at 50 Hz
print(receptive_field(3, cycle) / 16000.0)       # 0.0039 seconds at 16 kHz
print(receptive_field(2, [1, 2, 4, 8, 16, 32, 64, 128, 256, 512]))  # 1024

go deeper

for a junior

Be ready to apply the formula: start from one step and add (k-1) times the dilation for each layer. Know that width 3 with dilations 1, 2, 4, 8, 16 reaches 63 steps.

for a middle

Explain why doubling the dilation gives coverage that grows exponentially with depth while parameters grow linearly, and why the lower dilation levels keep the coverage gap-free.

for a senior

Show that you size the field from a named dependency measured in seconds, verify it against the sampling rate, and know that an undersized field is a ceiling no amount of data can lift.

for a principal

Own the tradeoff between span and deployment cost: a longer field means a longer history buffer and higher startup latency on a streaming device, so the horizon is a product constraint, not just an accuracy knob.

## The counting rule One convolution of width `k` and dilation `d` touches input offsets `{0, d, ..., (k-1)d}`, so it reaches `(k-1)*d` steps beyond its anchor. Stack layers and the reaches add: ``` rf = 1 + sum_over_layers( (k - 1) * d ) ``` (assuming stride 1 throughout - a stride multiplies the reach of everything above it, which is why temporal stacks usually keep stride 1 and buy their coverage with dilation instead). Work the canonical case: `k = 3`, dilations `1, 2, 4, 8, 16`. ``` rf = 1 + 2*(1 + 2 + 4 + 8 + 16) = 1 + 2*31 = 63 ``` Sixty-three input steps feed one output step. Repeat the same cycle a second time - dilations `1,2,4,8,16,1,2,4,8,16` - and you get `1 + 2*62 = 125`. ## Why doubling With dilation fixed at 1, `L` layers of width 3 give `rf = 1 + 2L`: coverage is linear in depth, so a 1000-step horizon needs 500 layers. With dilation doubling, the sum of a geometric series means `L` layers give roughly `2^L` coverage - exponential in depth, linear in parameters. That trade is the reason dilated stacks displaced deep undilated ones for long-range temporal modelling. A reasonable worry is whether doubling leaves holes: a layer at dilation 16 skips 15 of every 16 inputs. It does not leave holes, because the layers below it at dilations 1, 2, 4 and 8 have already blended every intermediate position into the values it reads. The gap-free property depends on that full cycle being present. Skipping levels - going 1, 4, 16 - genuinely does leave positions that contribute to some outputs and not others, which shows up as a periodic patterning of the output sometimes called gridding. ## Blocks, not layers Most practical temporal stacks use a residual block with **two** dilated convolutions at each dilation level rather than one. Each block then contributes `2*(k-1)*d`, so the same cycle `1,2,4,8,16` at width 3 covers `1 + 4*31 = 125` steps rather than 63. Count what your architecture actually contains; being off by a factor of two on the receptive field is the difference between covering the dependency and missing half of it. ## Convert to physical units Steps are not the unit anyone cares about. - 50 Hz wrist sensor: 63 steps is 1.26 s - comfortably more than a single arm swing, so a one-cycle stack is plausible for activity recognition on 2.56 s windows. - 16 kHz raw audio: 63 samples is 3.9 ms, which is nothing. A causal audio stack in the WaveNet style uses width 2 with dilations doubling `1, 2, 4, ... 512`; that gives `1 + 1*(1+2+...+512) = 1024` samples, or 64 ms, per cycle, and several cycles are stacked to reach a useful span. The same architecture drawing therefore means completely different things at different sampling rates. Always report the receptive field in seconds alongside steps. ## Sizing it The design procedure is: 1. Name the longest dependency the task requires - one gait cycle, one respiratory cycle, the lag at which a control action shows an effect. 2. Express it in time steps at your sampling rate. 3. Add dilation cycles until `rf` comfortably exceeds it. Comfortably, because a dependency sitting at the very edge of the field is seen through one tap of one kernel. Undershooting is a hard ceiling: information that never enters the field cannot be used, no matter how much data you collect or how long you train. Overshooting is a soft cost - more parameters and compute, and at inference in a streaming setting a longer history buffer must be retained before the model can emit anything at all, which is a real latency and memory constraint on device. If the field is short and you cannot add depth, the other levers are a wider kernel (linear in parameters), a larger maximum dilation, or downsampling the input so each retained step represents more time - which trades temporal resolution for span and is usually the right call when the signal is oversampled relative to the phenomenon.

  • Your receptive field covers 2 seconds but the dependency you need is 10 seconds long. What are your options?
    Four levers. Repeat the dilation cycle - each extra cycle adds its full sum again. Raise the maximum dilation so the top layers reach further. Widen the kernels, which buys reach linearly in parameters. Or downsample the input so each step carries more time, trading temporal resolution for span. Downsampling is usually right when the signal is oversampled relative to the phenomenon.
  • A residual block uses two dilated convolutions per dilation level. How does that change the count?
    Each level contributes twice its reach, so the per-layer term becomes 2*(k-1)*d. Width 3 over dilations 1, 2, 4, 8, 16 then covers 1 + 4*31 = 125 steps instead of 63. Counting layers rather than convolutions is a factor-of-two error, and it is the difference between covering your dependency and seeing half of it.
  • Is there a cost to making the receptive field far larger than you need?
    Yes, though a soft one. Extra depth means more parameters and compute, and in a streaming deployment the model cannot emit an output until it has buffered a full receptive field of history, which directly sets startup latency and memory on device. Undershooting is worse - it is a hard ceiling - but overshooting is not free.

A dilated stack is a set of nested rulers: each level measures with gaps twice as wide as the level below, so a handful of levels spans a distance that would take hundreds of evenly spaced marks.

saying these in an interview costs you the question

  • Adds kernel widths instead of (k-1) times dilation
  • Reports the receptive field in layers rather than time steps
  • Never converts steps into seconds for the sampling rate
  • Claims doubling dilations leaves gaps in coverage
  • Counts blocks as one convolution when they contain two

context