skip to content

A ten-layer stack of stride-1 3x3 convolutions misses a 300-pixel lesion; what is its receptive field and how do you grow it?

level: seniorimportance: must knowfreq 56%

answer

  1. structure only, weights never change it
  2. add (k - 1) times the current jump
  3. stride multiplies everything that comes after
  4. ten 3x3 stride-1 layers reach 21 pixels
  5. the nominal field is an optimistic ceiling

basics

~20 s

Each stride-1 3x3 layer widens the receptive field by two pixels, so ten of them see just 21 input pixels, far too few for a 300-pixel lesion. Grow it by downsampling, by adding depth, or with larger kernels.

solid answer

~50 s

The receptive field is the region of the input one output unit can be influenced by, and it accumulates by the recurrence `r <- r + (k - 1) * jump`, with `jump <- jump * s` after each layer. Starting at `r = 1, jump = 1`, a stride-1 3x3 layer adds exactly 2 pixels, so ten of them reach 21. No training fixes a 21-pixel unit that must classify a 300-pixel lesion; the evidence is not in its input. Three levers grow it. Depth is additive and slow: you would need roughly 150 such layers. Downsampling is multiplicative, because a stride-2 layer doubles the jump and every later layer then adds twice as many input pixels; interleaving four 2x2 stride-2 pools into the same ten convolutions takes the field from 21 to 140. Larger kernels add `k - 1` per layer instead of 2, at quadratic parameter cost. In practice you combine the first two.

code

python · 14 lines
python
def receptive_field(layers):
    r, jump = 1, 1
    for k, s in layers:
        r += (k - 1) * jump
        jump *= s
    return r, jump

# ten stride-1 3x3 convolutions
plain = [(3, 1)] * 10
print(receptive_field(plain))          # (21, 1)

# the same ten convolutions, with a 2x2 stride-2 pool after every second one
pooled = ([(3, 1), (3, 1), (2, 2)] * 4) + [(3, 1), (3, 1)]
print(receptive_field(pooled))         # (140, 16)

go deeper

for a junior

Know that a receptive field is the input region one output unit can see, and that a stride-1 3x3 layer widens it by two pixels. Be able to add that up over a stack.

for a middle

Explain the recurrence including the jump term, and why stride multiplies the effect of every later layer rather than only widening its own. Expect to compute a field for a mixed stack on the spot.

for a senior

Diagnose with it: given an object size and an architecture, say whether the network can physically see the object before blaming the optimiser or the data. Know the resolution cost of the downsampling you propose.

for a principal

Own the tradeoff between context and localisation as an architectural decision, and set the team's rule of thumb for how much nominal field headroom an object scale demands given that the effective field is much smaller.

## Definition The **receptive field** of a unit in a convolutional network is the set of input pixels that can affect that unit's value. It is a purely structural property: it depends on kernel sizes, strides and the layer ordering, and not at all on the learned weights. If a pixel lies outside a unit's receptive field, no training procedure and no amount of data can make that unit sensitive to it. Two quantities are tracked together as you walk the stack forward: - `r` — the current receptive field size, in input pixels. Starts at 1. - `jump` — how many input pixels separate two adjacent positions of the current feature map. Starts at 1. For each layer with kernel size `k` and stride `s`: ``` r <- r + (k - 1) * jump jump <- jump * s ``` Read the first line carefully: a layer widens the field by `(k - 1)` **steps of the current map**, and each of those steps is worth `jump` input pixels. This is why stride matters so much: it does not widen the field on the layer where it appears so much as it multiplies the value of every widening that comes after. ## The stated stack Ten stride-1 3x3 convolutions. `jump` stays 1 throughout, and each layer contributes `(3 - 1) * 1 = 2`. So ``` r = 1 + 10 * 2 = 21 ``` Twenty-one pixels. A unit at the top of this stack sees a 21x21 patch of the image. Asking it to decide the presence of a 300-pixel lesion is asking it to judge a structure from about one twentieth of the structure's area. The model will do what a model always does when the evidence is missing: latch onto whatever local texture happens to correlate with the label in the training set, then fail out of distribution. Depth alone gave the network parameters, not perception. A useful sanity check when a network underperforms on large objects: compute this number before touching the learning rate. ## Lever 1 — more depth Each additional stride-1 3x3 layer adds 2 pixels, so the field grows **linearly** in depth. Reaching 300 pixels needs roughly (300 - 1)/2 ≈ 150 such layers. That is an enormous amount of computation, memory and optimisation difficulty bought purely to widen a window. Depth is the least efficient lever there is for this problem, which is the point most candidates miss when they propose "make it deeper" as the whole answer. ## Lever 2 — downsample A layer with stride 2 doubles `jump`. Every subsequent layer's `(k - 1)` steps are then worth twice as many input pixels, so the growth becomes **multiplicative** in the number of downsampling stages. Take the same ten 3x3 convolutions and insert a 2x2 stride-2 pool after every second one — four reductions in total: ``` conv, conv -> r = 5, jump = 1 pool -> r = 6, jump = 2 conv, conv -> r = 14, jump = 2 pool -> r = 16, jump = 4 conv, conv -> r = 32, jump = 4 pool -> r = 36, jump = 8 conv, conv -> r = 68, jump = 8 pool -> r = 76, jump = 16 conv, conv -> r = 140, jump = 16 ``` The same ten convolutions now reach 140 pixels instead of 21, and they do so on grids that are progressively cheaper to compute. This is the main reason downsampling stages exist at all in a classification trunk; the compute saving is a bonus, not the motivation. The cost is resolution: the final grid is coarse, and `jump = 16` means neighbouring output positions are 16 input pixels apart, so anything that must be localised precisely has lost that precision. For a pure classification decision this is fine. For a task that must draw a boundary, downsampling to get context and then restoring resolution is the design tension the architecture has to resolve. ## Lever 3 — larger kernels A `k x k` layer adds `(k - 1) * jump` rather than `2 * jump`. A 7x7 kernel adds three times as much per layer as a 3x3. The cost is quadratic in `k` for parameters and multiply-adds at that layer, which is why deep stacks of small kernels historically displaced shallow stacks of large ones — but as a deliberate early-layer choice, a larger kernel is a cheap way to start the field off wider before the downsampling stages begin. ## The theoretical field is an upper bound The number the recurrence gives you counts every pixel that *can* influence the output. It does not say every pixel influences it equally. Contributions accumulate through many overlapping paths, and there are far more paths reaching the centre of the field than the corners; the resulting weighting is approximately Gaussian, densest at the centre and decaying outward. The **effective** receptive field — the region actually carrying meaningful influence — is therefore a fraction of the nominal one, and that fraction shrinks as the stack deepens, because the nominal field grows faster than the effective one does. The corner pixels of a nominal 140-pixel field may contribute almost nothing. The practical consequence: treat the computed number as an optimistic ceiling. If your object is 300 pixels wide, a computed field of exactly 300 is not comfortable — you want headroom, typically a nominal field several times the object's extent. ## How to check empirically Two cheap diagnostics. Occlude a patch of the input and see whether a chosen output unit moves at all — outside the field it cannot. Or take the gradient of a single centre unit with respect to the input pixels and look at the magnitude map: it is nonzero only inside the theoretical field, and its shape reveals the centre-weighted decay of the effective field directly.

  • Why is the effective receptive field smaller than the one your recurrence computes?
    The recurrence counts every pixel that can reach the output, but not how strongly. Many more paths connect the centre of the field to the output than connect its corners, so influence decays roughly like a Gaussian from the centre. The usable region is a fraction of the nominal one, and that fraction shrinks with depth — so treat the computed number as a ceiling and design with headroom.
  • How would you verify a network's receptive field empirically rather than by arithmetic?
    Take the gradient of a single central output unit with respect to the input and look at where it is nonzero: that region is the theoretical field, and the magnitude pattern inside it shows the centre-weighted decay. An occlusion sweep gives the same answer more crudely — slide a blanked patch over the input and note where the unit's value stops responding.
  • What does downsampling to reach a large receptive field cost you?
    Spatial precision. After four stride-2 stages, adjacent output positions are sixteen input pixels apart, so anything requiring finer localisation than that has to be recovered by later architectural means. For a classification decision the cost is usually acceptable; for a task that must place a boundary accurately it is the central design tension.
  • Does adding depth alone eventually solve an undersized receptive field?
    In principle yes, in practice no. Stride-1 3x3 layers add two pixels each, so reaching 300 pixels takes around 150 of them — a linear lever against a problem that downsampling solves multiplicatively. You would pay enormous compute and optimisation difficulty for a window that four reduction stages would have given you.

saying these in an interview costs you the question

  • Multiplies kernel sizes together to get the field
  • Ignores stride when accumulating the field
  • Thinks training can widen a receptive field
  • Assumes every pixel in the field contributes equally
  • Proposes only more depth to reach a large field

context