skip to content

Why does a transposed-convolution decoder produce checkerboard artifacts in its output?

level: seniorimportance: nice to knowfreq 38%

answer

  1. scatter, not gather
  2. each input writes a whole kernel
  3. overlap counts vary with position
  4. kernel not divisible by stride
  5. resample uniformly, then convolve

basics

~20 s

A transposed convolution writes a scaled copy of its kernel into the output at stride spacing. When the kernel size is not divisible by the stride, some output positions receive more overlapping writes than others, and that uneven overlap prints a periodic checkerboard.

solid answer

~50 s

Think of a transposed convolution as scatter rather than gather: for each input position it writes a scaled copy of the whole kernel into the output, and consecutive input positions write at offsets of one stride. Whether every output position ends up covered by the same number of writes depends entirely on the arithmetic. With a `4x4` kernel at stride 3, the writes overlap unevenly and a repeating pattern of positions receives twice as many contributions as their neighbours; in 2D that pattern is a checkerboard, and it appears in the logits as a periodic grid of over- and under-confident pixels, most visible near boundaries. Two fixes exist. Choose a kernel size divisible by the stride, such as `4x4` at stride 2, which makes the interior overlap uniform. Or decouple resampling from learning: upsample bilinearly by the required factor, then apply a `3x3` convolution. The second is the safer default because the resampling step is uniform by construction.

code

python · 12 lines
python
def overlap_counts(k, s, n_in):
    # each input position writes a scaled copy of the k-wide kernel,
    # consecutive writes offset by the stride s
    n_out = (n_in - 1) * s + k
    counts = [0] * n_out
    for i in range(n_in):
        for j in range(k):
            counts[i * s + j] += 1
    return counts

print(overlap_counts(4, 3, 5))  # kernel 4, stride 3 -> periodic 1,1,2 pattern
print(overlap_counts(4, 2, 5))  # kernel 4, stride 2 -> uniform interior

go deeper

for a junior

Know the name and the symptom: a regular grid pattern in an upsampled output, caused by how the upsampling operator overlaps its writes rather than by anything about the data.

for a middle

Be able to derive it. Work out the overlap counts for a given kernel size and stride, say why divisibility matters, and describe interpolation followed by a convolution as the alternative.

for a senior

Demonstrate the diagnosis. Recognise periodicity matching the decoder's upsampling factor as structural, rule out data and training causes quickly, and justify replacing the operator rather than tuning around it.

for a principal

Own the standard. Decide whether decoders in your codebase use learned upsampling at all, weigh the extra compute of resample-then-convolve against a class of bug that silently degrades boundary quality, and make it a default rather than a per-model debate.

## What a transposed convolution actually does A normal convolution *gathers*: each output position reads a window of the input and produces one number. A transposed convolution reverses the connection pattern, and the clearest way to picture it is as a *scatter*. For each input position, take the kernel, scale it by that input's value, and add it into the output at an offset. Consecutive input positions write at offsets separated by the stride. The output length in 1D is `(n_in - 1) * s + k` for kernel `k` and stride `s`, so a stride greater than one grows the map — which is why it is used to upsample in a decoder. Because the writes are additive and the kernel is wider than the stride, adjacent writes overlap. The critical question is whether every output position lies under the **same number** of writes. ## Where the unevenness comes from Output position `p` receives a contribution from input position `i` whenever `p - i*s` falls inside `[0, k)`. The count of such `i` values is periodic in `p` with period `s`, and it is constant only when `s` divides `k`. Concretely, with `k = 4` and `s = 3`, the interior overlap counts cycle as 1, 1, 2 — one position in every three collects double the contributions. With `k = 4` and `s = 2`, every interior position collects exactly two, and the pattern is flat. The classic bad case people meet in practice is a `3x3` kernel at stride 2. In two dimensions the count is the product of the row and column counts, so a 1D pattern of 1, 1, 2 becomes a 2D grid of 1, 2 and 4 — a literal checkerboard of magnitudes. Nothing about the learned weights causes this; it is baked into the geometry, and the network must spend capacity learning kernels that partially cancel it. ## What it looks like in a segmentation output In generative models the artifact is famous because you see it directly as texture. In segmentation it is easier to miss, because the argmax hides small logit wobbles inside confidently-classified regions. Where it shows is at boundaries and in low-confidence areas: a regular grid of misclassified or flickering pixels, thin periodic stripes along an edge, and a visible lattice if you look at a single class's probability map instead of the argmax. Stacking several transposed-convolution stages compounds the pattern, since each stage upsamples the previous stage's artifact along with the signal. ## The two fixes **Make the kernel divisible by the stride.** A `4x4` kernel at stride 2, or `2x2` at stride 2, gives uniform interior overlap. This removes the deterministic component of the artifact and is nearly free. It is not an absolute guarantee: the *learned* kernel can still be uneven across its own positions, so a weaker pattern can survive. Edges of the output also remain unevenly covered regardless of divisibility. **Separate resampling from learning.** Upsample with a fixed interpolation — bilinear or nearest neighbour — by the factor you need, then apply an ordinary `3x3` convolution at the new resolution. The interpolation treats every output position identically, so no uneven-overlap pattern can form, and the convolution still learns whatever refinement is needed. The cost is that the convolution now runs at the higher resolution, which is more compute than the transposed version, though in a decoder with few channels this is usually acceptable. This is the safer default, and it is why many segmentation decoders contain no transposed convolutions at all. ## Diagnosing it The tell is **periodicity with a period equal to the upsampling factor** or a product of the factors in the decoder. If you see a regular lattice whose spacing matches the decoder's stride arithmetic, the cause is structural, not a training problem, and more epochs will not remove it. If the pattern is irregular or aligned with image content, look elsewhere — at label noise, at augmentation, or at boundary annotation quality. Worth stating plainly: the artifact is a property of the upsampling operator, so it is independent of what loss you train with and independent of whether you have skip connections. Skips can mask it by supplying sharp features that dominate the merged result, which is why the same decoder can look clean in one architecture and speckled in another.

  • Does making the kernel size divisible by the stride guarantee no artifacts?
    No. It removes the deterministic uneven-overlap pattern in the interior, which is the dominant cause, but the learned kernel can still be uneven across its own positions and leave a weaker pattern, and the output borders stay unevenly covered. Fixed interpolation followed by a convolution avoids the mechanism entirely rather than balancing it.
  • What does upsample-then-convolve cost compared with a transposed convolution?
    The convolution now runs at the upsampled resolution, so its cost scales with the larger spatial size rather than the smaller input grid, and the interpolation itself is cheap but not free. In a decoder with modest channel counts this is usually a small share of total cost, and it buys an artifact-free operator.

saying these in an interview costs you the question

  • Blames checkerboard patterns on undertraining
  • Thinks transposed convolution reverses a convolution's values
  • Believes any stride-2 upsampling causes the artifact
  • Says the loss function is what creates the pattern
  • Cannot name a divisible kernel-stride pair as a fix

context