skip to content

Convolutional Networks

You will learn why convolutions beat dense layers on images, how ResNet's skip connections made very deep nets trainable, and what gets built on a backbone — detection scored by mAP, segmentation by U-Net. Expect to compute output shapes by hand; interviewers use that arithmetic as a fast competence check.

on this pageshow

explore

questions

page 2 of 2

A detector scores about 100,000 candidate boxes per image against roughly 10 objects - what breaks in training?

level: seniorimportance: should knowfreq 42%

basics

~20 s

With a thousand background candidates per object, the summed training loss is dominated by easy background and the model drifts toward predicting nothing. Fixes constrain which candidates contribute: a fixed positive-to-negative sampling ratio, hard-negative mining, and an ignore band for ambiguous candidates.

open as a page

Why can a detector win on mAP at IoU 0.5 but lose on mAP averaged over 0.50 to 0.95?

level: seniorimportance: should knowfreq 55%

basics

~20 s

mAP at IoU 0.5 forgives loose boxes, while averaging over the ten thresholds from 0.50 to 0.95 in 0.05 steps scores localisation quality directly. A model that classifies confidently but draws sloppy boxes tops the first and collapses on the second.

open as a page

Your detector finds buses reliably but misses 16-pixel traffic signs. Why does a single-scale head fail here?

level: seniorimportance: should knowfreq 56%

basics

~20 s

A head on the backbone's deepest feature map sees a stride near 32 pixels, so a 16-pixel sign occupies under one cell. Downsampling destroys its detail, and often no location is assigned to it. A feature pyramid fixes this.

open as a page

When is a convolution's shared-weight assumption wrong for the data you are modelling?

level: seniorimportance: should knowfreq 41%

basics

~10 s

Sharing asserts that the same local pattern means the same thing everywhere along an axis. It fails where position itself carries meaning: a spectrogram's frequency axis, absolute-coordinate targets, and tabular columns with no order.

open as a page

A segmentation model reports 95% pixel accuracy but 42 mIoU — what explains the gap?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Pixel accuracy is dominated by whichever classes own most of the pixels, while mIoU averages per-class intersection over union with every class weighted equally. Rare classes scoring near zero barely move accuracy but drag the mean down hard.

open as a page

A 300x300 image through five stride-2 stages breaks a decoder's skip concatenation — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

300 is not divisible by 32. The floor in the output-shape formula drops a row at every odd-sized stage, so the encoder produces 37 where the decoder's doubling produces 36, and concatenation needs identical height and width.

open as a page

How would you choose between a two-stage and a one-stage detector under a 30 ms per-frame budget?

level: principalimportance: should knowfreq 47%

basics

~20 s

Budget the worst case, not the average. A one-stage head costs the same on every frame; a two-stage head's second pass scales with proposal count, so busy scenes are its slowest. Decide on tail latency first, then on the accuracy you need.

open as a page

What does AlexNet's 11x11 stride-4 first layer discard that a 3x3 stem keeps?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Fine spatial detail. A stride-4 first layer samples the image on a coarse grid, aliasing away structure finer than four pixels, and one wide linear filter plus one nonlinearity is a weaker map than a stride-1 3x3 stack.

open as a page

You swapped 3x3 convolutions for depthwise-separable blocks, cut FLOPs 8x, but latency only halved — why?

level: seniorimportance: nice to knowfreq 34%

basics

~10 s

A depthwise stage removes arithmetic but not data movement: it streams the full activation tensor while doing only k-squared multiply-accumulates per element. Runtime becomes memory-bound, and one layer becoming two adds further passes.

open as a page

When merging instance masks and stuff maps into a panoptic output, how do you resolve contested pixels?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

A panoptic output allows one label per pixel, so conflicts follow a fixed precedence: instance masks are painted in confidence order, each keeping only unclaimed pixels, and stuff fills the remainder. Fragments left too small are dropped.

open as a page

What does a dilated (atrous) convolution buy you, and what does it cost?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Dilation spreads a kernel's taps apart, covering a wider region with no extra weights and no downsampling. The costs are sparse sampling between the taps, full-resolution activation memory, and gridding when layers repeat one rate.

open as a page

When would you replace every 2x2 max pooling layer with a stride-2 convolution, and what does it cost?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Replace it when the downsample itself should be learned: a stride-2 convolution trains a weighted summary per window instead of applying a fixed maximum. It costs parameters, multiply-adds and a useful prior, and pays off most when data is plentiful.

open as a page

When would you concatenate skip features as DenseNet does rather than add them as ResNet does?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Concatenation keeps every earlier feature map intact so later layers can select among them; addition merges them irreversibly but holds channel count fixed. Choose concatenation for feature reuse at moderate depth, addition for very deep throughput-sensitive backbones.

open as a page

Why does a transposed-convolution decoder produce checkerboard artifacts in its output?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

A transposed convolution writes a scaled copy of its kernel into the output at stride spacing. When the kernel size is not divisible by the stride, some output positions receive more overlapping writes than others, and that uneven overlap prints a periodic checkerboard.

open as a page

Why is a 3D convolution over a 16-frame video clip usually factorized into (2+1)D?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

A 3D kernel spans time, height and width at once, so weights and compute scale with clip length too. The (2+1)D form splits it into a spatial convolution then a temporal one: fewer weights, an extra nonlinearity, easier optimization.

open as a page

Per-class AP on a 12-instance defect class swings 8 points on one missed box - how do you report it?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

With 12 instances, recall moves in steps of about 8 points, so that class's AP is a coarse, high-variance statistic. Publish per-class AP beside its instance count with an uncertainty estimate, and never let the headline mean carry a ship decision it cannot support.

open as a page

showing 31–46 of 46