Convolutional Networks
You will learn why convolutions beat dense layers on images, how ResNet's skip connections made very deep nets trainable, and what gets built on a backbone — detection scored by mAP, segmentation by U-Net. Expect to compute output shapes by hand; interviewers use that arithmetic as a fast competence check.
on this pageshowhide
explore
- The Convolution Operator18 questions
- Kernels, Stride and Dilation4 questions
- Parameter Sharing and Locality3 questions
- Output Shape Arithmetic3 questions
- Pooling and Receptive Fields4 questions
- Temporal and 3D Convolutions4 questions
- Backbone Architectures11 questions
- Depth and Small Kernels3 questions
- Residual Connections4 questions
- Grouped and Depthwise Convolutions4 questions
- Detection and Segmentation17 questions
- Object Detection Heads4 questions
- IoU Losses and mAP5 questions
- Semantic Segmentation4 questions
- Instance and Panoptic Masks4 questions
questions
page 2 of 2A detector scores about 100,000 candidate boxes per image against roughly 10 objects - what breaks in training?
basics
~20 sWith a thousand background candidates per object, the summed training loss is dominated by easy background and the model drifts toward predicting nothing. Fixes constrain which candidates contribute: a fixed positive-to-negative sampling ratio, hard-negative mining, and an ignore band for ambiguous candidates.
Why can a detector win on mAP at IoU 0.5 but lose on mAP averaged over 0.50 to 0.95?
basics
~20 smAP at IoU 0.5 forgives loose boxes, while averaging over the ten thresholds from 0.50 to 0.95 in 0.05 steps scores localisation quality directly. A model that classifies confidently but draws sloppy boxes tops the first and collapses on the second.
Your detector finds buses reliably but misses 16-pixel traffic signs. Why does a single-scale head fail here?
basics
~20 sA head on the backbone's deepest feature map sees a stride near 32 pixels, so a 16-pixel sign occupies under one cell. Downsampling destroys its detail, and often no location is assigned to it. A feature pyramid fixes this.
When is a convolution's shared-weight assumption wrong for the data you are modelling?
basics
~10 sSharing asserts that the same local pattern means the same thing everywhere along an axis. It fails where position itself carries meaning: a spectrogram's frequency axis, absolute-coordinate targets, and tabular columns with no order.
A segmentation model reports 95% pixel accuracy but 42 mIoU — what explains the gap?
basics
~20 sPixel accuracy is dominated by whichever classes own most of the pixels, while mIoU averages per-class intersection over union with every class weighted equally. Rare classes scoring near zero barely move accuracy but drag the mean down hard.
A 300x300 image through five stride-2 stages breaks a decoder's skip concatenation — why?
basics
~20 s300 is not divisible by 32. The floor in the output-shape formula drops a row at every odd-sized stage, so the encoder produces 37 where the decoder's doubling produces 36, and concatenation needs identical height and width.
How would you choose between a two-stage and a one-stage detector under a 30 ms per-frame budget?
basics
~20 sBudget the worst case, not the average. A one-stage head costs the same on every frame; a two-stage head's second pass scales with proposal count, so busy scenes are its slowest. Decide on tail latency first, then on the accuracy you need.
What does AlexNet's 11x11 stride-4 first layer discard that a 3x3 stem keeps?
basics
~20 sFine spatial detail. A stride-4 first layer samples the image on a coarse grid, aliasing away structure finer than four pixels, and one wide linear filter plus one nonlinearity is a weaker map than a stride-1 3x3 stack.
You swapped 3x3 convolutions for depthwise-separable blocks, cut FLOPs 8x, but latency only halved — why?
basics
~10 sA depthwise stage removes arithmetic but not data movement: it streams the full activation tensor while doing only k-squared multiply-accumulates per element. Runtime becomes memory-bound, and one layer becoming two adds further passes.
When merging instance masks and stuff maps into a panoptic output, how do you resolve contested pixels?
basics
~20 sA panoptic output allows one label per pixel, so conflicts follow a fixed precedence: instance masks are painted in confidence order, each keeping only unclaimed pixels, and stuff fills the remainder. Fragments left too small are dropped.
What does a dilated (atrous) convolution buy you, and what does it cost?
basics
~20 sDilation spreads a kernel's taps apart, covering a wider region with no extra weights and no downsampling. The costs are sparse sampling between the taps, full-resolution activation memory, and gridding when layers repeat one rate.
When would you replace every 2x2 max pooling layer with a stride-2 convolution, and what does it cost?
basics
~20 sReplace it when the downsample itself should be learned: a stride-2 convolution trains a weighted summary per window instead of applying a fixed maximum. It costs parameters, multiply-adds and a useful prior, and pays off most when data is plentiful.
When would you concatenate skip features as DenseNet does rather than add them as ResNet does?
basics
~20 sConcatenation keeps every earlier feature map intact so later layers can select among them; addition merges them irreversibly but holds channel count fixed. Choose concatenation for feature reuse at moderate depth, addition for very deep throughput-sensitive backbones.
Why does a transposed-convolution decoder produce checkerboard artifacts in its output?
basics
~20 sA transposed convolution writes a scaled copy of its kernel into the output at stride spacing. When the kernel size is not divisible by the stride, some output positions receive more overlapping writes than others, and that uneven overlap prints a periodic checkerboard.
Why is a 3D convolution over a 16-frame video clip usually factorized into (2+1)D?
basics
~20 sA 3D kernel spans time, height and width at once, so weights and compute scale with clip length too. The (2+1)D form splits it into a spatial convolution then a temporal one: fewer weights, an extra nonlinearity, easier optimization.
Per-class AP on a 12-instance defect class swings 8 points on one missed box - how do you report it?
basics
~20 sWith 12 instances, recall moves in steps of about 8 points, so that class's AP is a coarse, high-variance statistic. Publish per-class AP beside its instance count with an uncertainty estimate, and never let the headline mean carry a ship decision it cannot support.
showing 31–46 of 46