Deep Learning
You will learn how neural networks actually work below the LLM layer: what backprop computes, why optimizers and normalization matter, the classic CNN/RNN/attention architectures, and how to diagnose training that goes wrong. Senior DL and ML interviews assume this foundation before any framework or LLM question.
on this pageshowhide
explore
- Neural Network Basics29 questions
- Neurons and Layers12 questions
- Depth and Expressivity6 questions
- Heads and Training Loop11 questions
- Activations and Loss Functions36 questions
- Hidden Unit Nonlinearities9 questions
- Matching Objective to Task11 questions
- Imbalance, Surrogates and Trade-offs16 questions
- Backpropagation and Autodiff30 questions
- Computational Graphs14 questions
- Reverse Sweep Mechanics6 questions
- Gradient Pathologies10 questions
- Optimizers and Learning-Rate Schedules38 questions
- Stochastic Steps and Momentum11 questions
- Adaptive Update Rules14 questions
- Schedules and Batch Size13 questions
- Regularization and Normalization28 questions
- Loss-Term Regularizers7 questions
- Penalty-Free Regularizers11 questions
- Normalization Layers10 questions
- Convolutional Networks46 questions
- The Convolution Operator18 questions
- Backbone Architectures11 questions
- Detection and Segmentation17 questions
- RNNs, LSTMs and Attention39 questions
- Recurrence and BPTT19 questions
- Gated Recurrent Cells8 questions
- Seq2Seq and Alignment12 questions
- Autoencoders, GANs and Generative Models45 questions
- Latent-Variable Models17 questions
- Generator and Discriminator17 questions
- Diffusion and Likelihood Families11 questions
- Training Dynamics and Debugging38 questions
- Starting the Run Right13 questions
- Loss Curve Reading11 questions
- Diagnosing a Broken Run14 questions
- Transfer Learning and Embeddings44 questions
- Learning Transferable Features19 questions
- Adapting a Pretrained Model17 questions
- Representation Quality8 questions
- GPU Training Fundamentals32 questions
- Arithmetic and Precision10 questions
- Memory Budget14 questions
- Scaling Across Devices8 questions
- Deep Reinforcement Learning34 questions
- Value Networks11 questions
- Policy Optimization13 questions
- Sample Cost and Reliability10 questions
- Model Efficiency and Compression35 questions
- Knowledge Distillation7 questions
- Pruning and Quantized Training17 questions
- Designing Small Networks11 questions
- Graph Neural Networks27 questions
- Graphs and Message Passing10 questions
- Architecture Families10 questions
- Tasks and Scale7 questions
questions
501 · 14 sectionsWhat does a single artificial neuron compute from its inputs?
basics
~20 sAn artificial neuron multiplies each input by a learned weight, sums the products, adds one learned bias, and passes that single number through a fixed nonlinear function. Only the weights and the bias are learned.
In a 12288-512-128-10 MLP fed a batch of 32 flattened aerial tiles, what shape is each activation?
basics
~10 sInput (32, 12288), then (32, 512), then (32, 128), and finally (32, 10). The batch size stays on the leading axis at every layer; only the trailing feature width changes, once per weight matrix.
How do you count the parameters of a 784-128-10 fully-connected neural network?
basics
~20 sEach fully-connected layer holds inputs times outputs weights plus one bias per output unit. For 784-128-10 that is 784 x 128 + 128 = 100,480 and 128 x 10 + 10 = 1,290, totalling 101,770 parameters.
Why use one sigmoid per label instead of a softmax over the same output layer?
basics
~20 sPer-label sigmoids fit tasks where one example can carry several labels at once: each output is its own independent probability and the outputs need not sum to one. A softmax makes the classes share a single probability budget, which only fits mutually exclusive classes.
Why must a softmax classification head have exactly one output unit per class?
basics
~20 sA softmax head emits one score per class and normalizes those scores into probabilities that sum to one. K classes need K units: with fewer, some class can never be predicted; extra units create classes no label ever selects.
When does a network's output head need binary cross-entropy rather than categorical cross-entropy?
basics
~20 sUse binary cross-entropy when each output is an independent yes/no decision, so several can be true at once. Use categorical cross-entropy when exactly one of the classes is correct and the outputs must compete for a single unit of probability.
What is a logit in a neural classifier, and why do losses take logits rather than probabilities?
basics
~20 sA logit is the raw, unbounded score a network's final layer emits before any squashing function turns it into a probability. Losses take logits so the squashing and the logarithm are computed together, which avoids overflow and log-of-zero.
Why is classification accuracy useless as a training loss for a neural network?
basics
~20 sAccuracy counts correct predictions, so it is a step function of the model's scores. Nudging a weight usually changes nothing at all, leaving a gradient of zero almost everywhere and no slope for gradient descent to follow.
What is ReLU, and why is it the default hidden-layer activation in deep networks?
basics
~20 sReLU computes max(0, z): positive pre-activations pass through unchanged and negative ones become zero. Its derivative is exactly 1 wherever the unit is active, so gradients flow back unshrunk, and evaluating it costs a single comparison.
What does per-class loss weighting change when one intent has 200,000 training examples and others have 20?
basics
~20 sPer-class weighting multiplies each example's loss by a factor set by its class, so errors on rare intents contribute more gradient. It rebalances the effective class prior the model fits. It adds no new information about rare classes.
Why does a training forward pass keep its intermediate tensors when an inference-only forward pass can discard them?
basics
~20 sTraining runs a backward pass, and each op's local gradient is a function of the tensors that op actually saw. Inference never runs backward, so an intermediate can be released the moment the next op has consumed it.
How does an elementwise max such as ReLU route gradient in backpropagation?
basics
~20 sA max passes the incoming gradient straight to whichever input won and gives exactly zero to the loser. For ReLU that is a 0/1 mask: positive inputs pass the upstream gradient through unchanged, non-positive ones cut it to zero.
In the backward pass, how does a max-pooling layer route gradients compared with average pooling?
basics
~20 sMax pooling sends a window's entire upstream gradient to the one input that was the maximum and exactly zero to the others. Average pooling spreads it evenly, so each input in a k-by-k window receives one k-squared-th of it.
In reverse-mode autodiff, why is a reused tensor's gradient the sum of its consumers' gradients?
basics
~20 sA reused tensor appears in several terms of the chain rule, so its gradient is the sum of one contribution per consumer. The reverse sweep adds each contribution into a single buffer as it reaches that consumer.
What is the computational graph a forward pass records, and in what order does the backward sweep replay it?
basics
~20 sThe forward pass records a directed acyclic graph whose nodes are ops and whose edges are the tensors between them. The backward sweep replays it in reverse topological order: a node is visited only after every op that consumed its output.
Why is a neural network's learning rate usually decayed over the course of training?
basics
~20 sA large step keeps stochastic gradient descent bouncing around a minimum instead of settling in it. Shrinking the step later in training lets the noisy gradient estimates average out, so the parameters settle and the loss stops oscillating.
In SGD with momentum, what does the velocity term add to a plain gradient step?
basics
~20 sMomentum keeps a running average of past gradients, the velocity, and steps along it rather than along the raw gradient. Consistent directions build up and move faster; components that flip sign each step cancel, so the path stops zig-zagging.
In AdaGrad, how does the running sum of squared gradients set each parameter's step size?
basics
~20 sAdaGrad divides the global learning rate by the square root of each parameter's own accumulated squared gradients. Consistently large gradients get small steps, quiet parameters keep large ones, and because the sum only grows, every step shrinks over time.
Why does Adam divide its moment estimates by one minus the decay rate raised to the step count?
basics
~20 sBoth moving averages start at zero, so early on they read too small — the squared-gradient average far more so. Dividing each by one minus its decay rate to the step count removes that startup bias.
What two running averages does Adam maintain, and how does its update rule combine them?
basics
~20 sAdam keeps two exponential moving averages per parameter: one of the gradient, one of the squared gradient. The update is the first divided by the square root of the second, so each parameter gets its own step size.
Why does a batch-normalized network give different predictions in training mode than in evaluation mode?
basics
~20 sIn training, batch normalization standardizes each channel with the current mini-batch's mean and variance. In evaluation it uses running averages collected during training, so predictions become deterministic and stop depending on which other samples share the batch.
Why does applying random label-preserving transforms to training images reduce overfitting?
basics
~20 sRandom label-preserving transforms show the network a different version of each image every epoch, so it cannot memorise exact pixels. This effectively enlarges the training set and bakes in invariances such as small shifts and rotations, cutting variance.
Why does dropout reduce overfitting in a fully connected network?
basics
~20 sDropout zeroes each hidden unit at random on every training step, so no unit can rely on a particular other unit being present. Training therefore fits many thinned, weight-sharing subnetworks whose behaviour is averaged at evaluation.
Which axis does LayerNorm normalize over, and why does batch size not change its output?
basics
~20 sLayerNorm computes a mean and variance across the feature dimensions of each sample on its own, then rescales with a learned per-feature gain and bias. No other sample enters the statistic, so batch size and batch composition cannot change the output.
What does weight decay do to a neural network's weights during training?
basics
~20 sWeight decay pulls every decayed weight slightly toward zero on each update, on top of the gradient step. The result is a smaller-norm solution: the network fits with less extreme weights, which usually improves held-out performance.
What does a 1x1 convolution compute in a CNN, and why would you add one?
basics
~10 sA 1x1 convolution mixes channels at one spatial position: each output channel is a learned linear combination of all input channels at that pixel. It changes channel count cheaply, leaving height and width untouched.
How do semantic, instance, and panoptic segmentation differ in what they label each pixel?
basics
~20 sSemantic segmentation labels each pixel with a class but no identity, so touching objects merge. Instance segmentation returns one mask per countable object. Panoptic gives every pixel exactly one class and, for countable classes, an instance id.
In object detection, how does IoU between a predicted and ground-truth box decide a detection is correct?
basics
~20 sIoU is the overlap area of a predicted and a ground-truth box divided by the area they jointly cover. A prediction counts as correct when its IoU with an unclaimed ground-truth box of the same class clears a threshold, commonly 0.5.
In a CNN, what does one 3x3 filter over a 4-channel input actually contain?
basics
~10 sA 3x3 filter over a 4-channel input is a 3x3x4 weight block plus one scalar bias. The kernel always spans every input channel, sums across them, and produces a single output channel.
How does an attention-based seq2seq decoder build its context vector at each output step?
basics
~20 sAt each output step the decoder scores its current state against every encoder state, softmaxes those scores over source positions into weights summing to one, then averages the encoder states with those weights — that average is the context vector.
What does backpropagation through time do to a recurrent network's computation graph?
basics
~10 sBackpropagation through time unrolls the recurrent loop into a chain with one copy of the cell per time step, then runs ordinary backpropagation over that finite graph. All copies share one set of weights.
In a vanilla encoder-decoder RNN for translation, what does the encoder pass to the decoder?
basics
~10 sOnly its final hidden state: one fixed-size vector, plus the cell state if the encoder is an LSTM. That single tensor initialises the decoder, and nothing else crosses between the two stacks.
Why must a batch of variable-length sequences be padded, and what does the mask do?
basics
~20 sA batch must be one rectangular array, so shorter sequences are padded out to the longest length. The mask marks which steps are real, keeping pad steps out of the loss, out of pooling, and out of the state you read.
In a GAN's minimax objective, what is the discriminator maximizing and the generator minimizing?
basics
~20 sThe discriminator maximizes the log-probability of labelling real data real and generated data fake. The generator minimizes that same quantity, pushing the discriminator toward calling its samples real. One shared value function, optimized in opposite directions.
What is mode collapse in GAN training, and how do you spot it in generated samples?
basics
~20 sMode collapse is when a GAN generator maps many different noise vectors to a few nearly identical outputs, covering only part of the real data. You spot it by decoding a fixed batch of noise vectors and seeing the samples repeat.
Why does a plain autoencoder need a bottleneck narrower than its input?
basics
~20 sAn autoencoder is trained to copy its input, so a wide enough code can simply learn the identity map: near-zero loss, nothing learned. A code narrower than the input makes exact copying impossible and forces the encoder to keep only rebuildable structure.
Why does a denoising autoencoder corrupt its input but score the reconstruction against the clean original?
basics
~20 sScoring against the clean original means copying the input no longer wins. The network has to use structure shared across the data to remove the corruption, so the code ends up describing that structure instead of the input itself.
In a conditional GAN, why does the discriminator receive the class label too?
basics
~20 sOnly the discriminator can create pressure to obey the label. If it judges images alone, any realistic image passes, so the generator's cheapest strategy is to ignore its label input and reproduce the overall data distribution.
Training loss oscillates with an identical pattern in every epoch — what data-loading bug does that suggest?
basics
~20 sThe batches are being read in the stored file order without shuffling. If that order is sorted by class, every batch holds one class, the loss swings as the class changes, and the same swings repeat each epoch.
How do you tell a too-high learning rate from a too-low one by the shape of the training loss curve?
basics
~20 sA learning rate that is too high drops the loss fast, then leaves it oscillating inside a band or diverging; one that is too low gives a smooth, near-linear descent still falling at the last epoch.
Why does initializing every weight in a network's hidden layers to zero break training?
basics
~20 sAll units in a layer then compute the same output and receive the same gradient, so they update identically and stay duplicates forever. The layer has the power of a single unit. Random asymmetric values break that tie.
How does Grad-CAM turn a convolutional network's feature maps into a class-specific heatmap?
basics
~20 sGrad-CAM averages the gradients of one class score over each feature map of a chosen convolutional layer to get per-channel weights, sums the feature maps with those weights, and applies ReLU. The coarse result is upsampled onto the image.
Which per-layer statistics reveal dead or saturated units in a neural network's training run?
basics
~20 sPer layer, log the fraction of units that output zero for every example in a fixed probe batch, plus the fraction of tanh or sigmoid outputs past 0.99 in magnitude. Both are per-unit across the batch, not per example.
What is catastrophic forgetting when a pretrained network is fine-tuned on a new task?
basics
~20 sCatastrophic forgetting is the sharp drop in a network's performance on its original task after it is trained on a new one. Fine-tuning minimises only the new task's loss, so the shared weights drift away from the old solution.
In unsupervised domain adaptation, what is covariate shift and why does it hurt a trained network?
basics
~20 sCovariate shift means the input distribution moves from source to target while the rule mapping input to label is unchanged. The network was fitted where source data lived, so target inputs land where its decision boundary was never pinned down.
Why is the fine-tuning learning rate for a pretrained network far smaller than its pretraining rate?
basics
~20 sPretraining already puts the weights in a good region, so fine-tuning only needs small corrections. A large step moves every weight far enough to destroy the learned features, and a small target dataset cannot rebuild them.
How do you replace a pretrained model's 1000-class head for a 7-class task?
basics
~20 sDiscard the old output layer and attach a new, randomly initialised one with seven outputs, keeping the layers beneath it. The old head maps to the wrong label set, so none of its weights are reusable.
Why does face verification use an embedding with a distance threshold instead of an N-way classifier?
basics
~20 sA classifier can only recognise identities it was trained on, so every new person means retraining. An embedding model learns a distance where same-person pairs land close together, so a new identity is enrolled by storing one vector.
What does activation checkpointing trade away to cut a training step's activation memory?
basics
~20 sActivation checkpointing trades compute for memory. Instead of keeping every layer's forward outputs alive until the backward pass, it keeps only a few segment boundaries and recomputes the rest on demand, costing roughly one extra forward pass per step.
A training run's GPU sits at 30% utilisation while the host CPU is pinned decoding image tiles — what limits step time?
basics
~20 sThe input pipeline, not the model. The GPU idles waiting for batches while the host decodes and resizes images. Fixes are overlapping data preparation with compute, adding host-side parallelism, and pre-processing tiles into a cheaper stored form.
Why can a gradient that is nonzero in 32-bit floats round to exactly zero in 16-bit?
basics
~20 sA 16-bit float's 5 exponent bits bottom out near 6e-5 for normal values and near 6e-8 once subnormals run out. A gradient smaller than that has no representation, so it stores as exactly zero and that weight stops moving.
What is mixed-precision training, and where do its speed and memory wins come from?
basics
~20 sMixed-precision training runs the forward and backward passes in a 16-bit float while keeping a single-precision copy of the weights. Wins come from moving half the bytes and from matrix units that multiply 16-bit inputs at much higher throughput.
Why does halving the training batch size cut device memory when the weights are unchanged?
basics
~20 sWeights, gradients and optimizer state are sized by the parameter count, not by batch size. Activations, the per-layer outputs held for the backward pass, exist once per sample, so halving the batch roughly halves that share.
In deep RL, why does a reward that is zero for hundreds of steps stall learning?
basics
~20 sWith no non-zero reward, every return and every value target is zero, so the policy gradient and the TD error are exactly zero. Nothing is learned slowly; there is no signal at all until the agent stumbles onto a reward.
Why does a deep Q-network train from a replay buffer instead of the transitions as they arrive?
basics
~20 sConsecutive transitions are almost identical, so learning from them in order gives correlated, unstable updates. A replay buffer stores past transitions and samples shuffled batches from them, breaking that correlation and letting each transition be reused many times.
In an advantage actor-critic, what does the critic's learned value function contribute to the policy update?
basics
~20 sThe critic learns a state-value estimate V(s). The actor is updated with the advantage - reward plus gamma times the next state's value, minus the current state's value - a signal centred on zero that reinforces only actions better than average.
How does PPO's clipped surrogate objective bound how far one update moves the policy?
basics
~20 sPPO maximises min(r*A, clip(r, 1-eps, 1+eps)*A), where r is the new-over-old action probability ratio and A the advantage. Once r passes the bound in the direction the advantage favours, the objective flattens and that sample stops pushing.
How does DDPG's deterministic actor get its gradient from the critic?
basics
~20 sDDPG's critic scores a state-action pair and is differentiable in the action, so the actor is trained by pushing its output uphill on that score: the gradient is dQ/da evaluated at the actor's own action, chained with the actor's parameter gradient.
In knowledge distillation, why train the student on the teacher's full probability vector instead of the one-hot label?
basics
~20 sThe teacher's wrong-class probabilities encode which classes resemble each other, structure a one-hot label throws away. Each example then supplies a whole similarity-ranked distribution rather than a single index, giving the student a richer and lower-variance training signal.
Why can a network with 5x fewer FLOPs still be slower than its rival on a mobile CPU?
basics
~20 sFLOPs counts arithmetic only. Latency also pays for moving data, for per-layer fixed costs, and for how well each operation uses the device's parallel units, so a low-FLOP network built from poorly-supported operations can lose on the stopwatch.
How does factorizing a 4096x4096 dense layer into two rank-256 matrices cut parameters?
basics
~20 sOne 4096x4096 matrix holds about 16.8M weights; replacing it with a 4096x256 matrix times a 256x4096 matrix holds about 2.1M, an 8x cut. The saving exists only while the rank stays below half the layer width.
How do you measure per-layer sensitivity before pruning a trained network?
basics
~20 sPrune one layer at a time to a fixed ratio, leave every other layer dense, and record the held-out drop. Repeating over several ratios gives a per-layer sensitivity curve that ranks which layers absorb cuts.
In unstructured magnitude pruning, which weights are zeroed and why must the model be retrained?
basics
~20 sUnstructured magnitude pruning zeroes the individual weights with the smallest absolute values, wherever they sit in the tensor, using |w| as a cheap proxy for importance. Retraining with those weights held at zero lets the survivors re-fit what was lost.
Why does a GCN normalize the neighbour sum by node degree instead of summing raw feature vectors?
basics
~20 sA raw neighbour sum scales with how many neighbours a node has, so hubs produce huge activations and ordinary nodes tiny ones. Dividing each edge's contribution by sqrt(d_i * d_j) keeps every node's aggregate on a comparable scale.
How is a graph encoded as the input to a graph neural network?
basics
~20 sA graph reaches the model as a node feature matrix — one row of features per node — plus an edge list of connected node id pairs, and optionally an edge feature matrix. Node ids are addresses, not values.
In graph neural networks, what separates transductive from inductive training?
basics
~20 sTransductive training assumes every node you will ever score was already in the graph when the model was fitted. Inductive training learns a function of node features and neighbourhoods, so a node added later can be scored without refitting.
What are the three steps of one message-passing round in a graph neural network?
basics
~20 sOne round has three steps: each neighbour builds a message from its state and the connecting edge; the node pools those messages with an order-free aggregator; a learned update then combines the pooled vector with the node's own previous state.
What is the difference between permutation invariance and equivariance for a graph model?
basics
~20 sPermutation invariance means relabelling the nodes leaves the output unchanged, which is what a single graph-level prediction requires. Equivariance means the output moves with the relabelling: per-node predictions keep their values but come back in the new node order.