skip to content

Neural Network Basics

You will learn what a network computes: a layer as a matrix multiply plus a nonlinearity, what depth buys over width, and which output head fits a task. Interviewers open here before any framework.

on this pageshow

explore

questions

29

What does a single artificial neuron compute from its inputs?

level: juniorimportance: must knowfreq 85%

answer

  1. one vector in, one number out
  2. two stages, not one
  3. a dot product, then a shift
  4. the shift is a learned threshold
  5. n inputs, n plus one learned numbers

basics

~20 s

An artificial neuron multiplies each input by a learned weight, sums the products, adds one learned bias, and passes that single number through a fixed nonlinear function. Only the weights and the bias are learned.

solid answer

~50 s

A neuron takes an input vector `x`, holds one weight per input plus a single bias, and computes the pre-activation `z = w1*x1 + ... + wn*xn + b` — a dot product plus a shift. That scalar goes through a fixed nonlinear function `f`, giving the unit's output `a = f(z)`. So a unit with `n` inputs owns exactly `n + 1` learned numbers, and `f` itself has nothing to learn. The weights say how strongly and in which direction each input pushes the unit; the bias is the learned threshold — it decides how large the weighted evidence must be before the unit responds. Without the bias the unit is pinned to respond around zero weighted input, which is rarely where the data sits. Everything else in a feed-forward network is this same unit repeated and stacked.

go deeper

for a junior

Be ready to write one unit's arithmetic on a whiteboard for a four-dimensional input, name the weights, the single bias and the activation, and say which of those three training changes.

for a middle

Expect to explain why the affine part and the nonlinearity are separate stages, why there is exactly one bias per unit, and how input scaling changes the magnitude of the weights you end up with.

for a senior

Show that you read the bias as a learned decision threshold when debugging a unit that never fires or always fires, and that you check input scaling before blaming the architecture.

for a principal

Own the vocabulary your team uses with non-ML stakeholders. Framing a unit as weighted evidence against a learned threshold is what makes model behaviour explainable and auditable rather than mystical.

## The unit An artificial neuron is the smallest trainable piece of a neural network. It takes a fixed-length vector of numbers as input and returns one number. It does this in two stages. **Stage 1 — the affine part.** The unit stores one *weight* per input, `w1 ... wn`, plus one extra number called the *bias*, `b`. It computes ``` z = w1*x1 + w2*x2 + ... + wn*xn + b ``` `z` is called the *pre-activation* or the *logit-like score* of the unit. Written compactly, `z = w · x + b` — a dot product between the weight vector and the input vector, then a shift. This map is *affine*, not merely linear: a linear map must send the all-zero input to zero, and the bias is precisely what removes that restriction. **Stage 2 — the nonlinearity.** The scalar `z` is pushed through a fixed function `f`, called the activation, and the unit outputs `a = f(z)`. Historically `f` was a hard threshold — output 1 when `z > 0`, otherwise 0 — which is the classical perceptron unit. Modern networks use smooth or piecewise-linear activations instead; *which* function to use is its own subject, and the important point here is structural: `f` is chosen by the engineer, has no trainable numbers of its own in the plain case, and is applied to the single scalar `z`, not to each input separately. ## What is learned and what is not A unit with `n` inputs has exactly `n + 1` learned parameters: `n` weights and one bias. Training changes those `n + 1` numbers and nothing else about the unit. Two consequences fall out of that immediately: - The *shape* of the unit's response is fixed by `f`; learning only moves where and how steeply it happens. - The unit is symmetric in a useless way if you feed it identical inputs — two inputs that always carry the same value can never be told apart by the unit, because only their weight *sum* is identifiable. ## Reading the weights and the bias The weight `wi` says how strongly input `i` pushes the unit, and its sign says in which direction. Doubling `wi` doubles that input's contribution to `z`. Because the weights only ever meet the inputs through a sum of products, the units and scale of each input matter: an input measured in thousands and an input measured in fractions will need weights of wildly different magnitude to contribute comparably, which is why input scaling changes what the weights look like. The bias is best read as a *learned firing threshold*. Consider one unit watching a single smoke-detector reading `s`, computing `z = w*s + b` with a threshold activation. The unit fires when `w*s + b > 0`, that is, when `s > -b/w` for positive `w`. The engineer never writes that cut-off down; training picks `b` so that the cut-off lands where the labelled data says it should. Remove the bias and the cut-off is forced to `s > 0`, meaning the unit fires for essentially any positive reading — a much poorer detector, and one that no amount of weight tuning can fix, because scaling `w` cannot move a threshold that is stuck at the origin. ## From one unit to a layer A fully-connected layer is simply many such units reading the same input vector, each with its own weight vector and its own bias, all applying the same activation. The layer's output is the vector of those units' scalars. Nothing about the individual unit changes; only the bookkeeping does. ## A worked scalar With weights `[0.5, -1.2, 0.3, 2.0]`, bias `-0.4` and input `[1.0, 0.0, -2.0, 0.5]`: ``` z = 0.5*1.0 + (-1.2)*0.0 + 0.3*(-2.0) + 2.0*0.5 - 0.4 = 0.5 + 0.0 - 0.6 + 1.0 - 0.4 = 0.5 ``` With a threshold activation the unit outputs 1. Drop the bias and `z` becomes `0.9` — the same qualitative answer here, but a unit whose threshold can no longer be placed anywhere except at zero weighted input. ## Common confusions - The pre-activation `z` is not the unit's output; the output is `f(z)`. - There is one bias *per unit*, not one per input. - The nonlinearity is applied after the summation, not to each product. - A unit with no nonlinearity is not a neural network building block in any interesting sense — it is a linear model wearing a different name.

  • What does a neuron lose if you remove its bias term?
    It loses the ability to place its threshold anywhere but at zero weighted input. With bias, the unit fires when `w · x > -b`, so training can put the cut-off wherever the data wants it; without it, the cut-off is pinned to `w · x > 0` and rescaling the weights cannot move it. That is a real restriction on the family of functions the unit can express, bought for one parameter per unit.
  • How does a single neuron relate to a classical linear model?
    Directly. Strip the nonlinearity and the unit computes `w · x + b` — that is linear regression's prediction. Keep a squashing activation that maps the score into (0, 1) and you get logistic regression's predicted probability. A neural network's novelty is not the unit; it is stacking many of them with nonlinearities in between so that later units read learned features rather than raw inputs.

Think of a panel of noisy advisors. Each advisor's opinion is scaled by how much you trust them, the opinions are added up, and you act only once the total clears a bar you have learned from experience — that bar is the bias.

saying these in an interview costs you the question

  • Says the neuron just adds its inputs, with no weights
  • Claims each input carries its own bias term
  • Calls the pre-activation the unit's output
  • Thinks the activation function has learned parameters here
  • Applies the nonlinearity to each product before summing

context

open as a page

In a 12288-512-128-10 MLP fed a batch of 32 flattened aerial tiles, what shape is each activation?

level: juniorimportance: must knowfreq 78%

basics

~10 s

Input (32, 12288), then (32, 512), then (32, 128), and finally (32, 10). The batch size stays on the leading axis at every layer; only the trailing feature width changes, once per weight matrix.

open as a page

How do you count the parameters of a 784-128-10 fully-connected neural network?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Each fully-connected layer holds inputs times outputs weights plus one bias per output unit. For 784-128-10 that is 784 x 128 + 128 = 100,480 and 128 x 10 + 10 = 1,290, totalling 101,770 parameters.

open as a page

Why use one sigmoid per label instead of a softmax over the same output layer?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Per-label sigmoids fit tasks where one example can carry several labels at once: each output is its own independent probability and the outputs need not sum to one. A softmax makes the classes share a single probability budget, which only fits mutually exclusive classes.

open as a page

Why must a softmax classification head have exactly one output unit per class?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A softmax head emits one score per class and normalizes those scores into probabilities that sum to one. K classes need K units: with fewer, some class can never be predicted; extra units create classes no label ever selects.

open as a page

With 50,000 training examples and batch size 128, how many iterations and updates do 3 epochs take?

level: juniorimportance: must knowfreq 78%

basics

~20 s

An epoch is one full pass over the data. At batch size 128, 50,000 examples give 390 full batches plus a ragged batch of 80, so 391 iterations per epoch, one update each, and 1,173 updates over 3 epochs.

open as a page

Why does a stack of fully-connected layers with no nonlinearity collapse to one layer?

level: middleimportance: must knowfreq 72%

basics

~20 s

Composing affine maps yields another affine map: W2(W1x + b1) + b2 equals (W2W1)x + (W2b1 + b2). Depth with no nonlinearity between the layers buys more parameters and a different optimisation path, never a larger family of functions.

open as a page

Why can a deep narrow network match a shallow wide one using far fewer parameters?

level: middleimportance: must knowfreq 62%

basics

~20 s

Depth composes. Each layer reuses every feature the layer below it built, so complexity compounds with depth, while a shallow net must build each feature independently from the raw input. For some functions that costs exponentially more units.

open as a page

Why is an embedding lookup equivalent to multiplying a one-hot vector by a weight matrix?

level: middleimportance: must knowfreq 70%

basics

~20 s

An embedding table is a weight matrix with one row per id. A one-hot vector times that matrix selects exactly that row and zeroes the rest, so the lookup is the same linear layer computed by indexing instead of multiplying.

open as a page

Narrate one training step for a batch of 64 — which tensors change and which do not?

level: middleimportance: must knowfreq 68%

basics

~20 s

The forward pass turns the 64 inputs into activations and one loss number, the backward pass fills a gradient for every parameter, and the update writes new parameter values. Only parameters, optimizer state and gradient buffers change; the batch, targets and architecture do not.

open as a page

What does the universal approximation theorem guarantee about a one-hidden-layer network?

level: middleimportance: must knowfreq 62%

basics

~20 s

It is an existence result: for any continuous target on a closed bounded region and any error tolerance, some one-hidden-layer network with a non-polynomial activation stays inside that tolerance. The width needed is unbounded and unspecified.

open as a page

An embedding lookup meets an id it never saw in training -- what happens at serving time?

level: seniorimportance: must knowfreq 58%

basics

~10 s

There is no row for it, so the lookup fails on an out-of-range index unless the vocabulary reserves an unknown row. That row only helps if rare ids were routed to it during training.

open as a page

How do two hidden units let a network express XOR over two binary flags?

level: middleimportance: should knowfreq 56%

basics

~20 s

Each hidden unit thresholds the same two flags at a different level: one fires when at least one flag is on, the other only when both are. The output unit fires when the first fires and the second does not.

open as a page

A layer's (8, 256) output adds a bias stored as (8, 1) instead of (256,) - what goes wrong?

level: middleimportance: should knowfreq 46%

basics

~20 s

Broadcasting aligns from the right, so a (256,) bias adds one value per unit to every row, while an (8, 1) bias adds one value per row across all units. Both run; only the first is a real bias.

open as a page

A 40-1024-1024-250 MLP holds 1.35M parameters; which single width do you cut to halve it?

level: middleimportance: should knowfreq 41%

basics

~10 s

Cut the second hidden width. The 1024-to-1024 matrix holds about 1.05M of the 1.35M parameters, so taking that layer to 512 units drops the total to roughly 695,000. The 40-input layer holds almost nothing.

open as a page

Your resale-price regression head predicts -30 for cheap items — how do you fix it?

level: middleimportance: should knowfreq 55%

basics

~20 s

A plain linear output head is unbounded by construction, so negative predictions are expected, not a bug. Fix it in the head: predict in log space and exponentiate, pass the output through a positive-valued transform, or accept the negatives and clip only at serving.

open as a page

Can you skip the softmax at inference and take the argmax of a classifier's raw scores?

level: middleimportance: should knowfreq 48%

basics

~20 s

Yes, for a top-1 or top-k label. Softmax exponentiates each score and divides them all by the same positive total, so it never changes their order. You need the normalized values only when something downstream consumes the number itself.

open as a page

Why can a network with far more parameters than training examples still generalize?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Parameter count bounds what a network could fit, not what training selects. Gradient descent picks a biased, simple subset of the functions that fit, and the architecture's priors narrow it further. Only held-out data settles it.

open as a page

How do you choose decision thresholds for a 50-tag multi-label audio tagging head?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Tune one threshold per tag on a held-out set rather than applying 0.5 everywhere. Tags differ in base rate, score distribution and the cost of a mistake, so 'live recording' may fire best at 0.2 while 'acoustic guitar' needs 0.7.

open as a page

A deployed 20-way product classifier needs a 21st category. What happens to the softmax head?

level: seniorimportance: should knowfreq 58%

basics

~20 s

The head gains one weight row and one bias for the new class; the trunk keeps its shape. All 21 classes then share one probability budget, so old outputs shift and old thresholds no longer transfer.

open as a page

What breaks when a validation pass is run in training mode instead of evaluation mode?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Dropout stays active, so every validation number is a noisy sample of a randomly thinned network, and normalization layers use the validation batch's own statistics and refresh their stored ones. The metric becomes irreproducible, batch-order dependent, and validation data leaks into the model.

open as a page

Why does universal approximation say nothing about a network's output far outside its fitted range?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The guarantee is stated on a closed bounded region chosen before the weights are picked, and closeness is claimed only on that region. Outside it the network is unconstrained and simply follows whatever its architecture does asymptotically.

open as a page

How do you choose the embedding width for a 2-million-item catalog id that dominates the model?

level: principalimportance: should knowfreq 44%

basics

~10 s

Size it from observations per id and a stated memory budget, not from cardinality. Two million rows at width 64 is 128 million parameters, so width is the model's main size decision.

open as a page

Universal approximation says one hidden layer suffices — so how do you answer a claim that depth is just fashion?

level: principalimportance: should knowfreq 38%

basics

~20 s

The theorem proves suitable weights exist; it never says any procedure finds them, bounds how many units they need, or says how much data pins them down. Density is a weak property that recommends no architecture on its own.

open as a page

A stakeholder asks what each hidden unit of your accelerometer model means — how do you answer?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Treat a name for a hidden unit as a hypothesis and test it: check the unit against labelled held-out conditions, and measure what ablating it costs. Hidden features are learned and distributed, not guaranteed to match human concepts.

open as a page

A transposed input batch trains without raising a shape error - how do you detect it?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

A product only checks the feature axis, so a transposed batch can still conform and run. Confirm the layout instead: vary the batch size and see which axis moves, and perturb one example to check only its own output changes.

open as a page

Given a fixed parameter budget for a feedforward net, how do you decide depth versus width?

level: principalimportance: nice to knowfreq 34%

basics

~10 s

Let the binding constraint decide. Depth buys parameter efficiency but adds sequential latency and harder optimization; width parallelizes well but costs full fan-in and fan-out per unit. Settle it with small matched-budget runs.

open as a page

Would you serve one shared-trunk model with steering and brake heads, or two separate models?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Share a trunk when the two tasks need the same perception features and you want one forward pass; keep them separate when they need independent release cadences. Sharing buys compute and transfer, and costs you a coupled artefact where retraining for one task can regress the other.

open as a page

Your model sees a 2-billion-token stream once — how do you plan and report the run without epochs?

level: principalimportance: nice to knowfreq 38%

basics

~20 s

Count the run in optimizer steps and tokens consumed, not epochs — an epoch that happens once carries no progress information. Fix a step budget up front, hang evaluation and checkpointing on step intervals, and compare runs at equal tokens seen.

open as a page