skip to content

Neurons and Layers

A neuron is a weighted sum plus a bias fed through a nonlinearity, and a fully-connected layer does that for many neurons as one matrix multiply. Interviewers probe the shape arithmetic first.

on this pageshow

explore

questions

12

What does a single artificial neuron compute from its inputs?

level: juniorimportance: must knowfreq 85%

answer

  1. one vector in, one number out
  2. two stages, not one
  3. a dot product, then a shift
  4. the shift is a learned threshold
  5. n inputs, n plus one learned numbers

basics

~20 s

An artificial neuron multiplies each input by a learned weight, sums the products, adds one learned bias, and passes that single number through a fixed nonlinear function. Only the weights and the bias are learned.

solid answer

~50 s

A neuron takes an input vector `x`, holds one weight per input plus a single bias, and computes the pre-activation `z = w1*x1 + ... + wn*xn + b` — a dot product plus a shift. That scalar goes through a fixed nonlinear function `f`, giving the unit's output `a = f(z)`. So a unit with `n` inputs owns exactly `n + 1` learned numbers, and `f` itself has nothing to learn. The weights say how strongly and in which direction each input pushes the unit; the bias is the learned threshold — it decides how large the weighted evidence must be before the unit responds. Without the bias the unit is pinned to respond around zero weighted input, which is rarely where the data sits. Everything else in a feed-forward network is this same unit repeated and stacked.

go deeper

for a junior

Be ready to write one unit's arithmetic on a whiteboard for a four-dimensional input, name the weights, the single bias and the activation, and say which of those three training changes.

for a middle

Expect to explain why the affine part and the nonlinearity are separate stages, why there is exactly one bias per unit, and how input scaling changes the magnitude of the weights you end up with.

for a senior

Show that you read the bias as a learned decision threshold when debugging a unit that never fires or always fires, and that you check input scaling before blaming the architecture.

for a principal

Own the vocabulary your team uses with non-ML stakeholders. Framing a unit as weighted evidence against a learned threshold is what makes model behaviour explainable and auditable rather than mystical.

## The unit An artificial neuron is the smallest trainable piece of a neural network. It takes a fixed-length vector of numbers as input and returns one number. It does this in two stages. **Stage 1 — the affine part.** The unit stores one *weight* per input, `w1 ... wn`, plus one extra number called the *bias*, `b`. It computes ``` z = w1*x1 + w2*x2 + ... + wn*xn + b ``` `z` is called the *pre-activation* or the *logit-like score* of the unit. Written compactly, `z = w · x + b` — a dot product between the weight vector and the input vector, then a shift. This map is *affine*, not merely linear: a linear map must send the all-zero input to zero, and the bias is precisely what removes that restriction. **Stage 2 — the nonlinearity.** The scalar `z` is pushed through a fixed function `f`, called the activation, and the unit outputs `a = f(z)`. Historically `f` was a hard threshold — output 1 when `z > 0`, otherwise 0 — which is the classical perceptron unit. Modern networks use smooth or piecewise-linear activations instead; *which* function to use is its own subject, and the important point here is structural: `f` is chosen by the engineer, has no trainable numbers of its own in the plain case, and is applied to the single scalar `z`, not to each input separately. ## What is learned and what is not A unit with `n` inputs has exactly `n + 1` learned parameters: `n` weights and one bias. Training changes those `n + 1` numbers and nothing else about the unit. Two consequences fall out of that immediately: - The *shape* of the unit's response is fixed by `f`; learning only moves where and how steeply it happens. - The unit is symmetric in a useless way if you feed it identical inputs — two inputs that always carry the same value can never be told apart by the unit, because only their weight *sum* is identifiable. ## Reading the weights and the bias The weight `wi` says how strongly input `i` pushes the unit, and its sign says in which direction. Doubling `wi` doubles that input's contribution to `z`. Because the weights only ever meet the inputs through a sum of products, the units and scale of each input matter: an input measured in thousands and an input measured in fractions will need weights of wildly different magnitude to contribute comparably, which is why input scaling changes what the weights look like. The bias is best read as a *learned firing threshold*. Consider one unit watching a single smoke-detector reading `s`, computing `z = w*s + b` with a threshold activation. The unit fires when `w*s + b > 0`, that is, when `s > -b/w` for positive `w`. The engineer never writes that cut-off down; training picks `b` so that the cut-off lands where the labelled data says it should. Remove the bias and the cut-off is forced to `s > 0`, meaning the unit fires for essentially any positive reading — a much poorer detector, and one that no amount of weight tuning can fix, because scaling `w` cannot move a threshold that is stuck at the origin. ## From one unit to a layer A fully-connected layer is simply many such units reading the same input vector, each with its own weight vector and its own bias, all applying the same activation. The layer's output is the vector of those units' scalars. Nothing about the individual unit changes; only the bookkeeping does. ## A worked scalar With weights `[0.5, -1.2, 0.3, 2.0]`, bias `-0.4` and input `[1.0, 0.0, -2.0, 0.5]`: ``` z = 0.5*1.0 + (-1.2)*0.0 + 0.3*(-2.0) + 2.0*0.5 - 0.4 = 0.5 + 0.0 - 0.6 + 1.0 - 0.4 = 0.5 ``` With a threshold activation the unit outputs 1. Drop the bias and `z` becomes `0.9` — the same qualitative answer here, but a unit whose threshold can no longer be placed anywhere except at zero weighted input. ## Common confusions - The pre-activation `z` is not the unit's output; the output is `f(z)`. - There is one bias *per unit*, not one per input. - The nonlinearity is applied after the summation, not to each product. - A unit with no nonlinearity is not a neural network building block in any interesting sense — it is a linear model wearing a different name.

  • What does a neuron lose if you remove its bias term?
    It loses the ability to place its threshold anywhere but at zero weighted input. With bias, the unit fires when `w · x > -b`, so training can put the cut-off wherever the data wants it; without it, the cut-off is pinned to `w · x > 0` and rescaling the weights cannot move it. That is a real restriction on the family of functions the unit can express, bought for one parameter per unit.
  • How does a single neuron relate to a classical linear model?
    Directly. Strip the nonlinearity and the unit computes `w · x + b` — that is linear regression's prediction. Keep a squashing activation that maps the score into (0, 1) and you get logistic regression's predicted probability. A neural network's novelty is not the unit; it is stacking many of them with nonlinearities in between so that later units read learned features rather than raw inputs.

Think of a panel of noisy advisors. Each advisor's opinion is scaled by how much you trust them, the opinions are added up, and you act only once the total clears a bar you have learned from experience — that bar is the bias.

saying these in an interview costs you the question

  • Says the neuron just adds its inputs, with no weights
  • Claims each input carries its own bias term
  • Calls the pre-activation the unit's output
  • Thinks the activation function has learned parameters here
  • Applies the nonlinearity to each product before summing

context

open as a page

In a 12288-512-128-10 MLP fed a batch of 32 flattened aerial tiles, what shape is each activation?

level: juniorimportance: must knowfreq 78%

basics

~10 s

Input (32, 12288), then (32, 512), then (32, 128), and finally (32, 10). The batch size stays on the leading axis at every layer; only the trailing feature width changes, once per weight matrix.

open as a page

How do you count the parameters of a 784-128-10 fully-connected neural network?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Each fully-connected layer holds inputs times outputs weights plus one bias per output unit. For 784-128-10 that is 784 x 128 + 128 = 100,480 and 128 x 10 + 10 = 1,290, totalling 101,770 parameters.

open as a page

Why does a stack of fully-connected layers with no nonlinearity collapse to one layer?

level: middleimportance: must knowfreq 72%

basics

~20 s

Composing affine maps yields another affine map: W2(W1x + b1) + b2 equals (W2W1)x + (W2b1 + b2). Depth with no nonlinearity between the layers buys more parameters and a different optimisation path, never a larger family of functions.

open as a page

Why is an embedding lookup equivalent to multiplying a one-hot vector by a weight matrix?

level: middleimportance: must knowfreq 70%

basics

~20 s

An embedding table is a weight matrix with one row per id. A one-hot vector times that matrix selects exactly that row and zeroes the rest, so the lookup is the same linear layer computed by indexing instead of multiplying.

open as a page

An embedding lookup meets an id it never saw in training -- what happens at serving time?

level: seniorimportance: must knowfreq 58%

basics

~10 s

There is no row for it, so the lookup fails on an out-of-range index unless the vocabulary reserves an unknown row. That row only helps if rare ids were routed to it during training.

open as a page

How do two hidden units let a network express XOR over two binary flags?

level: middleimportance: should knowfreq 56%

basics

~20 s

Each hidden unit thresholds the same two flags at a different level: one fires when at least one flag is on, the other only when both are. The output unit fires when the first fires and the second does not.

open as a page

A layer's (8, 256) output adds a bias stored as (8, 1) instead of (256,) - what goes wrong?

level: middleimportance: should knowfreq 46%

basics

~20 s

Broadcasting aligns from the right, so a (256,) bias adds one value per unit to every row, while an (8, 1) bias adds one value per row across all units. Both run; only the first is a real bias.

open as a page

A 40-1024-1024-250 MLP holds 1.35M parameters; which single width do you cut to halve it?

level: middleimportance: should knowfreq 41%

basics

~10 s

Cut the second hidden width. The 1024-to-1024 matrix holds about 1.05M of the 1.35M parameters, so taking that layer to 512 units drops the total to roughly 695,000. The 40-input layer holds almost nothing.

open as a page

How do you choose the embedding width for a 2-million-item catalog id that dominates the model?

level: principalimportance: should knowfreq 44%

basics

~10 s

Size it from observations per id and a stated memory budget, not from cardinality. Two million rows at width 64 is 128 million parameters, so width is the model's main size decision.

open as a page

A stakeholder asks what each hidden unit of your accelerometer model means — how do you answer?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Treat a name for a hidden unit as a hypothesis and test it: check the unit against labelled held-out conditions, and measure what ablating it costs. Hidden features are learned and distributed, not guaranteed to match human concepts.

open as a page

A transposed input batch trains without raising a shape error - how do you detect it?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

A product only checks the feature axis, so a transposed batch can still conform and run. Confirm the layout instead: vary the batch size and see which axis moves, and perturb one example to check only its own output changes.

open as a page