A stakeholder asks what each hidden unit of your accelerometer model means — how do you answer?
answer
- a name is a hypothesis, not a finding
- meaning may be spread across units
- the index carries no information
- probe on held-out labelled conditions
- supervise the feature if it must be guaranteed
basics
~20 sTreat a name for a hidden unit as a hypothesis and test it: check the unit against labelled held-out conditions, and measure what ablating it costs. Hidden features are learned and distributed, not guaranteed to match human concepts.
solid answer
~40 sI would answer with evidence, not with a label. Hidden units are learned intermediate features, and nothing in training forces one unit to match one human concept — features are often spread across several units, and one unit may respond to several unrelated conditions. So I form the hypothesis "unit 3 tracks device-upright", then test it: hold out labelled upright, shaking and at-rest segments and check whether the unit's activation separates them; zero the unit and measure the drop in task performance; and for a first hidden layer, read its weight vector directly, since that is a filter over the three raw axes. I would report a supported hypothesis with its counter-examples, and note that retraining gives a different but equally good set of units.
go deeper
Be ready to say hidden units are features the network learned for itself, not columns you designed, and that their names are guesses until someone checks them against labelled data.
Expect to explain why a first-layer weight vector is readable as a direction over raw inputs while a deeper unit's weights are not, and why a unit's index carries no meaning.
Demonstrate the method: probe on held-out labelled conditions, ablate and measure the task cost, hunt for counter-examples, and report a supported hypothesis with its limits rather than a confident label.
Own the boundary between explanation and guarantee. If a downstream team or a regulator needs a named signal, commit to supervising or engineering that feature explicitly instead of letting an interpretability story harden into a contract nobody can honour.
## What is actually being asked A small network reads raw three-axis accelerometer readings and predicts an activity label. A stakeholder points at the hidden layer and asks what each unit means. The honest answer is a method, not a list of names — and delivering the method well is the whole of this question. ## Why hidden units resist naming Hidden units *are* learned features: each computes a weighted sum of what the previous layer produced, plus a bias, through a nonlinearity, and training shapes those weights only to make the final loss small. Nothing in that objective rewards a unit for being individually meaningful. Three consequences follow. **Features are distributed.** The information the output layer needs about, say, orientation may be carried by a *combination* of several units rather than by one. There is no rule that says a concept must be concentrated in a single coordinate of the hidden vector. **Units can be polysemantic.** One unit may fire on two conditions a human would call unrelated, because the output layer can disambiguate them using other units. It is the layer's vector that carries meaning, not necessarily each of its components. **Unit identity is not stable.** Two training runs on the same data will land on different but equally good hidden layers, and the unit that looked like an orientation detector in run A may have no counterpart at the same index in run B. Unit *indices* certainly carry no meaning; they are bookkeeping. ## The method that does work Treat every proposed name as a falsifiable hypothesis about the unit's behaviour, and test it. **1. Behavioural probing on held-out labelled data.** Collect segments you have labelled independently — device flat and still, device upright and still, device shaking — and look at the unit's activation distribution on each. A unit that genuinely tracks "upright" should separate those groups on held-out recordings, not just on the ones that inspired the name. Report the overlap, not only the means. **2. Ablation.** Force the unit's output to a constant (its mean, or zero) and re-evaluate the model. If knocking it out barely moves the task metric, the unit is not carrying the concept you are attributing to it, whatever its activations look like. If it costs a great deal, the unit matters — though that still does not prove it means what you said. **3. Read the first-layer weights.** This is the one place where the interpretation is directly available. A first-layer unit computes a weighted sum of the *raw* axes, so its weight vector is a readable direction: weights near `[0, 0, 1]` with a bias placing the threshold around the resting gravity reading really is an orientation-style detector, and you can say so with confidence. Deeper units do not have this property — their inputs are already learned features, so their weights are only interpretable relative to those. **4. Find the counter-examples.** Search the held-out data for inputs where the unit fires strongly and the proposed concept is absent. If they are easy to find, downgrade the claim. ## What not to say Do not read meaning from a unit's index. Do not treat a large weight as proof of importance — magnitude depends on the scale of the input feeding it, and an input measured in different units would give a different weight for identical behaviour. Do not present a single striking activation example as evidence; cherry-picked examples are available for almost any hypothesis. And do not promise stability across retrains. ## When the business genuinely needs the feature If a downstream team needs a reliable "device is shaking" signal, stop hoping a hidden unit provides it and make it explicit. Two clean options: add an auxiliary output trained against labelled shaking/not-shaking targets, so the meaning is enforced by the objective rather than discovered by accident; or compute the feature directly from the raw signal and feed it in as an input. Both give a contract the stakeholder can rely on. An unsupervised hidden unit never gives you that contract, no matter how convincing its activation plot looks. ## How to close the answer The strong version of this answer has three beats: name the epistemic status (hypothesis, not fact), name the tests you would run and what each can and cannot establish, and name the alternative if a guarantee is required. That combination is what distinguishes an engineer who has actually had this conversation with a stakeholder from one who has read about feature visualisation.
- Two training runs on the same data give completely different hidden units. Is that a bug?No, it is expected. Many parameter settings achieve near-identical loss, and hidden units are exchangeable — the layer's coordinate system is arbitrary, so permuted, rescaled and genuinely different feature sets can all be equally good. Judge the model by its behaviour on held-out data, not by whether unit 3 looks the same twice. Instability of unit identity is only a problem if you built a downstream contract on it.
- Why is a first hidden layer's weight vector more interpretable than a deeper unit's?Because its inputs are the raw measured axes, so the weight vector is a direction in a space the stakeholder already understands — a filter over the three accelerometer channels. A deeper unit weights *learned* features whose own meaning is uncertain, so its weights are only interpretable relative to that uncertain basis, and the interpretation compounds errors at each level.
- The stakeholder needs a guaranteed 'device is shaking' signal. What do you do?Make it explicit rather than hoping a hidden unit supplies it. Either add an auxiliary output trained against labelled shaking targets, so the objective enforces the meaning, or compute the feature from the raw signal and feed it in as an input. Both yield a signal with a definition, a test set and a stated error rate — something a hidden unit's activation plot can never provide.
saying these in an interview costs you the question
- Assumes each hidden unit is one human-named concept
- Reads meaning from a unit's position or index
- Treats a large weight as proof the unit matters
- Offers one striking activation example as evidence
- Promises hidden features are stable across retraining