Your resale-price regression head predicts -30 for cheap items — how do you fix it?
answer
- A linear head has no activation
- Unbounded output spans the whole real line
- Log targets, or a positive-valued transform
- Sigmoid-scaled head for a genuine finite range
- Only impose bounds the domain guarantees
basics
~20 sA plain linear output head is unbounded by construction, so negative predictions are expected, not a bug. Fix it in the head: predict in log space and exponentiate, pass the output through a positive-valued transform, or accept the negatives and clip only at serving.
solid answer
~50 sA regression head is just a linear layer with no squashing function, so its output ranges over the whole real line — nothing in the architecture knows prices are positive. Clipping at serving time is the cheapest patch, but it leaves training pushing towards impossible values and it silently distorts the low end. Better options change the head itself: train on log-transformed targets and exponentiate at inference, which guarantees positivity and turns the objective into a relative-error one that suits prices spanning several orders of magnitude; or apply a positive-valued transform such as softplus to the head's scalar so the output cannot go below zero. When the quantity has a genuine finite range — an occupancy fraction that must land in [0, 1] — a sigmoid-scaled head is the natural choice, since it enforces both ends. The tradeoff is that a bounded head can never express a value outside the assumed range, so I only impose bounds the domain actually guarantees.
go deeper
Recall that a regression head is a linear layer with no activation, so its output can be any real number including negative ones. Know that a sigmoid on the output confines it to [0, 1].
Explain the menu and its mechanics: clipping, log-space targets with exponentiation, a positive-valued transform, and a scaled sigmoid for a genuine finite range. Say what each one does to the error geometry, not just to the sign.
Demonstrate judgment about which bounds are real. Show that you check whether a constraint is physical or merely observed, that you put the transform inside the served model, and that you know a bounded head hides violations as a flat line at the boundary.
Own the tradeoff between encoding domain constraints in the architecture and keeping the model able to surface data problems. Be ready to argue when a hard bound is a safety guarantee worth its rigidity and when it is a premature assumption.
## Why the negative appears A regression head is the simplest head there is: a linear layer of width 1 (or width d for a multi-output regression) applied to the trunk's features, with **no** activation on top. Its output is `w . h + b`, an affine function of the hidden vector, which ranges over all of the real line. There is no mechanism by which the architecture could know that resale prices are positive. If the trunk's features for a very cheap item extrapolate slightly past the cheapest examples the model saw, the linear head happily returns -30. A negative prediction is therefore a *modelling* omission, not a numerical failure. ## Four ways to respond **1. Leave it unbounded and clip at serving.** `max(pred, 0)`. Zero training cost, and it is the honest choice when negatives are rare, small, and confined to a region you do not care about. But the model is still being trained to produce impossible values, the gradient never learns that the region is forbidden, and everything clipped to exactly 0 becomes indistinguishable — you lose the ability to rank the cheapest items against each other. **2. Predict in log space.** Train on `log(y)` (or `log(1 + y)` if zeros occur) and exponentiate at inference. Positivity is guaranteed because `exp` cannot return a non-positive number. The deeper effect is on the error geometry: equal errors in log space are equal *ratio* errors in the original space, so being off by 10 currency units matters much more for a 20-unit item than for a 2000-unit one. For prices spanning orders of magnitude that is usually the behaviour you want, and it is the standard choice. The costs are that the exponentiated prediction is no longer an unbiased estimate of the mean in the original units, and that a large log-space error becomes an enormous absolute error after exponentiating. **3. Apply a positive-valued output transform.** Put softplus, `log(1 + exp(z))`, on the head's scalar so the output is positive by construction while staying close to linear for large `z`. This keeps the objective in the original units — useful when the business metric is an absolute currency error — while removing the impossible region. It is a lighter-touch fix than the log transform, and it does not change the relative weighting of cheap and expensive items. **4. Bound both ends with a scaled sigmoid.** For a quantity with a genuine finite range — an occupancy fraction constrained to [0, 1], a probability-like ratio, a normalised angle — apply a sigmoid to the head's scalar, and scale/shift it if the range is `[lo, hi]` rather than `[0, 1]`: `pred = lo + (hi - lo) * sigmoid(z)`. Now the prediction is structurally incapable of leaving the range, and every downstream consumer can rely on that without validation. ## What a bounded head costs you The bound is a hard architectural claim. If the true range is wider than you assumed — occupancy that can exceed 1.0 because of double-counting, or a price ceiling you set from last year's catalogue — the model cannot express the truth at all, and the error shows up as a flat line at the boundary rather than as a visible outlier. A bounded head also compresses: to output values very near either endpoint, the pre-activation must be driven far from zero, so predictions crowd towards the interior and the extremes are approached slowly. Impose bounds only where the domain genuinely guarantees them, and prefer a one-sided transform when only one side is physically constrained. ## Choosing between them Ask two questions. *Is the bound physical or empirical?* A fraction that is a ratio of counts is physically in [0, 1] — bound it. A price whose observed maximum happens to be 5000 is not bounded above at all — do not bound it. *What error geometry does the product want?* If a 10% error is equally bad at every price point, work in log space. If a fixed currency error is what the business loses, keep the original units and use a one-sided positive transform. ## Related head hygiene Standardising targets (subtracting the mean, dividing by the standard deviation) is orthogonal to all of this: it stabilises the scale the head must learn, but a standardised target is still unbounded, so it does not prevent negatives after you invert the transform. Multi-output regression heads follow the same rules per output — a head predicting width and height needs positivity on both, and each output may want a different transform. And whatever transform you pick belongs *inside* the served model, not in a caller's post-processing, so every consumer inherits the guarantee.
- Why does training on log-transformed prices change more than just positivity?It changes what counts as a big error. Equal errors in log space are equal ratio errors in the original units, so a 10-unit miss on a 20-unit item is penalised far more than the same miss on a 2000-unit item. That is usually right for prices spanning orders of magnitude, but it means the exponentiated prediction is no longer an unbiased estimate of the mean price.
- When is clipping negatives at serving time actually acceptable?When negatives are rare, tiny, and confined to a region nobody acts on. It is a patch, not a fix: the model is still trained towards impossible values, and everything clipped lands on exactly the same number, destroying any ordering among the cheapest items. If the low end matters to the product, change the head instead.
- Would you bound a head predicting an occupancy fraction that occasionally exceeds 1.0 in the data?No. Values above 1.0 mean the bound is not real — probably double-counting or a definitional issue in how occupancy is computed. A sigmoid-scaled head would make those cases inexpressible and hide the data problem behind a flat line at the boundary. Fix the label definition first, then decide.
saying these in an interview costs you the question
- Calls a negative regression output a numerical bug
- Believes standardising targets prevents impossible predictions
- Bounds an output using a range read off the training maximum
- Clips negatives in caller code rather than inside the model
- Thinks log-space training changes only the output's sign