Which skew transform handles a right-skewed column that contains zeros and negatives?
answer
- start with the transform's domain
- log(0) has no value
- add one before taking the log
- Box-Cox needs strictly positive input
- the power family that allows negatives
basics
~20 sPlain log fails here: log(0) is undefined and negatives have no log. For a non-negative column with zeros, use log(1+x). When values drop below zero, use Yeo-Johnson, the power transform defined on the whole real line.
solid answer
~50 sThe first question is the transform's domain, not its strength. Natural log needs strictly positive input, so a zero-inflated lifetime-spend column breaks it; `log(1 + x)` is the standard fix, since it is defined at zero, maps zero to zero, and behaves like log once x is large. Square root also tolerates zero and squeezes more gently. Box-Cox is the parametric family whose exponent lambda is estimated from the data (lambda = 0 recovers the log, 0.5 the square root), but it too requires strictly positive values, so it suits something like insurance claim severity. For a net-profit column with genuine negatives, Yeo-Johnson is the right member: it mirrors the power transform on the negative side, so it is defined for every real number. Shifting the column to force positivity is a last resort, because the fitted model then depends on the constant you picked.
go deeper
Be ready to say which transforms are undefined at zero or for negative values, and to name log(1 + x) as the standard fix for a column full of zeros.
Explain the Box-Cox family, that lambda is estimated rather than chosen, that lambda = 0 is the log, and why Yeo-Johnson is the version that accepts negative values.
Show the pipeline discipline: a fitted lambda belongs to the training fold, zeros that encode a different state deserve their own flag, and a tree ensemble gains nothing from any of this.
Own the tradeoff between predictive shape and explainability. Argue when a fitted, opaque power transform is worth it versus a fixed log the business can read, and set the team convention for zero-inflated columns.
## The problem Many real numeric columns are right-skewed: most rows are small, a thin tail runs a long way out. Squared-error losses, distance-based methods and penalised linear models all react badly to that shape, because a handful of huge values dominate the fit. A monotone squeeze of the large values is the standard remedy. The trap is that the popular squeezes are not defined everywhere, and the column you actually have often contains zeros or negatives. ## The transforms and where each is defined **Natural log, `log(x)`** — defined only for `x > 0`. It compresses the tail hard: multiplying a value by 10 adds a constant to the transformed value, so the transform turns multiplicative spread into additive spread. Undefined at 0 and for negatives. **`log1p(x) = log(1 + x)`** — defined for `x > -1`, so it covers the whole of a non-negative column. It maps 0 to 0 exactly, which is why a zero-inflated column such as `lifetime spend` (a large point mass at zero, then a long right tail among spenders) is its classic use. For large x, `log(1 + x)` is almost `log(x)`, so you lose nothing at the tail end. Note it is `log(1 + x)`, not `log(x) + 1` — a surprisingly common misreading. **Square root, `sqrt(x)`** — defined for `x >= 0`, mild compared with the log. It is the traditional variance-stabilising choice for count data, where the variance grows roughly with the mean. **Box-Cox** — a one-parameter family for strictly positive x: ``` T(x) = (x^lambda - 1) / lambda for lambda != 0 T(x) = log(x) for lambda = 0 ``` lambda is not chosen by hand: it is estimated by maximum likelihood, picking the value that makes the transformed column look as close to normal as the family allows. lambda = 1 is essentially no transform, 0.5 is the square root, 0 is the log, -1 is the reciprocal. Because it needs `x > 0`, Box-Cox suits a strictly positive quantity such as insurance claim severity but not a column with zeros. **Yeo-Johnson** — the extension of Box-Cox to the whole real line. It applies a shifted power transform to values at or above zero and a mirrored power transform to negative values, so it is defined for every real input while keeping a single lambda estimated the same way. This is the transform for a `net profit` column that legitimately goes negative for loss-making accounts. ## The bookkeeping that people get wrong **Fitted versus fixed.** `log`, `log1p` and `sqrt` are fixed functions: no parameter is learned, so applying them to the whole dataset leaks nothing. Box-Cox and Yeo-Johnson learn a lambda from data. That lambda is a fitted parameter and must be estimated on the training fold only, then reused unchanged on validation, test and production rows. Estimating it on everything is quiet leakage. **Shifting to force positivity.** The tempting hack for zeros and negatives is `log(x + c)` for some constant c. It works arithmetically but the result depends entirely on c: a tiny c stretches the near-zero values into a huge artificial left tail, while a large c makes the transform nearly linear and pointless. If you must shift, justify c from the domain (a known measurement floor), not from what happens to look pretty. **Zeros that mean something else.** A point mass at zero is often a different state, not a small value — `never purchased` rather than `purchased a little`. Squeezing it into the same continuous scale hides that. The stronger design is often a pair of features: a binary `spent anything` flag plus the transform applied to the positive part. **Which models care.** Skew transforms of an input feature matter to linear and penalised linear models, distance-based methods and anything with a squared-error loss on that input. A tree ensemble splits on order alone, so any strictly increasing transform of a single feature leaves its achievable splits unchanged and buys nothing. Reaching for a power transform to fix a tree model's accuracy signals a missing mental model of how trees split. **Interpretability cost.** After a log-family transform, a linear coefficient is read per unit of the transformed scale — roughly a proportional effect rather than an absolute one. That is often an advantage for money-like quantities, but you have to say so out loud when you present coefficients. ## A decision order that holds up 1. What is the column's range? Strictly positive, non-negative with zeros, or genuinely signed? 2. Does the model care about shape at all? If it is a tree ensemble, stop. 3. Non-negative with zeros: try `log1p` first, `sqrt` if the squeeze is too aggressive. 4. Strictly positive and you want the data to pick the strength: Box-Cox, lambda fitted on train. 5. Signed values: Yeo-Johnson. 6. Zeros that encode a different state: split into a flag plus a transformed positive part.
- How is the lambda in Box-Cox or Yeo-Johnson chosen?By maximum likelihood: the value of lambda that makes the transformed column closest to normal under the family's likelihood. It is a fitted parameter, so estimate it on the training fold and reuse it everywhere else. Practitioners often round it to a nearby interpretable value such as 0 (log) or 0.5 (square root) so the feature stays explainable.
- Why is shifting a column by a constant to make it positive, then logging, a weak fix?Because the result depends on the constant you invented. A very small shift blows the near-zero rows out into a long artificial left tail; a large shift flattens the transform until it is nearly linear and achieves nothing. Two analysts picking different constants get different models from identical data, with no data-driven way to choose between them.
- Would you apply a skew transform to features feeding a gradient-boosted tree ensemble?No, not for accuracy. Trees split on order, so any strictly increasing transform of a single feature leaves the reachable splits unchanged. The only reasons are downstream: making the feature readable on a plot, or feeding the same prepared column to a linear model as well.
saying these in an interview costs you the question
- Adds an arbitrary tiny epsilon before logging without noticing the inflated left tail
- Claims Box-Cox works on any real values including negatives
- Reads log1p as log(x) + 1 rather than log(1 + x)
- Fits the Box-Cox lambda on training and test data together
- Expects a power transform to improve a tree ensemble's accuracy
- Treats a large point mass at zero as just another small value