Why does taking logs of right-skewed values like house prices reduce the skewness?
answer
- the transform is concave
- equal ratios become equal distances
- large values get squeezed hardest
- the long right tail is pulled inward
- undefined at zero, so positives only
basics
~20 sThe logarithm compresses large values far more than small ones, so a long right tail is pulled in toward the body of the data. Applied to strictly positive right-skewed values such as house prices, it can bring skewness close to zero.
solid answer
~50 sA logarithm is a strongly concave function: the gap between 100 and 1000 becomes the same distance on the log scale as the gap between 10 and 100. That squeezes the far right tail much harder than the dense body of small values, which is exactly the asymmetry that produces positive skewness in the first place. On house prices or session durations — bounded below by zero, dense at the low end, with a thin tail of very large values — the log scale often produces a near-symmetric shape and a sample skewness near zero. Three cautions. The transform is undefined at zero and for negative values, so it needs strictly positive data or a shift. It only works on *right* skew; on left-skewed data it makes things worse. And it changes the scale you interpret on, so an effect on log values is a multiplicative effect on the original scale, and exponentiating the average of the logs does not recover the arithmetic mean of the original values.
go deeper
Recall the shape fact: logs squeeze large values more than small ones, so a long right tail shrinks. Know that the input must be strictly positive.
Explain why concavity does the work, name the Box-Cox family with lambda = 0 as the log, and be able to say what a fitted lambda of 0.5 or -1 corresponds to.
Show that you check the result rather than assume it, handle zeros deliberately, and can state what a coefficient or summary on the log scale means back in the original units.
Own the tradeoff between statistical convenience and decision clarity. Decide when a transformed metric belongs in a reporting standard and when the interpretability cost outweighs the cleaner shape.
## Why the log pulls in a right tail Positive skewness on a strictly positive variable usually has a common structure: a hard floor at zero, a dense cluster of ordinary values, and a thin tail of values that are many times larger. House prices, session durations, file sizes, and claim amounts all look like this. The natural logarithm is **concave and increasing**. Increasing means it preserves order — the most expensive house stays the most expensive. Concave means it compresses larger inputs more aggressively than smaller ones: `log(1000) - log(100) = log(10) = log(100) - log(10)`. A tenfold step is the same distance no matter where it starts. So the stretch from 200,000 to 2,000,000 in the tail collapses to the same span as the stretch from 20,000 to 200,000 near the floor, and the tail stops dominating. Because skewness is the average of *cubed* z-scores, the far tail is where nearly all of the positive contribution comes from. Shrinking that region is precisely what drives the statistic toward zero. ## What the numbers do There is no guarantee, only a strong tendency. For data whose logs are close to symmetric, the log transform lands sample skewness near zero and typically pulls excess kurtosis down as well, because it also thins the fourth-power contributions from the tail. On a variable with a mild right skew, the log can **overcorrect** and leave you with a left-skewed shape and a negative skewness. Always recompute skewness on the transformed values rather than assuming the transform worked. ## The Box-Cox family The log is one member of a one-parameter family of power transforms. For strictly positive `y`, the Box-Cox transform is `y(lambda) = (y^lambda - 1) / lambda` for `lambda != 0` `y(lambda) = log(y)` for `lambda = 0` The `lambda = 0` case is defined as the log precisely because `(y^lambda - 1) / lambda` tends to `log(y)` as `lambda` goes to zero, which makes the family continuous in `lambda`. Useful landmarks: `lambda = 1` is essentially no transform (a shift and rescale), `lambda = 0.5` is a square root, `lambda = 0` is the log, and `lambda = -1` is a reciprocal. Smaller `lambda` means a more aggressive squeeze of the right tail. `lambda` is normally chosen by maximising a profile log-likelihood over a grid — in effect, picking the power that makes the transformed values look most like a symmetric bell. Two practical points: the fitted `lambda` usually comes with a wide interval, so it is common and sensible to round to a nearby interpretable value such as 0 or 0.5 rather than reporting `lambda = 0.07`; and the standard Box-Cox transform requires **strictly positive** input, which is the single most common reason it cannot be applied directly. ## Zeros and negatives The log is undefined at 0 and for negative numbers, and a variable such as session duration or revenue per user is full of exact zeros. Options in practice: - Add a small constant: transform `log(x + c)`. This works but is not innocent — the choice of `c` visibly changes the shape of the low end, and results can be sensitive to it. Choose `c` on substantive grounds (a natural measurement floor) and state it. - Use a transform defined for zero and negative values, such as the signed power family that extends Box-Cox, or the inverse hyperbolic sine. - Model the zeros separately from the positive part, if the zeros mean something structurally different (no session started at all versus a very short session). ## Left skew The log is a *right*-tail tool. Applied to left-skewed data it lengthens the left tail and makes skewness more negative. For left skew, reflect first (transform `max(x) + 1 - x`) and then apply a power transform, or use a power greater than 1 such as squaring or cubing, which stretches the upper end. ## Interpretation cost This is the tradeoff that separates a good answer from a mechanical one. After the transform you are describing a different quantity. Differences on a log scale are ratios on the original scale: a gap of 0.7 in natural logs is roughly a 2x multiplicative difference. And exponentiating the mean of the logs does **not** return the arithmetic mean of the original values — the exponential of an average is not the average of exponentials. So a summary produced on the log scale must either be reported on the log scale with that stated plainly, or back-transformed with an explicit and correct statement of what the back-transformed number represents. If the audience makes decisions in dollars, hiding a log transform inside the pipeline is a communication failure even when the statistics are cleaner.
- In the Box-Cox family, why is lambda = 0 defined as the logarithm?Because `(y^lambda - 1) / lambda` converges to `log(y)` as `lambda` approaches zero, so defining the `lambda = 0` case as the log makes the family continuous in the parameter. It also gives useful landmarks: `lambda = 0.5` is a square root, `lambda = 1` is no real transform, and `lambda = -1` is a reciprocal.
- What do you do when the right-skewed variable contains exact zeros?A plain log is undefined there. You can shift and use `log(x + c)` with `c` justified by the measurement floor and stated openly, use a transform defined at zero such as the inverse hyperbolic sine, or model the zeros as a separate component if they mean something structurally different from a small positive value.
- Can a log transform make skewness worse?Yes, in two ways. On left-skewed data it lengthens the already-long left tail and drives skewness further negative. On mildly right-skewed data it can overcorrect and leave a left-skewed result. Always recompute the skewness on the transformed values rather than assuming the transform did what you wanted.
- What is the main cost of reporting on the log scale?Interpretability. Differences in logs are ratios in the original units, so the audience must translate every number. Exponentiating the average of the logs does not give back the arithmetic mean of the original values, so any back-transformed figure needs an explicit statement of what quantity it actually represents.
A log scale is a camera lens that zooms in on the cheap end of the market and zooms out on the mansions, so the whole price range fits in one balanced frame.
saying these in an interview costs you the question
- Applies a log to data containing zeros or negatives
- Assumes the transform always removes skew without rechecking
- Uses a log on left-skewed data
- Exponentiates the mean of logs and calls it the mean
- Treats a transform as a fix for a few bad records