You log-transform right-skewed income before a t-test - which hypothesis are you now testing?
answer
- the transform changes the question
- means of logs are not logs of means
- exponentiating yields a ratio
- the long tail is what you flattened
- zeros have no logarithm
basics
~20 sOne about geometric means, not arithmetic ones. A t-test on logged values compares mean log income between groups, and exponentiating that difference gives a ratio of geometric means - not the difference in average income the business usually asked about.
solid answer
~50 sTaking logs before the test changes the estimand. The comparison now asks whether the mean of log income differs between the groups, and back-transforming the difference gives a ratio of geometric means rather than a difference of arithmetic means. That is often a perfectly sensible quantity: it is a multiplicative, percentage-style comparison, and on right-skewed income it is stable and interpretable. But if the decision depends on total or average money, geometric and arithmetic means can move differently, because the log deliberately compresses exactly the large values that dominate the total. Two practicalities go with it. Zeros and negatives have no logarithm, so any additive constant you introduce is an arbitrary modelling choice that changes the answer. And the transform should be declared before you see the outcome - trying several and keeping the significant one destroys the error rate you are reporting.
go deeper
Know that a logarithm compresses a long right tail and that comparing logged values is not the same as comparing the original values. Be able to say what a geometric mean is.
Explain the back-transformation: the difference in mean logs exponentiates to a ratio of geometric means, and that ratio minus one is a percentage difference in typical values.
Show that you track the estimand, not just the assumption. Give a case where a transformed analysis answers a different business question from the one asked, and say how you would report it honestly.
Own the analysis-plan rule: which scale a metric is compared on is declared up front and tied to the decision it feeds, so nobody swaps in a transform after seeing a disappointing result.
## Why anyone logs income in the first place Income, revenue per customer, session duration and waiting times share a shape: a dense mass of small values and a long right tail. That shape makes the arithmetic mean unstable, gives a few observations enormous leverage on the sample standard deviation, and slows the convergence that normally makes a mean-comparison well behaved at moderate sample sizes. A logarithm compresses the upper tail hard and stretches the lower one, and multiplicative processes - things that grow by percentages - become additive on the log scale, which is often close to symmetric. So the transform genuinely fixes the shape problem it was applied for. ## What it does to the question The cost is that the question changes. On the log scale you are comparing `mean(log X)` between groups. Exponentiate a group's mean log and you get the geometric mean: the n-th root of the product of the values, equivalently `exp(mean(log X))`. Exponentiate the difference between the two mean logs and you get the ratio of the two geometric means. So the confidence interval you report on the log scale becomes, after back-transformation, an interval for a ratio - a multiplicative effect - not for a difference in currency. That matters because geometric and arithmetic means are different quantities. By Jensen's inequality the geometric mean is never larger than the arithmetic mean, with equality only when every value is identical. For a lognormal variable with log-scale parameters `mu` and `sigma`, the geometric mean is `exp(mu)` while the arithmetic mean is `exp(mu + sigma^2 / 2)`. Two groups can therefore share a geometric mean and differ in arithmetic mean purely because one has a wider spread in logs - and that extra spread sits in the upper tail, which is exactly where total revenue comes from. A transformed analysis can honestly report 'no difference' while the totals differ substantially. ## Back-transformation is not a formality A related trap is exponentiating a mean of logs and calling the result an estimate of the mean. It is not: it estimates the geometric mean, which for skewed data is systematically smaller. For a lognormal it corresponds to the median rather than the mean. If a stakeholder asks 'how much more does group B earn on average', an exponentiated log difference does not answer that question, and presenting it as though it does is a real reporting error. Report it as what it is: a ratio or a percentage difference in typical values, since `exp(difference) - 1` is the proportional change. ## Zeros, negatives and arbitrary constants The logarithm is undefined at zero and negative. Real income data have zeros. The common workaround of adding a small constant before logging is a modelling decision disguised as a data-cleaning step: the smaller the constant, the more extreme the logged zeros become and the more leverage they gain over the result, and different analysts choosing different constants will reach different conclusions. If zeros represent a genuinely distinct state - unemployed rather than low-earning - the more honest treatments are to model the presence of income and its amount separately, or to work on the original scale with a method that tolerates skew. ## When transforming is the right call Transform when the multiplicative comparison is the one you actually want. Percentage-style effects, elasticities, and outcomes generated by proportional processes all argue for the log scale, and there the geometric mean is not a compromise but the correct target. Transform, too, when the sample is small and the raw skew is severe enough that a mean comparison would be unreliable, and say clearly in the write-up which quantity you compared. Do not transform reflexively. At a large sample size the central limit theorem already handles moderate skew for a comparison of means, so the transform buys little and costs interpretability. And if the decision is about a total - budget, revenue, cost - the arithmetic mean is the decision-relevant quantity, and it is better to face the skew directly than to answer a cleaner question nobody asked. ## The discipline point Choose the scale before seeing the outcome. Trying the raw comparison, then the logged one, and reporting whichever crosses a threshold inflates the false-positive rate in a way no assumption check will reveal. State the transform in the analysis plan, justify it by the shape you expected and the estimand you want, and report the effect on the scale a decision-maker can act on. An interviewer asking this question is checking whether you know that an assumption fix can quietly replace the hypothesis, not just clean the data.
- Can two groups have equal geometric means but different arithmetic means?Yes, easily. For a lognormal variable the geometric mean is `exp(mu)` while the arithmetic mean is `exp(mu + sigma^2 / 2)`, so a group with the same `mu` but a wider log-scale spread has the same geometric mean and a larger arithmetic mean. That is not a rounding artefact: the extra variance sits in the upper tail, which is precisely where totals such as revenue come from.
- How should you report the result of a t-test run on logged values?As a multiplicative effect. Exponentiate the difference in mean logs and its confidence limits, then present it as a ratio or percentage difference between typical values - 'group B is about 12% higher' rather than a currency amount. Do not exponentiate and then describe the result as a difference in average income, because that is a claim the analysis never made.
- What if the outcome contains zeros?Then the logarithm is undefined and every workaround becomes a modelling decision. Adding a small constant makes the answer depend on that constant, and the smaller it is the more leverage the zeros gain. It is usually better to ask what the zeros mean: if they are a distinct state, model presence and amount separately, or stay on the original scale with a method that tolerates the skew.
saying these in an interview costs you the question
- Says a log transform just tidies the data and changes nothing
- Exponentiates the result and calls it a difference in average income
- Adds an arbitrary constant to zeros without reporting the choice
- Tries several transforms and keeps the one that reaches significance
- Assumes equal geometric means implies equal arithmetic means