skip to content

Insurance claim sizes have a long right tail; how do you choose between log-normal and Pareto?

level: seniorimportance: should knowfreq 44%

answer

  1. log the data first and look
  2. does the logged shape become a bell?
  3. straight line on a log-log tail plot
  4. ask whether the mean even exists

basics

~10 s

Take logs of the claim sizes: a symmetric bell after logging points to log-normal. A tail that traces a straight line on a log-log survival plot points to Pareto. A normal fits neither.

solid answer

~50 s

Both candidates are positive and right-skewed, so the choice is decided in the tail. Two diagnostics do most of the work. First, take logarithms of the claim amounts: log-normal means the logged values are approximately normal, so a symmetric bell after logging is direct evidence for it. Second, plot the survival probability against claim size on log-log axes: a Pareto tail is a straight line there, while a log-normal curves progressively downward. The distinction is not cosmetic. A Pareto with tail index at or below 2 has infinite variance, and at or below 1 has no finite mean at all, so averages keep drifting upward as data accumulates and any planning built on a sample mean is unstable. Log-normal has all moments finite. A common working answer fits a log-normal body and a Pareto tail above a threshold.

go deeper

for a junior

Be ready to say that claim sizes are positive and right-skewed, so a symmetric normal is the wrong shape, and that logging the data is the first thing to look at.

for a middle

Explain what log-normal means as a definition, that the logarithm is normal, and describe how a log-log survival plot separates a power-law tail from a log-normal one.

for a senior

Show the decision consequences: which moments exist at a given tail index, why averages stop being usable, and how you would split the fit into a body and a tail with a justified threshold.

for a principal

Own the risk framing: whether the organisation should be steering on expected loss at all under a heavy tail, and what reinsurance, reserving or pricing decision the chosen family and threshold are feeding.

## Why the normal is off the table first Claim amounts are positive, strongly right-skewed, and dominated by a handful of enormous values. A normal distribution is symmetric, defined on the whole real line, and has a tail that decays extremely fast. Fitting one to claim sizes fails on all three counts: it assigns real probability to negative claims, it places the bulk symmetrically around a mean that no typical claim resembles, and it understates the probability of a very large claim by orders of magnitude. The last failure is the expensive one, because reserving and reinsurance decisions are driven almost entirely by the tail. A candidate who invokes the central limit theorem here is misapplying it. That theorem concerns the behaviour of an *average* of many claims, not the distribution of an *individual* claim, and for very heavy tails even the average converges slowly or not at all. ## The two serious candidates **Log-normal.** A quantity is log-normal when its logarithm is normally distributed. It arises naturally from multiplicative processes: a claim that is the product of many independent proportional effects, a revenue figure that is a base amount scaled by several independent multipliers. Support is strictly positive, the shape is right-skewed, and crucially every moment is finite, however large the spread. **Pareto.** A power-law family whose survival probability above a level `x` falls off like `x` raised to the negative tail index. Its defining feature is scale invariance in the tail: the ratio of the chance of exceeding 2 million to the chance of exceeding 1 million is the same as the ratio for 20 million versus 10 million. That is a much heavier tail than log-normal at extreme levels. ## Diagnostic one: log the data Log every claim amount and look at the shape. If the result is roughly symmetric and bell-like, log-normal is the natural fit, and this is the cleanest evidence you can present. If the logged data is still visibly right-skewed, with a tail that refuses to come in even on the log scale, you are looking at something heavier, and Pareto becomes the candidate. Note that the log here is being used as *evidence about the generating family*, not as a cosmetic fix applied to a sample. ## Diagnostic two: the log-log tail plot Sort the claims, compute for each level the fraction of claims that exceed it, and plot that fraction against the level with both axes on log scales. A Pareto tail appears as a straight line whose slope is the negative tail index. A log-normal tail curves steadily downward, dropping away from any straight line as you move right. This plot is the single most informative picture in heavy-tail work, and being able to describe it is a strong signal in an interview. ## Why the tail index is a business fact, not a statistical nicety For a Pareto with tail index `a`: - If `a` is greater than 2, both the mean and the variance are finite. - If `a` lies between 1 and 2, the mean is finite but the variance is infinite. Sample standard deviations will not settle down; they grow with sample size. - If `a` is at or below 1, even the mean is infinite. Running averages drift upward without limit as more data arrives, and a reported average is a statement about your sample size rather than about the process. Moments are nested: if a given moment is finite, all lower-order moments are finite too. A distribution cannot have an infinite mean but a finite variance. The operational consequence is that when the tail index is low, you stop planning on averages and plan on quantiles and exceedance probabilities instead: the chance of a claim above 10 million, the level exceeded once per decade. Those quantities remain meaningful when the mean does not. ## Fitting body and tail separately Real claim data often looks log-normal through the bulk and heavier than log-normal at the very top. The standard resolution is a split fit: choose a threshold, model everything below it with a log-normal, and model the excesses above it with a power-law tail. The threshold is chosen where the estimated tail index stops moving as the threshold rises, which is the point at which the tail behaviour has stabilised. Say this explicitly and you have shown you know that families are modelling tools rather than laws of nature. ## Sample-size honesty The extreme tail is, by construction, where you have almost no data. Ten thousand claims might contain three above 5 million, and those three carry the whole tail estimate. Report the uncertainty in the tail index, check how the estimate moves when you drop the single largest claim, and state whether the decision you support would change across that range. A senior answer names this fragility rather than presenting a fitted index as a measured constant. ## What a strong answer covers Rule out the normal for support, symmetry and tail decay. Name the two diagnostics. Connect the tail index to whether the mean and variance exist. Offer the body-plus-tail split. Close on how little data the tail estimate rests on.

  • Why is a normal a bad choice for claim sizes even with a very large sample?
    Sample size does not change the shape of the family you fit. A normal is symmetric, allows negative claims, and its tail decays far too fast, so it understates the chance of a very large claim by orders of magnitude no matter how much data you have. The central limit theorem describes averages of claims, not individual claims.
  • What does a fitted tail index below 1 imply for capital planning?
    It implies the theoretical mean does not exist, so running averages keep drifting upward as data accumulates and any plan anchored on an average claim is unstable. Plan instead on exceedance probabilities and high quantiles, and state clearly that the tail estimate rests on a handful of observations.
  • The body looks log-normal but the extreme tail looks heavier. What do you do?
    Split the fit. Model claims below a threshold with the log-normal and model the excesses above the threshold with a power-law tail. Pick the threshold at the point where the estimated tail index stops moving as you raise it, and report how sensitive the result is to that choice.

saying these in an interview costs you the question

  • Fits a normal because the sample is large
  • Invokes the central limit theorem to justify normal individual claims
  • Treats the sample mean as stable under a very heavy tail
  • Claims a distribution can have infinite mean but finite variance
  • Presents a tail index estimated from three claims as a firm constant

context