skip to content

How would you test whether leading digits in an expense dataset follow Benford's law?

level: seniorimportance: nice to knowfreq 20%

answer

  1. leading digit, nine categories
  2. logarithmic, not uniform
  3. one is far commoner than nine
  4. nine categories minus one
  5. needs several orders of magnitude

basics

~20 s

Tally the first significant digit of every amount into nine categories, compute expected counts as N times log10(1 + 1/d), and run a chi-square goodness-of-fit test on 8 degrees of freedom. Rejection flags the data for review, never proves fraud.

solid answer

~50 s

Benford's law gives the leading digit d a probability of `log10(1 + 1/d)`, so digit 1 appears about 30.1% of the time and digit 9 about 4.6%. Extract the first significant digit of every expense amount, tally the nine categories, set `E_d = N * log10(1 + 1/d)`, and compute `sum((O - E)^2 / E)` against a chi-square distribution on `9 - 1 = 8` degrees of freedom. Before trusting the result, check that the data are the kind Benford's law describes: many values, spanning several orders of magnitude, with no imposed floors, ceilings, round-number thresholds or sequential numbering. Then check the sample size cuts the other way too — with millions of rows an audit-irrelevant deviation will reject. Report which digits are over-represented alongside the p-value, and treat a rejection as a pointer to investigate, not a verdict.

go deeper

for a junior

Know that Benford's law makes leading digit 1 the most common at roughly 30% and digit 9 the rarest, and that comparing tallies against it is a goodness-of-fit problem, not a test about averages.

for a middle

Set the test up correctly: nine categories, expected counts of N times log10(1 + 1/d), and 8 degrees of freedom because no parameter is estimated from the data. Be able to say why it is not 9.

for a senior

Demonstrate that you validate eligibility before testing — spread across orders of magnitude, no caps, thresholds or assigned identifiers — and that you separate statistical significance from audit-relevant deviation on very large ledgers.

for a principal

Own how the result feeds a process. Decide the granularity at which the test is run, what deviation triggers human review, and how the organisation avoids treating an automated flag as an accusation.

## The law Benford's law describes the distribution of the **first significant digit** — the leftmost non-zero digit — in many naturally occurring collections of numbers. It states P(first digit = d) = log10(1 + 1/d), for d = 1, 2, ..., 9 which gives roughly 30.1% for 1, 17.6% for 2, 12.5% for 3, and declines to about 4.6% for 9. The intuition is scale invariance: if a collection of amounts spans several orders of magnitude and there is no privileged unit, the distribution of the digits must be unchanged when you convert currencies or change units, and the logarithmic law is the distribution with that property. Values grow multiplicatively, so a quantity spends longer climbing from 100 to 200 than from 900 to 1000, and its leading digit is 1 more often. The law is used in audit because most people fabricating amounts spread their leading digits far too evenly, or cluster them just under approval thresholds. ## The test, step by step 1. **Define the categories.** Nine of them, one per leading digit. Strip signs, currency symbols and leading zeros; 0.0472 has leading digit 4, as does 47,200. 2. **Tally observed counts.** One tally per amount. These are the O values. 3. **Compute expected counts.** `E_d = N * log10(1 + 1/d)`. With N = 20,000 expenses, the digit-1 cell expects about 6,021 and the digit-9 cell about 916. Leave the expectations fractional. 4. **Form the statistic.** `sum over d of (O_d - E_d)^2 / E_d`. 5. **Degrees of freedom.** Nine categories, no parameters estimated from the data — the probabilities come from the law itself — so `df = 9 - 1 = 8`. 6. **Check the expected counts** are large enough for the chi-square approximation. With realistic audit volumes this is rarely the binding constraint; the digit-9 cell is the smallest and needs N of only a few hundred to clear the usual minimum. ## What must be true of the data This is where a strong answer separates itself. Benford's law is not a law of nature that all numbers obey, and testing against it when the data could not possibly follow it produces a guaranteed rejection that means nothing. Look for: - **Sufficient spread.** The amounts should cover several orders of magnitude. A dataset of daily meal claims that all fall between 8 and 40 has no reason to be Benford-distributed. - **No imposed bounds.** Hard floors and caps — a reimbursement ceiling, a minimum invoice size — distort the leading digits directly. - **No assigned or sequential numbers.** Invoice IDs, employee numbers, postcodes and telephone numbers are labels, not measured quantities, and follow whatever scheme assigned them. - **No structural rounding.** Amounts habitually recorded to the nearest 50 or 100 change the digit distribution mechanically. - **A single homogeneous population.** Mixing categories with wildly different price structures can produce a rejection driven by composition, not by any individual anomaly. ## Reading the outcome Audit datasets are often enormous, and the chi-square statistic grows with N. With a million rows, a deviation of a fraction of a percentage point in the digit-1 share will reject decisively while being of no audit interest whatsoever. So the p-value alone is a poor decision rule. Report the **observed and expected proportions per digit side by side**, so the reader can see whether digit 1 is under-represented by 0.2 points or by 8, and inspect standardised residuals `(O - E)/sqrt(E)` to identify which digits drive the result — while keeping in mind you are scanning nine cells at once. Direction matters interpretively. A shortfall of leading 1s with a bulge in the middle digits is the classic fabrication signature. A spike at a specific digit often points instead at a threshold effect: if expenses above 5,000 need extra approval, expect an unnatural pile of amounts starting with 4. ## What the result licenses A rejection says the leading-digit distribution does not match the law. It does **not** identify fraud, a person, or even an anomaly of any particular kind — mundane causes such as a pricing change, a threshold policy, or a mixed population explain most rejections. The correct output of the test is a prioritised list of sub-populations to examine with domain methods. Equally, a non-rejection is weak evidence of cleanliness: a small number of fabricated entries inside a large clean population will not move the digit distribution at all, so the test has almost no power against targeted manipulation. A sensible practice is to run the test at the level where an anomaly would be actionable — by cost centre, vendor or period rather than across the whole ledger — so that a rejection points somewhere specific, while accepting that this multiplies the number of tests and requires you to interpret the collection of p-values, not each one in isolation.

  • Why are the degrees of freedom 8 rather than 7 for this test?
    There are nine leading-digit categories and one degree of freedom is lost because the counts must total N, giving 9 - 1 = 8. You would subtract further only if a parameter of the hypothesised distribution had been estimated from the same data, and Benford's probabilities are fixed by the law, not fitted.
  • The test rejects on a ledger of two million rows. What do you check before escalating?
    The size of the deviation, not just its significance. Put observed against expected proportions per digit: at that sample size a fraction of a percentage point rejects. Then check the data are Benford-eligible at all — spread over orders of magnitude, no caps, no thresholds, no assigned numbers — and look for a digit spike suggesting an approval threshold.
  • Why can a Benford test pass on data that does contain fabricated entries?
    Because power against a small contaminated subset is very low. A few hundred invented amounts inside a million genuine ones barely shift the aggregate digit distribution. That is why the test is run per vendor, cost centre or period, where the fabricated share of the population is large enough to register.
  • Which datasets should not be tested against Benford's law at all?
    Anything whose values are assigned rather than measured, or whose range is tightly bounded: invoice or employee identifiers, postcodes, telephone numbers, ages, percentage scores, and amounts hemmed in by a floor and a cap. These cannot follow the law, so a rejection carries no information about the process that produced them.

Think of an amount growing steadily like a bank balance compounding. It lingers a long while in the hundreds before crossing two hundred, but sprints from nine hundred to a thousand, so it is caught with a leading 1 far more often than with a 9.

saying these in an interview costs you the question

  • Expects leading digits to be uniformly distributed
  • Uses nine degrees of freedom for nine digit categories
  • Treats a rejection as proof of fraud
  • Applies the law to invoice IDs or postcodes
  • Escalates on a tiny deviation because the sample is huge
  • Reads a non-rejection as evidence the ledger is clean

context