In a Fermi estimate, how do you find the assumption that dominates the answer's uncertainty?
answer
- in a product, relative errors add
- rank factors by ratio, not by size
- widest plausible high-over-low wins
- halve it and double it, re-run
- spend remaining effort only on that factor
basics
~20 sRank factors by the width of their plausible range as a ratio, not by their size. In a product relative errors add, so a factor uncertain by 3x dominates one known to 10 percent. Halve and double it.
solid answer
~50 sIn a multiplicative chain the relative errors roughly add, so the dominant assumption is the one with the widest plausible range as a ratio, regardless of how large the number itself is. In the piano-tuner chain, the city population is known within maybe 20 percent and is irrelevant to the uncertainty; tunings per piano per year could plausibly be 0.3 or 1.5, a 5x spread, and that factor drives everything. So I test it directly: at half the assumed rate the answer falls from about 50 tuners to about 25, at double it rises to about 100. I quote the estimate as a range built from those swings, and if more work is available I spend it entirely on that one factor. Refining a factor you already know to 10 percent cannot move an answer whose spread is 5x.
go deeper
Be ready to name which single assumption you are least sure of and to re-run the chain with it halved and doubled. Producing a low-high range instead of one number is the habit being checked.
An interviewer at this level expects you to explain why relative errors add in a product, and therefore why a small-magnitude factor with a wide plausible ratio outweighs a large one you know well.
Show that you triage effort by influence: rank every factor by its high-over-low ratio, state which one you would go measure, and handle correlated assumptions as joint scenarios rather than independent sweeps.
Own the limits of the technique. Argue where sensitivity analysis stops working, namely structural omissions it cannot see, and set the expectation that a headline estimate always travels with its dominant assumption named.
## Why one assumption usually owns the answer A Fermi estimate is almost always a product of factors. When you multiply quantities, the *relative* uncertainties combine, and a useful rule of thumb is that the relative error of a product is approximately the sum of the relative errors of its factors. Concretely: if factor A might be off by 10 percent and factor B might be off by a factor of three, the product's uncertainty is essentially B's uncertainty. A is noise. This has a consequence that surprises people: **the magnitude of a factor tells you nothing about how much uncertainty it contributes**. A population of 3 million and a rate of 1 per year contribute uncertainty purely according to how confidently you know each of them, not according to how big they are. ## Finding the dominant factor Go through the chain and for each factor write a low and a high value you would genuinely defend, then compute the *ratio* high/low. For the piano-tuner chain: | Factor | Low | High | Ratio | |---|---|---|---| | City population | 2.5M | 3.5M | 1.4 | | People per household | 2.2 | 2.8 | 1.3 | | Share of households with a piano | 2% | 8% | 4 | | Tunings per piano per year | 0.3 | 1.5 | 5 | | Tunings per tuner per year | 700 | 1,400 | 2 | The two right-hand rows dominate, and tunings per piano per year dominates most. Population and household size, which feel like the *big* numbers, contribute almost nothing to the spread. This ranking is the deliverable. It tells you where the answer's honesty lives. ## Testing the dominant assumption Once identified, vary it and re-run: - At **half** the assumed tuning rate (one tuning every two years), annual demand falls from 50,000 to 25,000, and the answer drops from about 50 tuners to about **25**. - At **double** (two tunings a year, closer to serious players and institutions), demand rises to 100,000 and the answer rises to about **100**. So the honest statement is: *about 50, plausibly 25 to 100, driven almost entirely by how often a piano actually gets tuned*. That sentence is worth far more to a decision-maker than the bare number 50, because it says both what you believe and what would change your mind. ## Additive terms behave differently The adding-relative-errors rule applies to products. Where the chain **adds** terms — household pianos plus institutional pianos, for example — a term contributes uncertainty in proportion to its **share of the total**. A term that is 5 percent of the sum can be off by 100 percent and still move the answer by only 5 percent, which is why small additive terms can simply be dropped with a one-line justification. The mirror-image mistake is spending half your time carefully estimating the institutional pianos while leaving the dominant household tuning rate as an unexamined guess. ## Where the rule breaks **Correlated assumptions.** If two factors move together, you cannot vary one while holding the other fixed. Suppose you assume a high piano ownership rate because you pictured an affluent city; those same households probably also tune more often. Doubling both is a coherent scenario; doubling one alone is not. When factors are linked, vary them jointly as scenarios rather than one at a time. **Structural error.** Sensitivity analysis explores uncertainty *inside* your model. It cannot detect that the model is missing a term entirely — for instance, forgetting that many tuners also repair and sell instruments, so tuning is only part of their year. No amount of varying the existing factors reveals an omitted one. The defence is a separate cross-check from a different direction, or against an independently known total. **Asymmetric ranges.** Some factors are bounded on one side. A participation rate cannot exceed 1, and a count cannot go below zero, so a symmetric halve-and-double sweep can propose impossible values. Clip the sweep at the physical bounds and say you did. ## Turning it into a decision about effort The practical payoff of the ranking is triage. If you get one more hour to improve the estimate, it goes entirely to the top factor. Researching city population to three digits when the tuning rate is uncertain by 5x is effort that provably cannot change the answer. Interviewers ask this question precisely to see whether you allocate effort by influence rather than by ease. A short script that works: *this factor drives the answer, here is what it does when I halve and double it, and here is the one thing I would go find out.* ## What a weak answer looks like A weak answer varies every factor a little, reports that the answer moved a little, and concludes the estimate is robust. That conclusion is an artifact of testing every factor over the same narrow window rather than over each factor's own honest range. The ranges have to be defensible per factor, and they are almost never equal.
- Why does the largest number in the chain not automatically dominate the uncertainty?Because in a product the uncertainties combine as ratios, not as absolute amounts. A population of 3 million known to within 20 percent contributes a factor of 1.2 to the spread; a rate of 1 per year that could plausibly be 0.3 or 1.5 contributes a factor of 5. Magnitude is irrelevant once every factor is expressed relatively, which is why ranking by high-over-low is the correct move.
- How does a term that is added rather than multiplied change the analysis?An additive term contributes uncertainty in proportion to its share of the sum. A term making up 5 percent of the total can be wrong by 100 percent and shift the answer by only 5 percent, so it can be dropped with a stated justification. That is the opposite of a multiplicative factor, where even a small-magnitude factor with a wide ratio can dominate everything.
- What can a sensitivity sweep never tell you?That the model is missing a term. Varying the factors you wrote down explores uncertainty inside your structure and is blind to anything outside it — a whole category of demand you forgot, or a workforce that spends only part of its time on the activity you counted. Catching structural error requires an independent estimate from a different direction or a comparison against a known total.
saying these in an interview costs you the question
- Ranks factors by magnitude instead of relative range
- Varies every factor by the same fixed percentage
- Refines a factor already known to within ten percent
- Calls the estimate robust after a uniformly narrow sweep
- Varies correlated factors independently as if unlinked
- Treats sensitivity analysis as proof the model is complete