skip to content

How would you estimate a factory's total output from a sample of observed serial numbers?

level: seniorimportance: nice to knowfreq 24%

answer

  1. the largest one seen can only be too small
  2. add the average gap between observations
  3. invert the expectation of the maximum
  4. compare against an estimator built from the mean
  5. check it cannot fall below what you saw

basics

~10 s

The largest serial number observed always underestimates the true total, since it can never exceed it. Scale it up: with n serials seen and maximum m, the estimate m*(n+1)/n - 1 is unbiased.

solid answer

~50 s

Model the serials as the numbers 1 through N with your n observations a random sample from them. The natural guess, the largest serial seen, is biased low by construction — it can never exceed N — and in fact `E[max] = n*(N + 1)/(n + 1)`. Inverting that gives an unbiased estimator, `N_hat = max * (n + 1)/n - 1`. A second unbiased option comes from the mean, since `E[xbar] = (N + 1)/2` yields `N_hat = 2*xbar - 1`. Both are centred on N, so the choice is efficiency, and the gap is dramatic: the maximum-based estimator's spread shrinks like `N/n` while the mean-based one shrinks only like `N/sqrt(n)`, and the latter can even return a value below a serial you have already seen. The real interview content is the assumptions: numbering from one, no gaps or restarts, and a sample uncorrelated with production order.

code

python · 13 lines
python
import random, statistics

N, n, trials = 1000, 10, 20000
from_mean, from_max = [], []
for _ in range(trials):
    s = random.sample(range(1, N + 1), n)
    from_mean.append(2 * statistics.mean(s) - 1)
    from_max.append(max(s) * (n + 1) / n - 1)

for name, est in (("2*xbar - 1", from_mean), ("max*(n+1)/n - 1", from_max)):
    print(name,
          "mean estimate", round(statistics.mean(est), 1),
          "sd", round(statistics.stdev(est), 1))

go deeper

for a junior

Be ready to explain why the largest serial you observed must underestimate the total, and that the fix is to scale it up by a factor depending on how many serials you saw.

for a middle

Expect to derive it: the expected maximum is n times (N+1) over (n+1), so inverting gives max times (n+1)/n minus one. Check the formula at a sample size of one.

for a senior

Show you compare rival unbiased estimators on variance rather than stopping at unbiasedness, and that you stress-test the model — numbering scheme, voided blocks, and whether the sample is random with respect to production order.

for a principal

Own the decision this feeds: how much a tenth-of-total error band changes the call being made, whether the assumptions are defensible to a sceptical audience, and when to invest in better sampling instead of a better estimator.

## The model You see a handful of serial numbers stamped on units and want the total produced. Assume units are numbered `1, 2, ..., N` with `N` unknown, and that your `n` observed serials are a simple random sample of distinct values from those N. Everything below follows from that model, and the model is where the real risk lives. ## Why the observed maximum is the wrong answer alone Let `m` be the largest serial you observed. Because `m <= N` always, `m` can never overshoot; it can only equal N or fall short. An estimator that errs in one direction only is biased, and here the exact expectation is ``` E[m] = n*(N + 1)/(n + 1) ``` With n = 4, that is about 0.8 of `N + 1` — a 20% systematic shortfall. This is a rare, clean case where the sign of the bias can be argued in a sentence without algebra, which is exactly why interviewers like it. ## The corrected estimator Invert the expectation. Since `E[m] = n*(N+1)/(n+1)`, setting ``` N_hat = m * (n + 1)/n - 1 ``` gives `E[N_hat] = (N + 1) - 1 = N`. Unbiased, exactly, at every sample size. A useful way to read the formula is `N_hat = m + (m/n - 1)`: take the largest serial seen and add the average gap between consecutive observed serials, since on average one such gap of unseen units sits above your maximum. Sanity check with n = 1: a single observed serial `m` gives `N_hat = 2m - 1`, and since `E[m] = (N+1)/2`, that is indeed unbiased. ## The competing unbiased estimator The sample mean offers a different route. The mean of `1, ..., N` is `(N + 1)/2`, so `E[xbar] = (N + 1)/2` and ``` N_hat_mean = 2*xbar - 1 ``` is also unbiased. Two unbiased estimators of the same parameter — so the choice is decided by variance. ## Why the maximum-based estimator wins The maximum-based estimator's standard deviation shrinks roughly like `N/n`, while the mean-based one shrinks like `N/sqrt(n)`. At n = 10 that is roughly a twofold difference in typical error, and the gap widens with every extra observation. The reason is informational: for this uniform-range problem the maximum carries essentially all the information about the upper limit, while the sample mean spends its precision estimating a centre and then extrapolates to the edge. There is also a hard logical defect in the mean-based version. Suppose n = 4, the observed serials are 3, 12, 37 and 60. Then `xbar = 28` and `2*xbar - 1 = 55` — an estimate of the total output that is smaller than a serial number you are holding in your hand. An unbiased estimator can still produce impossible values, which is a memorable demonstration that unbiasedness is not the same as sensible. The maximum-based estimator on the same data gives `60 * 5/4 - 1 = 74`, which at least respects what was observed. ## The assumptions are the interview The arithmetic is the easy half. What distinguishes a strong answer is naming what would break it: - **Numbering starts at 1 and increments by 1.** Many real schemes start at 1000, skip blocks, or embed dates and plant codes. If the scheme is `plant-year-sequence`, you must parse out the sequence field first, and then you are estimating that plant-year's output, not total output. - **No gaps or restarts.** Serials voided in QA, or counters reset each year, break the range assumption in ways that bias the answer in unpredictable directions. - **The sample is random with respect to serial order.** This is the assumption that fails most often. Units captured in one place at one time are typically consecutive production, so the observed maximum is not a draw from the whole range — it reflects when you sampled, not how much exists. - **A fixed N.** If production continues while you sample, N is a moving target and the question becomes a rate estimate rather than a total. - **Serials are visible only on units that survived.** If failed or destroyed units carry the highest numbers or the lowest, the sample is not representative of the range. ## Saying it in an interview State the model, argue the downward bias of the maximum in one sentence, produce `m*(n+1)/n - 1` and check it at n = 1. Offer the mean-based rival, note both are unbiased, and separate them on variance — `N/n` versus `N/sqrt(n)` — plus the impossible-estimate defect. Then spend real time on the assumptions, because in practice the estimator is never the thing that fails; the numbering scheme and the sampling mechanism are.

  • Why does the maximum carry more information about N than the sample mean does?
    The unknown is the upper limit of the range, and the largest observed value is the closest evidence you have to that limit — every other observation only tells you the limit is at least that big. The sample mean spends its precision locating the centre and then doubles out to the edge, so all the noise in the centre estimate is magnified. That is why its spread shrinks like one over the square root of n rather than one over n.
  • What is the single assumption most likely to break this in the field?
    That the observed serials are random with respect to production order. Units seen together usually come from the same batch, so the observed maximum reflects when you sampled rather than how much was produced. Sequential numbering from one, no voided blocks and no annual counter resets are also frequently violated, but a clustered sample is the failure that quietly produces a confident wrong answer.
  • How would you report uncertainty around such an estimate?
    The estimator's spread scales roughly like N over n, so with n = 10 a typical error is around a tenth of the total — that alone tells the reader how seriously to take it. Because the estimator's distribution is skewed and bounded below by the observed maximum, a symmetric plus-or-minus interval is inappropriate; any interval should respect the hard floor that N cannot be smaller than the largest serial actually seen.

Reading four page numbers torn from a book, the highest one you hold is a floor on the book's length, never the length itself. You add a plausible number of unseen pages above it, and how many depends on how sparse your scraps are.

saying these in an interview costs you the question

  • Reports the observed maximum as the estimate itself
  • Cannot say which direction the maximum errs in
  • Picks the mean-based estimator purely because it is unbiased
  • Accepts an estimate below an already-observed serial number
  • Never questions whether serials start at one
  • Assumes units seen together are a random sample of the range

context