Your resampled 95% interval for p99 latency from 500 requests is enormous - why?
answer
- count how many points beat the p99
- about five out of five hundred
- resamples reuse the same tail values
- endpoints snap to observed data points
- separate replicate noise from data noise
basics
~20 sOnly about five of 500 observations exceed the p99, so the estimate rests on two or three sorted values. Resampling only reshuffles those, giving a wide, chunky interval. The width is real information, not a defect.
solid answer
~50 sWith n = 500, roughly five observations exceed the p99 and the estimate is effectively determined by the two or three largest sorted values. Resampling draws from those same observations, so the resampled p99 can only ever take one of a small number of distinct values and its upper endpoint can never exceed the sample maximum. That produces an interval that is both very wide and visibly chunky. Separate the two noise sources before reacting: Monte Carlo error, from using too few resamples, shrinks for free if you raise the number of replicates, while sampling error, from having almost no data in the tail, does not - it needs more observations, a longer window, or a lower quantile. Usually the right response is to report p95 instead of p99, or widen the aggregation window. A wide interval is the honest statement that 500 requests do not pin down a p99.
code
python · 21 linesimport random
random.seed(7)
def p99(xs):
s = sorted(xs)
return s[int(0.99 * len(s)) - 1]
def draw(n=500):
return [random.lognormvariate(0.0, 1.0) for _ in range(n)]
print('p99 of five fresh 500-point samples:',
[round(p99(draw()), 1) for _ in range(5)])
sample = draw()
boot = sorted(p99([random.choice(sample) for _ in range(500)])
for _ in range(2000))
print('point estimate:', round(p99(sample), 1),
'interval:', round(boot[50], 1), 'to', round(boot[1949], 1))
print('distinct values the endpoints can take:', len(set(boot)))
print('six largest points in the sample:', [round(x, 1) for x in sorted(sample)[-6:]])go deeper
Know that a p99 from a few hundred observations rests on only a handful of points, so it is a noisy number. Being able to say the interval should be reported alongside it is enough here.
Explain why a resampled tail quantile is discrete: resamples only ever contain values already in the sample, so the estimate snaps to observed points and the upper endpoint is capped by the sample maximum.
Show you can triage in production - distinguish Monte Carlo error from sampling error, choose between reporting p95, widening the window, or modelling the tail, and refuse to declare a regression that sits inside the interval.
Set the policy. Decide what quantiles the organisation may commit to at each traffic volume, require intervals on tail metrics in dashboards and alerts, and take the argument that a noisy p99 target creates false pages rather than reliability.
## Count the data that actually drives the estimate Start by counting. The sample p99 from 500 observations is one of the top few sorted values; only about 1% of the sample, five observations, lies above it. Whatever machinery you wrap around that, the number you are reporting is a function of two or three specific requests. The estimate inherits their randomness wholesale, which is why the interval around it is wide: the interval is telling you the truth about how little evidence there is. ## Why the interval also looks lumpy A resampling interval is built from many resampled datasets drawn from the observations you already have. A resampled p99 is again a top-few order statistic of that draw, so it can only be one of the values present in the original sample. Across thousands of resamples the p99 therefore takes only a handful of distinct values, and both interval endpoints snap to observed data points. Two consequences follow. Endpoints move in visible steps rather than smoothly, and the upper endpoint is hard-capped by the sample maximum - resampling cannot invent a slower request than the slowest one you saw, even though the population certainly contains some. ## Two different noises, two different fixes When the interval moves between runs, decide which noise you are looking at. **Monte Carlo error** comes from using too few resamples. It is a property of your computation, not of your data, and it shrinks as you add replicates. If two runs on the *same* dataset give different endpoints, raise the replicate count until they stabilise; that is cheap and it is the only free fix here. **Sampling error** comes from having almost no observations above the quantile. If two runs on *different* 500-request samples give different estimates, that is genuine and no amount of computation removes it. Only more data, a wider aggregation window, more traffic pooled together, or a less extreme quantile helps. Candidates who conflate these will try to fix a data problem by turning up the replicate count and conclude that the method is broken when nothing improves. ## Why no formula rescues you either The closed-form standard error of a sample quantile is `sqrt(p(1-p)/n) / f(q_p)`, which depends on the probability density at the quantile. At p99 in a right-skewed latency distribution that density is tiny and must itself be estimated from the same handful of tail points, so the plug-in route is, if anything, less trustworthy than resampling. There is no formula that manufactures information the sample does not contain. ## What to do instead 1. **Report a lower quantile.** p95 from the same 500 observations rests on roughly 25 points above it instead of five, and its interval will be dramatically narrower. If the operational question tolerates p95, this is the cheapest real fix. 2. **Aggregate more.** A longer time window, or pooling comparable traffic, raises n directly. Watch that the pooled population is still homogeneous - stitching together peak and off-peak traffic changes the distribution you are estimating. 3. **Model the tail.** If you must speak about p99 or beyond from limited data, fit a tail model above a high threshold rather than reading an empirical order statistic. This buys extrapolation at the cost of a modelling assumption you must defend. 4. **Always publish the interval.** The single most valuable habit is refusing to quote a bare tail number. A p99 of 340ms with an interval from 250ms to 600ms stops a team from chasing a 30ms "regression" that is inside the noise. ## Ties and rounding make it worse If latency is recorded in whole milliseconds, the top of the sample may contain repeated values. The resampled quantile then sits on a single repeated number in a large share of resamples, and the interval can collapse to a point or to two adjacent values - a spuriously confident-looking result produced by measurement granularity rather than by evidence. Check the number of distinct values near the quantile before believing a narrow tail interval. ## The one-line version A p99 from 500 requests is an estimate built from about five observations; the enormous interval is the method reporting that fact honestly, and only more data or a less extreme quantile changes it.
- Would raising the number of resamples from 200 to 20,000 narrow the interval?No. More replicates reduce Monte Carlo error, so the endpoints stop wobbling between runs on the same data, but they converge to a width set by the data itself. That width reflects the handful of tail observations and only more data - or a less extreme quantile - reduces it.
- What does it mean when the interval's upper endpoint equals the sample maximum?It means the procedure has run out of evidence. Resampling can only reuse observed values, so it cannot represent the possibility that the population tail extends well beyond your slowest recorded request. Treat that endpoint as a floor on the true upper limit, and switch to an explicit tail model if you need to speak about values past the observed range.
- Your p99 interval comes back suspiciously narrow on millisecond-rounded data - what do you check?Count distinct values near the quantile. Heavy rounding creates ties, so most resamples return the same repeated value and the interval can collapse to one or two adjacent numbers. That narrowness is measurement granularity, not precision; re-examine at a finer timing resolution before reporting it.
It is like estimating the height of the tallest tree in a forest after measuring five trees; measuring those five more carefully does not tell you about the ones you never walked past.
saying these in an interview costs you the question
- Blames the wide interval on the method rather than the data
- Adds resamples expecting the interval to narrow
- Quotes a bare p99 with no uncertainty at all
- Treats a tail move inside the interval as a regression
- Assumes resampling can produce values beyond the sample maximum