skip to content

Choosing a Distribution

Reading a described process and naming the law it implies: counts per hour, time to failure, a bounded proportion, a sum of many small effects. It tests modelling judgement rather than formula recall.

on this pageshow

questions

5

Which distributions model 'outages per week' and 'hours between outages' for the same incident stream?

level: middleimportance: must knowfreq 70%

answer

  1. same stream, two ways to measure it
  2. counts in a window versus gaps between events
  3. one rate parameter links both views
  4. mean gap is one over the rate

basics

~10 s

Counts in a fixed window, outages per week, are Poisson. The gaps between consecutive outages, hours between them, are exponential. Both describe one event stream: one view counts events, the other measures waiting times.

solid answer

~50 s

They are two views of the same model. If outages arrive independently at a roughly constant average rate `r`, then the number in any fixed window is Poisson with mean `r * window`, and the time between consecutive outages is exponential with mean `1 / r`. At 2 outages per week the mean gap is 0.5 week, or 84 hours. Which one you pick is decided by what you actually measured: aggregated counts per bucket point at the count family, raw timestamps whose differences you take point at the waiting-time family. The shared assumption is the fragile part. If outages cluster into storms, or the rate swings between business hours and nights, neither fit is honest, and you move to a time-varying rate or a clustered arrival model rather than forcing a single constant rate.

go deeper

for a junior

Be ready to say which quantity is a count and which is a duration, and to name the matching family for each without hesitating.

for a middle

An interviewer expects you to connect the two views through a single rate, convert between rate and mean gap with the units carried correctly, and state the independence and constant-rate assumptions unprompted.

for a senior

Show that you check the assumptions against real operational data: bursts, time-of-day swings, batched alerts. Say what you would switch to when a single constant rate does not survive contact with the incident log.

for a principal

Own the framing decision: whether the organisation should model individual alerts, deduplicated incidents, or clusters at all, and what downstream capacity or on-call decision the chosen family is actually feeding.

## One stream, two measurement axes An incident stream is nothing but a set of timestamps: the instants at which outages begin. There are two natural things to record about such a stream, and each has its own natural distribution family. **Counting on the time axis.** Chop time into equal buckets, one week each, and count how many outages fall in each bucket. You now have a sequence of non-negative integers: 1, 0, 3, 2, 0. The family for counts of independent events in a fixed window is the Poisson family, with one parameter, the expected number per window. **Measuring between events.** Instead, sort the timestamps and take successive differences: 31 hours, 190 hours, 12 hours. You now have positive real numbers with no upper bound. The family for the waiting time between events that arrive at a constant average rate is the exponential family, again with one parameter. The two are locked together. If events arrive at rate `r` per unit time, the count in a window of length `w` has mean `r * w`, and the gap between events has mean `1 / r`. Estimate `r` from either view and you can express the other. Ten alerts per day means a mean gap of 2.4 hours; a mean gap of 3 minutes means 20 events per hour. ## Choosing between them in an interview answer The choice is not a matter of taste, it is a matter of what the data records. - The data is already bucketed (tickets per hour, defects per batch, arrivals per minute) and the buckets are equal width: you are modelling counts. - The data is raw event times, or already a list of durations between events: you are modelling gaps. - The question you must answer decides too. *What is the chance of more than five outages next week* is a count question. *What is the chance the next outage comes within an hour* is a waiting-time question. Both are answerable from one estimated rate. ## The assumptions that make either fit legitimate Three conditions sit under both families: 1. **Events are independent.** One outage does not make the next more likely. 2. **The rate is constant** over the period you model. 3. **Events do not coincide.** Two outages do not begin at the exact same instant; if they do, you are counting something that arrives in batches. In real operational data the second and third fail constantly. A cascading failure produces five alerts in ninety seconds, then silence for a week. Traffic-driven errors are ten times more frequent at midday than at 4 a.m. Both break the constant-rate picture, and both leave a recognisable footprint: far more quiet buckets *and* far more extreme buckets than a single-rate count model predicts, and gaps that are either very short or very long with few in between. ## What to do when the assumptions fail Do not abandon the framing, refine it. If the rate varies predictably with time of day or day of week, model the rate as a function of time and keep the count family for each bucket. If failures genuinely cluster, treat the cluster as the event: model cluster arrivals as counts and cluster size separately. If the counts are simply more spread out than a one-parameter count model can accommodate, a two-parameter count family such as the negative binomial absorbs the extra spread. Each of these is still a *choice of family driven by the process*, which is the actual interview skill being probed. ## Units, the quiet trap Rates and mean gaps are reciprocals, and every reciprocal drags a unit with it. A rate of 2 per week is a mean gap of half a week; if the interviewer asked in hours, that is 84 hours, not 0.5 and not 12. Candidates who write `1 / r` and stop, without carrying the unit through, produce answers wrong by a factor of 24 or 168. State the unit of the rate first, then invert. ## Sanity checks before you commit Plot the bucketed counts and check whether the shape is right-leaning with a hard floor at zero rather than a symmetric bell. Plot the gaps and check for a mass of very short gaps with a long thin tail rather than a hump around the average. Split the data by hour of day and see whether the rate is recognisably different. Two minutes of that work is what separates a fitted family from an assumed one.

  • If the outage rate doubles, what happens to the mean time between outages?
    It halves. The mean gap is the reciprocal of the rate, so going from 2 to 4 outages per week moves the mean gap from half a week to a quarter of a week, roughly 42 hours. The whole gap distribution compresses toward zero, it does not just shift.
  • What would you look for in the data to decide the constant-rate assumption is broken?
    Bursts: several events within minutes, then long silence. Bucketed counts with far more empty buckets and far more extreme buckets than a single-rate fit predicts. A rate that visibly differs by hour of day or day of week. Any of those means you model a time-varying rate or treat clusters as the unit of arrival.
  • You only have weekly counts, not timestamps. Can you still say anything about the gaps?
    You can estimate the rate from the counts and report an implied mean gap, but only under the assumption that arrivals are independent at a constant rate. You cannot verify the shape of the gap distribution, or detect clustering inside a week, without the timestamps. Say the assumption out loud rather than presenting the implied gap as measured.

It is one bus route described two ways: buses per hour, or minutes between buses. Same timetable, two readings of it.

saying these in an interview costs you the question

  • Treats counts and gaps as two unrelated, unconnected models
  • Fits a normal distribution to a weekly count of 0 to 3 outages
  • Assumes a constant rate when incidents clearly arrive in bursts
  • Confuses the rate per week with the mean gap in hours
  • Picks the family from habit rather than from what was measured

context

open as a page

Which distribution models the number of retries before a flaky request finally succeeds?

level: juniorimportance: should knowfreq 48%

basics

~20 s

The geometric distribution. It counts repeated independent attempts, each succeeding with the same probability p, up to the first success. It is not Poisson, because Poisson counts events inside a fixed window rather than trials.

open as a page

For disk time-to-failure, when is an exponential model wrong and a Weibull right?

level: seniorimportance: should knowfreq 34%

basics

~20 s

An exponential lifetime assumes the failure rate never changes with age, so an old disk is as likely to fail next month as a new one. When wear-out makes failures rise with age, use a Weibull.

open as a page

Insurance claim sizes have a long right tail; how do you choose between log-normal and Pareto?

level: seniorimportance: should knowfreq 44%

basics

~10 s

Take logs of the claim sizes: a symmetric bell after logging points to log-normal. A tail that traces a straight line on a log-log survival plot points to Pareto. A normal fits neither.

open as a page

Which family models a per-user conversion propensity that must lie between 0 and 1?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

The beta distribution. Its support is exactly the interval from 0 to 1, and two shape parameters let it be flat, bell-shaped or piled at both ends. A normal would put probability on impossible values.

open as a page