skip to content

PMF, PDF and CDF

How a distribution is written down: mass for discrete variables, density for continuous ones, and the CDF that links them and yields quantiles. Many candidates wrongly read a density as a probability.

on this pageshow

questions

5

For a continuous latency variable, what is the probability that response time is exactly 200.000 ms?

level: juniorimportance: must knowfreq 68%

answer

  1. a single point has no width
  2. probability is area over an interval
  3. zero area, not impossible
  4. F(200) minus F(200) is zero
  5. only intervals carry probability

basics

~10 s

Exactly zero. Under a continuous model, any single point has zero width and therefore zero area under the density, so only intervals carry probability. Ask instead for P(199.5 < X <= 200.5).

solid answer

~40 s

It is zero. For a continuous random variable, probability is area under the density, and a single point has no width, so `P(X = 200.000) = 0`. Equivalently, using the CDF, `P(X = 200) = F(200) - F(200^-) = 0` because a continuous CDF has no jump. That is not the same as impossible: the experiment always produces some value, and every value it produces had probability zero beforehand. To ask something answerable you have to nominate an interval, for example `P(199.5 < X <= 200.5) = F(200.5) - F(199.5)`. A useful consequence is that `P(X < 200)` and `P(X <= 200)` are equal for a continuous variable, so open versus closed endpoints never matter — unlike the discrete case, where they differ by the point mass at 200.

go deeper

for a junior

Be ready to answer "zero" without hesitating and to say why: a point has no width, so no area under the density. Then offer an interval version of the question, such as the probability of landing between 199.5 and 200.5 ms.

for a middle

An interviewer expects the CDF framing: the probability of a single value is the size of the jump in F there, and a continuous CDF has no jumps. Contrast with a discrete variable, whose PMF value is exactly that jump.

for a senior

Show that you know when the continuous model is a fiction. Recorded timings are rounded, so the measured variable is discrete on a grid; explain that the answer for the recorded value is the probability of the rounding interval, and that the continuous model is chosen for convenience.

for a principal

Own the modelling call: decide when a metric should be treated as continuous, as discrete counts, or as a mixture, and make sure downstream reporting and alerting are phrased over intervals and thresholds rather than exact values, so nobody builds a metric on a quantity whose probability is structurally zero.

## The short answer, then the reason Under a continuous model the answer is exactly 0, and the reason is structural rather than a matter of "very unlikely". A continuous random variable X is described by a density f, and the defining property of a density is that probability equals **area**: ``` P(a <= X <= b) = area under f between a and b ``` Now shrink the interval onto a single point. Set a = b = 200. The region under the curve has width 0 and finite height, so its area is 0. There is no probability left to assign. This holds for every point, including the one you eventually observe. ## The same statement in CDF language The cumulative distribution function is `F(t) = P(X <= t)`. Any point mass shows up as a **jump** in F: the probability of the single value 200 is the size of the jump of F at 200, `F(200) - F(200^-)`, where `F(200^-)` is the limit from the left. A continuous variable has a continuous CDF — no jumps — so every such difference is 0. This is also the cleanest way to see the contrast with a discrete variable. If X counts retries and `P(X = 3) = 0.12`, the CDF of X steps up by 0.12 at 3, and that step *is* the PMF value. A PMF assigns positive mass to individual points; a PDF never does. ## Probability zero is not impossibility Candidates often flinch here, because "probability 0" sounds like "cannot happen". The precise statements are: - The **impossible** event is the empty set: no outcome belongs to it. - A **probability-zero** event can be non-empty and can occur. Any single latency value has probability zero, yet the request does return some latency. The model says only that no individual value is singled out with positive mass; the mass lives on intervals. Analogously, picking a point at random on a line segment lands somewhere, but every particular point had probability zero of being the one. ## What to ask instead The question becomes well-posed the moment you name an interval or a threshold: - `P(199.5 < X <= 200.5) = F(200.5) - F(199.5)` — a band around 200. - `P(X > 500) = 1 - F(500)` — the tail beyond an SLO threshold. - `P(a < X <= b) = F(b) - F(a)` — the general form. The density value `f(200)` is *not* one of these. It is probability per unit of time, not a probability. It is useful for comparing relative plausibility near different values, and `f(200) * w` approximates `P(200 - w/2 < X <= 200 + w/2)` for a small width w, but on its own it answers nothing about the event "exactly 200". ## Open versus closed endpoints Because each point contributes 0, all four of these coincide for a continuous X: ``` P(a < X < b) = P(a <= X < b) = P(a < X <= b) = P(a <= X <= b) ``` For a discrete variable they differ by the PMF values at the endpoints, which is exactly why interview problems on discrete variables are fussy about `<` versus `<=` and continuous ones are not. ## The honest caveat about real data Real timers report rounded values — say whole milliseconds — so the recorded quantity is genuinely discrete and "exactly 200" has positive probability at the level of the *recording*. The continuous density is a model of the underlying quantity, chosen because it is analytically convenient and because the rounding grid is fine relative to the spread. A strong answer says both things: zero under the continuous model, and if you truly need the probability of the recorded value 200, you are asking about the interval the rounding maps onto, `P(199.5 < X <= 200.5)`. ## How to say it in an interview "Zero — a point has no width, so no area. Probability zero is not impossibility; the request still returns a value. If you want a number, give me an interval: `F(200.5) - F(199.5)`." Then, if there is time, add the CDF-jump framing and the discrete contrast.

  • Does probability zero mean the event is impossible?
    No. The impossible event is the empty set, which contains no outcome at all. A probability-zero event can be non-empty and can occur: the request does return some latency, and whatever value it returned had probability zero beforehand. Under a continuous model, every individual value is in that position, so "zero probability" cannot mean "cannot happen".
  • How would you rewrite the question so it has a non-zero answer?
    Name an interval instead of a point. For a band around 200 ms, `P(199.5 < X <= 200.5) = F(200.5) - F(199.5)`. For an SLO-style question, `P(X > 500) = 1 - F(500)`. Both are differences of CDF values, which is the general recipe: `P(a < X <= b) = F(b) - F(a)`.
  • For a continuous variable, does P(X < 200) differ from P(X <= 200)?
    No, they are equal, because they differ only by `P(X = 200)`, which is zero. So open and closed endpoints never change the answer for a continuous variable. For a discrete variable they do differ, by exactly the PMF value at 200, which is why discrete problems require care about strict versus non-strict inequalities.
  • Recorded latencies are rounded to whole milliseconds — does that change the answer?
    It changes the object you are talking about. The recorded value is discrete, so "exactly 200" has positive probability at the level of the recording, equal to the probability that the underlying latency fell in the rounding bin `199.5 < X <= 200.5`. The continuous model is an idealisation of the underlying quantity; the honest answer states both.

Picking a random spot on a one-metre ruler: the pencil lands somewhere, but the chance of hitting one exact mathematical point is zero. You can only bet on segments.

saying these in an interview costs you the question

  • Says the probability is small but non-zero
  • Reports the density value f(200) as the probability
  • Claims probability zero means the event is impossible
  • Distinguishes P(X < 200) from P(X <= 200) for a continuous variable
  • Divides one by a count of possible values to get an answer

context

open as a page

Why can a probability density function take values greater than 1 when a probability cannot?

level: middleimportance: must knowfreq 62%

basics

~20 s

A density is probability per unit of x, not a probability. Probability is the area under the curve, so a tall density over a narrow interval still has total area 1. The uniform density on [0, 0.5] equals 2 everywhere.

open as a page

How do you read the median and the p90 latency off a theoretical CDF F(t)?

level: middleimportance: should knowfreq 52%

basics

~10 s

Invert the CDF rather than reading heights: the median is the smallest t with F(t) at least 0.5, and the p90 the smallest t with F(t) at least 0.9.

open as a page

What normalising constant c makes f(x) = c*x a valid density on [0, 2]?

level: middleimportance: should knowfreq 45%

basics

~20 s

c = 1/2. The area under c*x from 0 to 2 is c times 2, and a density's total area must equal 1, so c = 1/2. The sign is fixed by requiring the density to be non-negative.

open as a page

How do you describe a customer spend variable with a point mass at 0 and continuous spend above it?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

As a mixed distribution: no single PMF or PDF describes it. Use the CDF, which jumps by the non-buyer share at 0 then rises smoothly, or a mixture of an atom and a spend distribution.

open as a page