skip to content

Probability & Distributions

You will learn to reason with conditional probability and Bayes' theorem, work with the common distributions, and explain why the central limit theorem makes inference possible. This is the most heavily drilled area in data interviews — brainteasers, dice-and-coin questions, and 'what distribution models this?' all live here.

on this pageshow

explore

questions

92 · 5 sections

What is the difference between a permutation and a combination?

level: juniorimportance: must knowfreq 82%
basics
~10 s

A permutation counts ordered arrangements; a combination counts unordered selections. Taking k items from n gives n!/(n-k)! permutations and n!/(k!(n-k)!) combinations. The extra k! divides out the orderings of each selected group.

open as a page

In probability, what is a sample space and what counts as an event?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A sample space is the set of all possible outcomes of a random experiment, listed so exactly one occurs. An event is any subset of it. Two dice give 36 outcomes; "the sum is 7" is a 6-outcome event.

open as a page

In counting problems, how do you decide between n^k, n!/(n-k)! and n choose k?

level: middleimportance: must knowfreq 67%
basics
~10 s

Answer two yes/no questions: does order matter, and may items repeat? Ordered with repeats is n^k; ordered without repeats is n!/(n-k)!; unordered without repeats is n choose k; unordered with repeats is C(n+k-1, k).

open as a page

Why does P(A or B) = P(A) + P(B) give the wrong answer when A and B overlap?

level: middleimportance: must knowfreq 68%
basics
~20 s

Adding P(A) and P(B) counts outcomes in both events twice. The addition rule subtracts the overlap: P(A or B) = P(A) + P(B) - P(A and B). For one card, P(red or face) = 26/52 + 12/52 - 6/52 = 8/13.

open as a page

Why is the chance of at least one six in four dice rolls not 4/6?

level: juniorimportance: should knowfreq 56%
basics
~20 s

Adding 1/6 four times double-counts rolls containing more than one six, and would give a probability above 1 for seven rolls. Use the complement: P(no six) = (5/6)^4, so P(at least one six) is about 0.518.

open as a page

In the Monty Hall problem, why does switching doors win two-thirds of the time?

level: juniorimportance: must knowfreq 64%
basics
~20 s

Your first pick wins only one time in three, so two times in three the car is behind another door. The host, who knows where it is, opens a losing door and concentrates that 2/3 onto the single door left.

open as a page

Why can two mutually exclusive events with nonzero probability never be independent?

level: juniorimportance: must knowfreq 80%
basics
~20 s

Mutually exclusive means the two events cannot both happen, so P(A and B) = 0. Independence requires P(A and B) = P(A) times P(B), which is strictly positive when both probabilities are. Exclusivity therefore forces maximal dependence, not independence.

open as a page

Using the law of total probability, what is P(defective) if line A makes 60% of units at 2% defective and line B 40% at 5%?

level: juniorimportance: must knowfreq 78%
basics
~10 s

The overall defect rate is 3.2%. The law of total probability weights each line's defect rate by that line's share of production: 0.60 * 0.02 + 0.40 * 0.05 = 0.032.

open as a page

A test with 99% sensitivity and 95% specificity flags a disease with 1% prevalence: how likely is a positive to be real?

level: middleimportance: must knowfreq 78%
basics
~10 s

About 17%. In 10,000 people, 100 have the disease and 99 of them test positive, while 495 of the 9,900 healthy people also test positive. Rare conditions make positives mostly false positives.

open as a page

How does the chain rule factor P(landed, started, finished) for a three-step signup funnel?

level: middleimportance: must knowfreq 62%
basics
~20 s

P(landed) times P(started given landed) times P(finished given landed and started). The chain rule turns a joint probability into a product of conditionals, each measured on the survivors of the previous step, no independence needed.

open as a page

What does Shannon entropy measure for a discrete distribution, in bits?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Shannon entropy is a distribution's average uncertainty: H = -sum p log2 p, the expected number of yes/no questions needed to pin down one draw. It is largest for equally likely outcomes and zero when one outcome is certain.

open as a page

How do you compute the expected value of a $1 bet on a single roulette number?

level: juniorimportance: must knowfreq 84%
basics
~20 s

Multiply each outcome by its probability and add the pieces up. On a 38-pocket wheel a $1 straight-up bet nets +$35 with probability 1/38 and -$1 with probability 37/38, giving about -$0.053 per dollar staked.

open as a page

For a continuous latency variable, what is the probability that response time is exactly 200.000 ms?

level: juniorimportance: must knowfreq 68%
basics
~10 s

Exactly zero. Under a continuous model, any single point has zero width and therefore zero area under the density, so only intervals carry probability. Ask instead for P(199.5 < X <= 200.5).

open as a page

What is the KL divergence KL(P||Q), and why is it not a distance metric?

level: middleimportance: must knowfreq 66%
basics
~20 s

KL(P||Q) = sum P(x) log(P(x)/Q(x)) is the cost of describing draws from P as if they came from Q. It is never negative and zero only when P equals Q, but it is asymmetric, so it is a divergence, not a metric.

open as a page

Why does linearity of expectation hold even when the random variables are dependent?

level: middleimportance: must knowfreq 76%
basics
~20 s

Linearity of expectation is proved by regrouping a sum over the joint distribution, never by multiplying probabilities. E[X + Y] = E[X] + E[Y] holds for any variables with finite means. Independence is needed for products, not sums.

open as a page

Test scores are normal with mean 100 and SD 15: by the 68-95-99.7 rule, what share exceeds 130?

level: juniorimportance: must knowfreq 82%
basics
~20 s

About 2.5%. A score of 130 is two standard deviations above the mean, so its z-score is 2; the rule puts roughly 95% of values within two SDs, leaving about 5% split evenly between the two tails.

open as a page

What are the mean and variance of a binomial random variable with n trials and success probability p?

level: juniorimportance: must knowfreq 78%
basics
~10 s

A binomial count over n independent trials with success probability p has mean np and variance np(1-p). It is the sum of n Bernoulli(p) indicators, each contributing mean p and variance p(1-p).

open as a page

For a Poisson count averaging 3 support tickets per hour, what is the probability of zero tickets in an hour?

level: juniorimportance: must knowfreq 64%
basics
~20 s

The Poisson mass function is P(X = k) = e^(-λ) λ^k / k!. With λ = 3 tickets per hour, P(X = 0) = e^(-3), about 0.0498, so roughly a 5 percent chance of a completely quiet hour.

open as a page

Which distributions model 'outages per week' and 'hours between outages' for the same incident stream?

level: middleimportance: must knowfreq 70%
basics
~10 s

Counts in a fixed window, outages per week, are Poisson. The gaps between consecutive outages, hours between them, are exponential. Both describe one event stream: one view counts events, the other measures waiting times.

open as a page

If bus waiting time is exponential with mean 10 minutes, why does waiting 10 minutes not shorten the expected remaining wait?

level: middleimportance: must knowfreq 64%
basics
~20 s

The exponential distribution is memoryless: P(T > s + t | T > s) = P(T > t). Time already spent waiting carries no information about what remains, so the expected remaining wait is still the full 10 minutes.

open as a page

What does the Central Limit Theorem say about the average of many independent samples?

level: juniorimportance: must knowfreq 85%
basics
~10 s

The Central Limit Theorem says that averaging many independent draws from almost any distribution with finite variance produces an average whose distribution is approximately normal, even when the individual observations are not remotely normal.

open as a page

Does the law of large numbers make black due after a roulette wheel lands red eight times?

level: juniorimportance: must knowfreq 70%
basics
~20 s

No. Spins are independent, so the chance of black is exactly what it was before the streak. The law of large numbers dilutes early results under a huge volume of later ones; it never reaches back to correct them.

open as a page

What does the law of large numbers say about the sample mean as observations accumulate?

level: juniorimportance: must knowfreq 78%
basics
~20 s

The law of large numbers says that for independent draws from one distribution with a finite mean, the average of the observations converges to that distribution's true expected value as the number of draws grows.

open as a page

What does the Markov property mean for a chain that models users as trial, paid or churned?

level: juniorimportance: must knowfreq 82%
basics
~20 s

The Markov property says the next state depends only on the current state, not on the path taken to reach it. A user in the paid state has the same transition probabilities regardless of how long ago they upgraded.

open as a page

How would you estimate pi by Monte Carlo simulation with uniform random points?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Draw many independent points uniformly in the unit square and count the fraction that land inside the quarter circle of radius 1. That fraction estimates pi/4, so multiply it by 4. Accuracy improves like 1/sqrt(n).

open as a page