Probability & Distributions
You will learn to reason with conditional probability and Bayes' theorem, work with the common distributions, and explain why the central limit theorem makes inference possible. This is the most heavily drilled area in data interviews — brainteasers, dice-and-coin questions, and 'what distribution models this?' all live here.
on this pageshowhide
explore
- Combinatorics & Event Rules11 questions
- Sample Spaces and Events5 questions
- Permutations and Combinations6 questions
- Conditioning & Bayes' Rule15 questions
- Independence vs Exclusivity5 questions
- Law of Total Probability5 questions
- Base Rates and False Positives5 questions
- Random Variables & Moments27 questions
- PMF, PDF and CDF5 questions
- Expectation and Variance6 questions
- Joint Distributions and Covariance5 questions
- Sums and Transformations5 questions
- Entropy and KL Divergence6 questions
- Distribution Families16 questions
- Bernoulli, Binomial and Poisson6 questions
- Uniform, Normal and Exponential5 questions
- Choosing a Distribution5 questions
- Limit Theorems & Simulation23 questions
- Law of Large Numbers6 questions
- Central Limit Theorem5 questions
- Monte Carlo Simulation6 questions
- Markov Chains6 questions
questions
92 · 5 sectionsWhat is the difference between a permutation and a combination?
basics
~10 sA permutation counts ordered arrangements; a combination counts unordered selections. Taking k items from n gives n!/(n-k)! permutations and n!/(k!(n-k)!) combinations. The extra k! divides out the orderings of each selected group.
In probability, what is a sample space and what counts as an event?
basics
~20 sA sample space is the set of all possible outcomes of a random experiment, listed so exactly one occurs. An event is any subset of it. Two dice give 36 outcomes; "the sum is 7" is a 6-outcome event.
In counting problems, how do you decide between n^k, n!/(n-k)! and n choose k?
basics
~10 sAnswer two yes/no questions: does order matter, and may items repeat? Ordered with repeats is n^k; ordered without repeats is n!/(n-k)!; unordered without repeats is n choose k; unordered with repeats is C(n+k-1, k).
Why does P(A or B) = P(A) + P(B) give the wrong answer when A and B overlap?
basics
~20 sAdding P(A) and P(B) counts outcomes in both events twice. The addition rule subtracts the overlap: P(A or B) = P(A) + P(B) - P(A and B). For one card, P(red or face) = 26/52 + 12/52 - 6/52 = 8/13.
Why is the chance of at least one six in four dice rolls not 4/6?
basics
~20 sAdding 1/6 four times double-counts rolls containing more than one six, and would give a probability above 1 for seven rolls. Use the complement: P(no six) = (5/6)^4, so P(at least one six) is about 0.518.
In the Monty Hall problem, why does switching doors win two-thirds of the time?
basics
~20 sYour first pick wins only one time in three, so two times in three the car is behind another door. The host, who knows where it is, opens a losing door and concentrates that 2/3 onto the single door left.
Why can two mutually exclusive events with nonzero probability never be independent?
basics
~20 sMutually exclusive means the two events cannot both happen, so P(A and B) = 0. Independence requires P(A and B) = P(A) times P(B), which is strictly positive when both probabilities are. Exclusivity therefore forces maximal dependence, not independence.
Using the law of total probability, what is P(defective) if line A makes 60% of units at 2% defective and line B 40% at 5%?
basics
~10 sThe overall defect rate is 3.2%. The law of total probability weights each line's defect rate by that line's share of production: 0.60 * 0.02 + 0.40 * 0.05 = 0.032.
A test with 99% sensitivity and 95% specificity flags a disease with 1% prevalence: how likely is a positive to be real?
basics
~10 sAbout 17%. In 10,000 people, 100 have the disease and 99 of them test positive, while 495 of the 9,900 healthy people also test positive. Rare conditions make positives mostly false positives.
How does the chain rule factor P(landed, started, finished) for a three-step signup funnel?
basics
~20 sP(landed) times P(started given landed) times P(finished given landed and started). The chain rule turns a joint probability into a product of conditionals, each measured on the survivors of the previous step, no independence needed.
What does Shannon entropy measure for a discrete distribution, in bits?
basics
~20 sShannon entropy is a distribution's average uncertainty: H = -sum p log2 p, the expected number of yes/no questions needed to pin down one draw. It is largest for equally likely outcomes and zero when one outcome is certain.
How do you compute the expected value of a $1 bet on a single roulette number?
basics
~20 sMultiply each outcome by its probability and add the pieces up. On a 38-pocket wheel a $1 straight-up bet nets +$35 with probability 1/38 and -$1 with probability 37/38, giving about -$0.053 per dollar staked.
For a continuous latency variable, what is the probability that response time is exactly 200.000 ms?
basics
~10 sExactly zero. Under a continuous model, any single point has zero width and therefore zero area under the density, so only intervals carry probability. Ask instead for P(199.5 < X <= 200.5).
What is the KL divergence KL(P||Q), and why is it not a distance metric?
basics
~20 sKL(P||Q) = sum P(x) log(P(x)/Q(x)) is the cost of describing draws from P as if they came from Q. It is never negative and zero only when P equals Q, but it is asymmetric, so it is a divergence, not a metric.
Why does linearity of expectation hold even when the random variables are dependent?
basics
~20 sLinearity of expectation is proved by regrouping a sum over the joint distribution, never by multiplying probabilities. E[X + Y] = E[X] + E[Y] holds for any variables with finite means. Independence is needed for products, not sums.
Test scores are normal with mean 100 and SD 15: by the 68-95-99.7 rule, what share exceeds 130?
basics
~20 sAbout 2.5%. A score of 130 is two standard deviations above the mean, so its z-score is 2; the rule puts roughly 95% of values within two SDs, leaving about 5% split evenly between the two tails.
What are the mean and variance of a binomial random variable with n trials and success probability p?
basics
~10 sA binomial count over n independent trials with success probability p has mean np and variance np(1-p). It is the sum of n Bernoulli(p) indicators, each contributing mean p and variance p(1-p).
For a Poisson count averaging 3 support tickets per hour, what is the probability of zero tickets in an hour?
basics
~20 sThe Poisson mass function is P(X = k) = e^(-λ) λ^k / k!. With λ = 3 tickets per hour, P(X = 0) = e^(-3), about 0.0498, so roughly a 5 percent chance of a completely quiet hour.
Which distributions model 'outages per week' and 'hours between outages' for the same incident stream?
basics
~10 sCounts in a fixed window, outages per week, are Poisson. The gaps between consecutive outages, hours between them, are exponential. Both describe one event stream: one view counts events, the other measures waiting times.
If bus waiting time is exponential with mean 10 minutes, why does waiting 10 minutes not shorten the expected remaining wait?
basics
~20 sThe exponential distribution is memoryless: P(T > s + t | T > s) = P(T > t). Time already spent waiting carries no information about what remains, so the expected remaining wait is still the full 10 minutes.
What does the Central Limit Theorem say about the average of many independent samples?
basics
~10 sThe Central Limit Theorem says that averaging many independent draws from almost any distribution with finite variance produces an average whose distribution is approximately normal, even when the individual observations are not remotely normal.
Does the law of large numbers make black due after a roulette wheel lands red eight times?
basics
~20 sNo. Spins are independent, so the chance of black is exactly what it was before the streak. The law of large numbers dilutes early results under a huge volume of later ones; it never reaches back to correct them.
What does the law of large numbers say about the sample mean as observations accumulate?
basics
~20 sThe law of large numbers says that for independent draws from one distribution with a finite mean, the average of the observations converges to that distribution's true expected value as the number of draws grows.
What does the Markov property mean for a chain that models users as trial, paid or churned?
basics
~20 sThe Markov property says the next state depends only on the current state, not on the path taken to reach it. A user in the paid state has the same transition probabilities regardless of how long ago they upgraded.
How would you estimate pi by Monte Carlo simulation with uniform random points?
basics
~20 sDraw many independent points uniformly in the unit square and count the fraction that land inside the quarter circle of radius 1. That fraction estimates pi/4, so multiply it by 4. Accuracy improves like 1/sqrt(n).