How do you choose between statistics.mean, median, stdev and pstdev in Python?
answer
- Which summary survives one wild value
- Every point votes, or only the ordering
- Two spread functions, two divisors
- n minus one means it is a sample
- Too-few points raises StatisticsError
basics
~20 sstatistics.mean averages every value, so one outlier drags it; statistics.median returns the middle value and ignores how extreme the tail is. statistics.stdev divides by n-1 for a sample, statistics.pstdev by n for a whole population.
solid answer
~50 sReach for `statistics.mean` when the data is roughly symmetric and every point should count: it sums all values and divides by the count, so one extreme measurement shifts it noticeably. Reach for `statistics.median` when the data is skewed or contains outliers, because it sorts the values and takes the middle one (the average of the two middle values on an even-length input), so extremes change the ordering but not the answer. For spread, `statistics.stdev` applies Bessel's correction and divides the summed squared deviations by `n - 1`, which is what you want when the list is a sample drawn from a larger population; `statistics.pstdev` divides by `n` and is correct only when the list *is* the whole population. `statistics.variance` and `statistics.pvariance` are the same pair unsquared, and all of them raise `statistics.StatisticsError` when the input is too small.
code
pycon · 6 lines>>> import statistics
>>> salaries = [42_000, 45_000, 47_000, 50_000, 900_000]
>>> statistics.mean(salaries)
216800
>>> statistics.median(salaries)
47000go deeper
Be ready to name the four functions and say what each returns: mean averages, median takes the middle, stdev and pstdev measure spread with different divisors. Knowing that one outlier wrecks the mean but not the median is the point of the question.
Explain the mechanics: the n-1 divisor and why a sample needs it, what median does on an even-length list, and that these calls raise StatisticsError rather than returning a sentinel on too-small input.
Show judgement about which summary reaches a dashboard or an alert. Reporting a mean latency next to a median that disagrees with it is the finding, not a formatting choice, and guarding empty batches at the call site is production hygiene.
Own the convention across services: which summary the platform publishes for a given metric, whether stored series are samples or populations, and how using the wrong divisor or the wrong centre quietly changes the meaning of every downstream comparison.
## The module and the choice it forces `statistics` is the standard library's descriptive-summary module: pure Python, no dependency, working on any iterable of numbers. Its API splits along two axes that beginners routinely collapse into one. The first axis is *what kind of centre* you want, `mean` or `median`. The second is *what the data represents*: a sample (`stdev`, `variance`) or a complete population (`pstdev`, `pvariance`). Picking the wrong function on either axis produces a number that does not look wrong, which is exactly why the choice gets asked about. ## mean: every value votes `statistics.mean(data)` adds every value and divides by the count. Because every point contributes, one extreme value moves the result by roughly its own size divided by n. On five salaries where one is twenty times the rest, the mean lands above every ordinary value: ```pycon >>> import statistics >>> salaries = [42_000, 45_000, 47_000, 50_000, 900_000] >>> statistics.mean(salaries) 216800 ``` Two implementation details matter. `mean` does its arithmetic exactly, converting the input to `fractions.Fraction` internally, so it returns an `int` when the division comes out exact, a `float` when it does not, and a `Decimal` when fed `Decimal` input. That accuracy costs speed: `statistics.fmean(data)` converts to `float` up front and is far faster, and it accepts a `weights` argument. If what you actually want is an accurate *total* of many floats rather than an average, the tool next door on this subject is `math.fsum`. ## median: only the ordering votes `statistics.median(data)` sorts a copy and returns the middle value. On an even-length input it returns the average of the two middle values, which is why `statistics.median([1, 3, 5, 7])` is `4.0`, a value that is not in the data. When the answer must be a real data point, `statistics.median_low` and `statistics.median_high` return the lower or the upper of the two middles instead. Because only rank matters, replacing the largest value with something ten times larger does not move the median at all. That is the property you are buying: a summary that still describes the typical case when the tail is long. Request latencies with one timeout, response sizes with one huge payload, salaries with one founder in the list, all read honestly through the median and misleadingly through the mean. The rule of thumb to say out loud: roughly symmetric data where every point is meaningful, use `mean`; skewed data or known outliers, use `median`. And if the two disagree noticeably on the same list, that disagreement is itself the finding, so report both rather than picking the flattering one. ## stdev versus pstdev: the divisor is the whole difference Both compute the square root of the average squared deviation from the mean, and they differ only in what they divide by. `statistics.pstdev` divides the summed squared deviations by `n`. `statistics.stdev` divides by `n - 1`, known as Bessel's correction, because when the data is a *sample* drawn from a larger population the deviations are measured from the sample's own mean and therefore systematically understate how spread out the population is. `statistics.variance` and `statistics.pvariance` are the same pair without the square root. ```pycon >>> data = [2, 4, 4, 4, 5, 5, 7, 9] >>> statistics.pstdev(data) 2.0 >>> statistics.stdev(data) 2.138089935299395 ``` The gap shrinks as n grows and is negligible past a few hundred points, but at n = 8 it is about 7%. Choose by what the list literally *is*: every request the service handled in the last minute is the whole population of that minute, so `pstdev`; a thousand sampled requests standing in for millions is a sample, so `stdev`. Both functions accept a precomputed centre, `xbar=` for `stdev` and `mu=` for `pstdev`, so you do not recompute the mean when you already hold it; pass a centre that does not belong to the data and the result is silently wrong rather than an error. ## Errors, and the calls next to these four Every one of these functions raises `statistics.StatisticsError` instead of returning a sentinel when the input is too small. `statistics.mean([])` reports that it requires at least one data point, and `statistics.stdev([x])` reports that it requires at least two, because the `n - 1` divisor would be zero. `statistics.pstdev([x])` is legal and returns `0.0`. Guard the empty-list case at the call site rather than wrapping every call in `try`. Nearby in the same module and worth knowing so you do not reimplement them: `statistics.mode` returns the most common value, `statistics.multimode` returns every tied winner as a list, and `statistics.quantiles(data, n=4)` returns the cut points that split the data into n equally sized groups, which is what people often mean when they ask for the median and the quartiles together.
- What does statistics.fmean give you that statistics.mean does not?Speed, and weighting. `statistics.mean` works through `fractions.Fraction` internally so it stays exact and can return an `int`, a `float` or a `Decimal` depending on the input type. `statistics.fmean` converts everything to `float` up front, which makes it substantially faster on large lists and always returns a `float`, and it accepts a `weights` argument for a weighted average. Use `fmean` for bulk float data and `mean` when the input type or exactness matters.
- On an even-length list, what does statistics.median return, and when would median_low be the better call?It returns the average of the two middle values, so `statistics.median([1, 3, 5, 7])` is `4.0`, a number that does not appear in the data. Use `statistics.median_low` (or `statistics.median_high`) when the answer has to be an actual observation, for example when you will look the value up elsewhere, or when averaging the two middles would be meaningless for the kind of value you are holding.
- Why does statistics.stdev raise on a one-element list when statistics.pstdev does not?`statistics.stdev` divides by `n - 1`, which is zero for a single point, so it raises `statistics.StatisticsError` saying it requires at least two data points. `statistics.pstdev` divides by `n`, and the spread of a one-element population is genuinely zero, so it returns `0.0`. In code that summarizes arbitrary input, check the length before calling rather than letting the exception define the contract.
The mean is everyone in the room shouting a number and you averaging the noise; the median is asking them to line up in order and reading the one in the middle. The loudest shouter changes the first answer and not the second.
saying these in an interview costs you the question
- Treats mean and median as interchangeable summaries
- Thinks stdev and pstdev differ only in speed
- Picks pstdev on a sample because the number is smaller
- Expects median to always return a value present in the data
- Assumes statistics.mean of an empty list returns 0
- Believes statistics.mode raises whenever values tie