How can a histogram's bin width change the shape the data appear to have?
answer
- The analyst chooses the picture
- Smoothing dial with two ends
- Wide hides humps, narrow invents them
- Same width, shifted origin, different plot
- Bias-variance, same as any smoother
basics
~20 sBin width is a smoothing dial the analyst chooses. Wide bins merge separate clusters into one hump; narrow bins turn random sampling noise into spurious spikes and gaps. Where the bin edges start can also split or merge a cluster.
solid answer
~50 sA histogram counts observations into intervals, so the picture depends on two arbitrary choices: how wide the bins are and where the first edge sits. Draw the same 1,000 values with 5 bins and you may see a single smooth mound; draw them with 50 and the same data can show two clear humps plus jagged noise between them. Too wide hides real structure, too narrow shows structure that is only sampling variation. Bin placement matters independently: with the same width, shifting the origin can put a boundary through the middle of a cluster and split one hump into two, or straddle two clusters and merge them. The practical habit is to redraw at several bin widths before believing any shape, and to use a rule such as Freedman-Diaconis or Sturges as a starting point rather than an answer.
go deeper
Be ready to say that a histogram counts observations into intervals and that the number of bins is a choice, not a fact about the data. Knowing that too few and too many bins both mislead is enough here.
Explain both failure directions concretely and name at least one bin-width rule with the assumption it encodes. Be able to describe how shifting the bin origin at fixed width can split or merge a cluster.
Show the verification habit: how you confirm a mode is real before someone builds a decision on it, when you switch to an unbinned view, and how you keep bin edges consistent when comparing groups or time periods.
Own the standard. Decide what distribution views your organisation publishes and whether a single default bin count across dashboards is buying comparability at the cost of hiding structure in specific metrics.
## What a histogram actually does A histogram partitions the value axis into consecutive intervals called bins, counts how many observations fall into each, and draws a bar whose height (or, with unequal widths, whose area) is that count. The data are fixed; the partition is not. Bin width and the position of the first bin edge are choices made by whoever drew the plot, and different choices produce different pictures of the same numbers. That makes a histogram a *smoothing* device with a tuning parameter, and every smoothing parameter sits on a bias-variance tradeoff. ## Too wide: real structure disappears With very few, very wide bins, each bar aggregates a broad range of values. Detail inside a bar is lost by construction. Two genuine clusters, say response times around 40 ms for cache hits and around 400 ms for cache misses, can land inside the same wide bin or in two adjacent bins whose heights do not visibly separate. The plot then shows a single mound and the analyst concludes the data are unimodal, which is the wrong conclusion about the system: there are two populations, not one. This is the high-bias end. The estimate is stable across resamples, but it is stably wrong. ## Too narrow: noise becomes structure With very many very narrow bins, each bar holds only a handful of observations. Counts of 3, 0, 4, 1, 5 are what random sampling produces in that regime even from perfectly smooth data, and drawn as bars they look like spikes and holes. An inexperienced reader interprets those as multiple modes or as a gap in the data. Redraw the same distribution from a fresh sample and the spikes appear in different places, which is the giveaway. This is the high-variance end. The estimate is unstable and every wiggle is an artifact. ## Bin edges, not just bin width A subtler failure is bin *placement*. Fix the width and slide all the edges by half a bin. Observations that were split between two bars now fall in one, and vice versa. A tight cluster sitting astride a boundary is drawn as two mid-height bars; shift the origin and it becomes one tall bar. Real, published examples exist where a histogram of the same data with the same bin width looks unimodal under one anchor and bimodal under another. If a claimed feature of your data vanishes when you shift the origin by half a bin, it was an artifact of the grid, not a property of the data. ## Rules of thumb and what they assume Several formulas give a starting width for n observations: - **Square-root rule**: use about sqrt(n) bins. Crude, common as a default, ignores the data's spread and shape. - **Sturges' rule**: number of bins = ceiling(log2(n)) + 1. Derived assuming roughly normal data; it produces too few bins for large samples and over-smooths skewed data. - **Scott's rule**: bin width = 3.49 * s * n^(-1/3), where s is the sample standard deviation. Optimal for normal-shaped data, so it over-smooths when the data are stretched or clustered. - **Freedman-Diaconis rule**: bin width = 2 * (Q3 - Q1) * n^(-1/3), driven by the quartile spread rather than the standard deviation, which makes it less sensitive to a few extreme values. All of these are *starting points*. They target a single global width and they encode an assumption about shape, so none of them can be trusted to reveal multimodality on its own. The professional habit is to draw the histogram at three or four widths and report the feature only if it survives all of them. ## The alternative that has no bin choice If you want a view of the distribution with no arbitrary parameter, plot the **empirical CDF**. For a sample of n values, the empirical CDF at x is the fraction of observations less than or equal to x: F(x) = (count of observations <= x) / n. It is a step function that rises by 1/n at every observation, uses every data point exactly where it fell, and requires no binning or smoothing at all. You read any percentile straight off it: go to 0.9 on the vertical axis and across to the curve. Its cost is that shape is harder to see by eye. A flat stretch means a gap in the data and a steep stretch means a dense cluster, which is a less intuitive encoding than a hump, so multimodality reads as changes in slope rather than as separate peaks. Many analysts draw both: the histogram to see shape, the empirical CDF to check that the shape is not an artifact of the bins. ## What to say in an interview The complete answer names both knobs (width and edge placement), places width on a bias-variance tradeoff, gives the concrete failure in each direction, and finishes with the mitigation: vary the width, shift the origin, and cross-check against a representation that does not bin.
- You see two humps in a histogram. How do you check the bimodality is real?Redraw at several bin widths and shift the bin origin by half a width; a genuine feature survives all of them while a grid artifact does not. Then cross-check with a representation that does not bin, such as an empirical CDF, where two clusters show as two steep stretches separated by a flat one. Finally look for a variable that would explain two populations and split the data by it.
- Why does the Freedman-Diaconis rule use the quartile spread rather than the standard deviation?The standard deviation is computed from squared deviations, so a few very large values inflate it and the rule then picks bins far too wide for the bulk of the data. The distance between the first and third quartiles is a positional quantity that a handful of extreme observations barely move, so the resulting width tracks where most of the data actually sit.
- How would you compare the distributions of two groups of very different sizes?Do not compare raw counts, because the larger group's bars dominate purely by volume. Plot densities or relative frequencies so each group's bars sum to the same total, and use identical bin edges for both so the bars are comparable position by position. Overlaid empirical CDFs are often cleaner still, since they are already on a 0-to-1 scale.
Bin width is the zoom on a camera: too far out and two people merge into one blur, too far in and skin texture looks like a landscape.
saying these in an interview costs you the question
- Treats the default bin count as a property of the data
- Reports a bimodal shape without varying bin width
- Believes only bin width matters, not edge placement
- Says more bins always means a more accurate picture
- Compares two groups with raw counts and different bin edges