skip to content

A distribution chart shows one peak at the tool's default bucket width and three peaks at half that width - which shape is real?

level: juniorimportance: must knowfreq 68%

answer

  1. the picture is not the rows
  2. one mark per bucket, not per row
  3. who chose the bucket width
  4. redraw at several widths

basics

~20 s

Neither shape is a property of the data alone. The chart counted rows into buckets before drawing, so the picture is of those counts, and the bucket width - which the tool chose if you did not - decides how many peaks appear.

solid answer

~40 s

A distribution chart of one numeric column does not draw your rows. It computes a statistic first: it cuts the value range into buckets and counts how many rows fall into each, then draws one mark per bucket. Two settings decide those counts - the bucket width, and where the first boundary sits - and if you did not state them the tool applied a rule of its own. Wide buckets merge distinct groups into one hump; narrow buckets turn ordinary variation into peaks. There is no correct default, because the right width depends on how many rows there are, how spread out they are and how finely the values were recorded. So redraw at several widths, go and look at the rows under any interesting peak, and put the width on the chart.

go deeper

for a junior

Know that a distribution chart counts rows into buckets and draws those counts. Saying that a bar's height is a count of rows in that bucket, and that the bucket width was a setting somebody or something chose, is the whole of the junior answer.

for a middle

Explain the trade the width makes: wide buckets merge distinct groups into one hump, narrow buckets turn small count wobbles into peaks, and no rule is right for every column. Say what you would redraw in order to test a shape.

for a senior

Show the habit rather than the theory. Redraw at a spread of widths before quoting a shape, count the rows under the interesting region, check the recording precision, and put the width on the chart so a reviewer can reproduce your reading of it.

for a principal

Treat an unstated bucket width as a reporting risk, not a plotting detail: decisions get taken on modes that vanish at another width. Decide whether charts that leave your team must carry their settings, and who maintains that.

## The chart aggregated before it drew A distribution chart of a single numeric column does not draw your rows. It performs an aggregation first - **the chart's own statistic, computed before any mark exists** - by cutting the value range into buckets of some width, counting how many rows fall into each bucket, and drawing one mark per bucket whose height is that count. Forty thousand rows become perhaps thirty numbers, and everything you read off the picture is a statement about those thirty numbers. Two settings decide what those numbers are: - **the bucket width** - how wide each bucket is, equivalently how many buckets span the data; - **where the buckets start** - the value the first boundary sits on, which pushes every row near a boundary into one side or the other. If you stated neither, the tool chose both with a rule of its own. That is the whole answer to *which shape is real*: **neither shape is a property of the data alone**, because each is a set of counts under a different bucketing, and both sets of counts are correct. ## Why no default can be correct Bucket width trades two errors against each other, and the exchange rate depends on the column in front of you: - **Too wide and distinct groups merge.** Two clusters of values forty units apart fall in one bucket, and the chart shows a single hump that no amount of staring recovers. - **Too narrow and ordinary variation becomes structure.** With a few hundred rows spread across many buckets, each count is small and moves by one or two for no reason worth naming; the eye reads those wobbles as modes. - **Too narrow against the recording precision and you get a comb.** A column rounded to whole units cannot place a value at 3.5, so buckets that can never catch a recorded value sit empty between tall neighbours. That pattern comes from the bucketing meeting the recording, not from the world. Automatic rules exist and every tool applies one, because a chart has to appear. They are heuristics built on assumptions about how many rows there are and roughly what shape the values take. Your column may not satisfy those assumptions, and nothing on the chart will say so. | what you see | what it may be a property of | how to tell | |---|---|---| | one smooth hump | the data, or buckets wide enough to merge groups | halve the width and see whether it splits | | several sharp peaks | the data, or too few rows per bucket | widen the width; count the rows under one peak | | tall and empty buckets alternating | the recording precision against the width | check the distinct values; widen to a multiple of the unit | | a gap in the middle | the data, or where the first boundary landed | shift the start by half a bucket | ## Reading a shape honestly 1. **Redraw at a spread of widths** - at least one clearly wider and one clearly narrower than the tool's pick. A feature that survives them all is robust to the setting, which is worth saying out loud; one that appears at a single width is a property of that width until something else establishes it. 2. **Go to the rows.** Under an interesting peak there are specific records. Count them, and check whether a handful of repeated values, one import batch or one test account produced the whole feature. 3. **Check the recording granularity** before believing any fine structure at all. 4. **State the width on the chart**, in the subtitle or the caption. A shape whose width is unstated cannot be redrawn by a reader and therefore cannot be checked by one. ## What the chart does not tell you - **How many rows it was built from.** You can sum bucket heights by eye only when the vertical axis carries counts; some charts rescale that axis so the total area comes to one, which is another decision the chart made and which erases the sample size completely. - **The individual values.** Every row inside a bucket has become a tally mark. Exact duplicates, ties on a boundary and the single extreme value in the tail all look alike. - **Which rows.** The picture carries no way back to the records, which is why going to the rows is a step and not a flourish. ## The same trap under another name Replace the buckets with a smooth curve fitted over the values and the problem has not gone away - it has been renamed. Each part of such a curve averages over a neighbourhood of nearby values, and the width of that neighbourhood is exactly as unchosen as the bucket width was: wide erases structure, narrow invents it. The curve is also more persuasive than a row of bars, because smoothness reads as a finding rather than as a setting. Ask it the same two questions: who chose the width, and does the feature survive another one?

  • The same column is drawn as a smooth curve instead of buckets - which setting are you accountable for now?
    The bucket width is replaced by the width of the neighbourhood each part of the curve averages over. Wide smooths real structure away, narrow follows noise, and it is just as unchosen if you did not set it. The curve also usually hides how many rows produced it, which the bucket heights at least implied.
  • Where does a distribution chart tell you how many rows went into it?
    Usually nowhere. If the vertical axis carries counts you can sum them by eye; if it has been rescaled so the total area is one, the sample size is gone entirely. When the row count matters to the claim, write it into the chart or its caption yourself rather than assuming a reader will ask.
  • A column recorded to the nearest whole unit produces alternating tall and empty buckets - what happened?
    The bucket width fell below the granularity of the recorded values, so some buckets sit between the values that can actually occur and can never catch one. It is an artefact of the bucketing against the recording precision. Widen the buckets to a whole multiple of the recording unit and the comb disappears.

saying these in an interview costs you the question

  • Treats the shape at the default bucket width as the data's own structure.
  • Assumes the automatic bucket rule is appropriate for any column it is given.
  • Reports a peak that appears at only one bucket width as a finding.
  • Thinks a distribution chart draws one mark per row of the table.
  • Picks the width that shows the hoped-for effect and never states it.