skip to content

A bar chart of revenue by region is drawn straight from 40,000 rows and shows five bars - what is each bar?

level: middleimportance: must knowfreq 62%

answer

  1. five marks, forty thousand rows
  2. somebody folded the rows
  3. which statistic, over how many rows
  4. reproduce one bar by hand

basics

~20 s

Each bar is a statistic the chart computed over the rows sharing that region, not a row. Which statistic - a total, a mean, a count - depends on the mark and the interface, so establish it before reading the bars.

solid answer

~50 s

Five bars from forty thousand rows means the rows sharing each region were folded into one number, and nobody wrote that aggregation - it came attached to the mark. Charting interfaces differ here: where a mark carries a default pre-draw statistic, handing it a table and two column names is enough and the aggregation is invisible; where the interface consumes coordinates you prepared, five bars are five numbers you produced yourself. So the first question is which kind drew this one. If the chart aggregated, three things are missing from the page: which statistic the bars show, since a total and a per-record mean rank the same regions differently; how many rows are behind each bar; and where any interval drawn over the bars came from. A bar over twelve rows and a bar over thirty thousand look equally authoritative.

go deeper

for a junior

Recall that one mark per category over a table of many rows means the rows were combined. Say what a bar stands for - all the records of that region reduced to one number - rather than reading it as a value from the data.

for a middle

Explain the mechanics: which statistic was used, that a total and a per-record mean order categories differently because the categories hold different row counts, and that some interfaces attach the statistic to the mark while others take numbers you prepared.

for a senior

Demonstrate the check. Reproduce a bar by hand to identify the statistic, put the row count per category back on the page, and flag a category built from a handful of records before anyone compares its height to a well-supported one.

for a principal

The tradeoff is who owns the default. Leaving each author on their tool's per-mark statistic means two charts of one dataset can make opposite claims with no one at fault; standardising the axis label and the row-count annotation costs discipline but buys comparability.

## A mark can carry a statistic of its own Five bars from a forty-thousand-row table means one thing before it means anything else: **the rows sharing each region were folded into a single number, and that number is what was drawn**. Nobody wrote the aggregation. It arrived attached to the mark, which is what makes this the characteristic failure of the surface - you are accountable for a quantity that no expression in your code produced. So the first job is to establish what the number is, and the second is to establish what is missing from the page. ## Which statistic, and why the choice changes the claim | what a bar can be | the question it answers | how it misleads | |---|---|---| | a total over the region's rows | how much altogether | a region of many small records towers over one of few large ones | | a mean over the region's rows | how much per record | twelve records sit at the same height as thirty thousand | | a count of the region's rows | how many records | reads as money when the axis label says only *revenue* | | a middle value | what a typical record looks like | hides that the total is carried by a handful of records | A total and a per-record mean routinely order the same five regions differently, because the regions hold different numbers of rows. If the chart picked between them for you, then **the chart picked which of two different claims the picture is making**, and a reader has no way to tell which from the marks alone. ## Two designs, and how to tell which one drew this The interfaces in this space genuinely disagree about mechanism, and the disagreement is not cosmetic: - **Where a mark carries a default pre-draw statistic** - the aggregation a chart performs on the rows before any mark is drawn - a table and two column names are enough. The mark aggregates, the row count appears nowhere, and whether the default is a total, a mean or a count is a property of that mark in that interface. - **Where the interface consumes prepared coordinates**, five bars mean five numbers that already existed. You made them, you know what they are, and nothing was computed behind your back. Two tests tell you which you have, and neither needs the source of the tool: 1. **Reproduce one bar.** Compute the candidate statistics for a single region yourself and see which reproduces the height. Where a total and a mean order two regions differently, the picture matches exactly one of them. 2. **Look for a quantity you never produced.** If the value axis carries a number that exists nowhere in the table or in your own steps, a statistic was computed for you. ## What the page does not show - **How many rows are behind each bar.** Bar height carries no weight of evidence with it, so a category built from twelve records is drawn with the same ink as one built from thirty thousand. - **Whether the rows behind a bar are alike.** One dominant record and a long tail of small ones produce the same bar as a tight cluster around the same mean. - **Where an interval came from.** Some interfaces compute an interval and draw it over each bar as part of the same default; others draw nothing unless you compute one and hand it in. A band that arrived by itself still has a rule behind it, and that rule is yours to state even though you did not choose it. ## Putting the missing facts back on the page - **Write the statistic into the axis label.** *Mean revenue per order* is a different chart from *total revenue*, and the label is the cheapest place to say which you drew. - **Show the row count per category** - as part of each tick label, or as a small companion chart of counts beside the bars. - **Treat a thinly supported category as provisional.** A summary over a dozen rows moves when one record changes, and nothing in the bar's height warns a reader of that. - **Decide whether the reader needs rows or a summary at all.** A mark that draws one per row answers a different question from a mark that draws one per category, and choosing between them is choosing what the reader is permitted to see. The habit that survives every interface is small: before reading a category chart, say out loud what one mark stands for and how many records are inside it. If you cannot answer both from the page, the page is incomplete, whoever drew it.

  • Two bars are nearly equal, but one summarises twelve rows and the other thirty thousand - what do you do?
    Say so on the page. Put the row count behind each category into the tick label or a small companion chart of counts, and treat the thinly supported category as provisional rather than comparable: one record changes it. Equal height is not equal evidence, and nothing about the mark communicates the difference.
  • How do you find out which statistic a chart computed, without reading the code that drew it?
    Reproduce one category by hand. Compute the candidates - total, mean, count, middle value - for a single region and see which matches the bar. Where two candidates order two regions differently, only one ordering can match the picture, which settles it in a single comparison.
  • Why do a total and a per-record mean rank the same regions differently?
    Because the categories hold different numbers of rows. A region of many small records can lead on a total and sit near the bottom on a per-record mean. Both are true statements about the same rows; they are answers to different questions, and the chart chose which question to answer.

saying these in an interview costs you the question

  • Reads a bar's height as a row's value rather than a computed statistic.
  • Assumes every charting interface totals the rows behind a category by default.
  • Compares two bars without knowing how many rows are behind each.
  • Treats an interval that appeared by default as needing no explanation.
  • Insists no aggregation happened because no aggregation was written.