skip to content

Two million points drawn on one pair of axes form a solid silhouette — which remedies are honest, and what does each distort?

level: seniorimportance: must knowfreq 62%

answer

  1. the outline is not the rows
  2. every fix trades one honesty for another
  3. transparency saturates past a few dozen
  4. cell counts keep density, lose rows
  5. scattering falsifies position

basics

~20 s

At that density the eye reads the outline of the ink, not the rows beneath it. Each remedy trades one honesty for another: transparency saturates, cell counts lose individual rows, scattering falsifies position, and a sample must be labelled.

solid answer

~50 s

The picture has stopped being about rows: everything inside the shape is equally black, so the only thing readable is the boundary. The remedies are not ranked, they are traded. **Mark transparency** — drawing each mark partly see-through so overlaps accumulate into visible shade — works while overlaps are few and saturates back into a silhouette past a few dozen coincident marks. **Binned two-dimensional counts** — dividing the plane into cells and shading each by how many rows fell into it — report density honestly on a legend but dissolve individual rows, and the cell size changes the picture. **Scattering marks slightly** makes ties visible and makes every position wrong by the offset. **Drawing a sample** keeps the chart readable and must be labelled, because density can no longer be read off it. Say which one you used and what it cost.

go deeper

for a junior

Recall that heavy overlap means the eye is reading ink, not rows, and that a dark region does not tell you how many rows are in it.

for a middle

Explain why transparency has a working range: it encodes density through accumulated shade only until the plane saturates, after which the silhouette is back.

for a senior

Show the trade explicitly for the density in front of you — which remedy, what it distorts, and how you state that on the chart — rather than naming a favourite.

for a principal

The standard worth setting is that a chart says what it did to its rows: sampled, binned, offset or summarised. Without it no reviewer can tell an honest density picture from a flattering one.

## What the silhouette is A **mark** is the thing drawn for each row — here, one point per row. When two million of them land in a region a few hundred pixels across, the marks stop being separable. Every pixel inside the cloud is painted by many rows, so it looks exactly like a pixel painted by one. The only information the eye can still extract is the **outline**: roughly where the extremes are. Everything the chart was drawn to show — where the mass sits, whether there are two clumps or one, whether a relationship bends — is behind the ink. This matters more than it sounds, because a saturated cloud looks confident. It is dark, it is dense, it fills the frame, and nothing about it signals that its interior carries no information. There is a cheap first move before any of the remedies below: draw smaller marks. Shrinking the mark buys back some separability at no cost to honesty, and at high density it merely postpones the problem. ## The remedies, and what each one costs | remedy | what it preserves | what it distorts | where it stops working | |---|---|---|---| | **mark transparency** — each mark partly see-through, so overlaps accumulate | every row is still drawn; shade begins to stand for density | shade is not a calibrated count, and an isolated row becomes nearly invisible | past a few dozen coincident marks the plane saturates and the silhouette returns | | **binned two-dimensional counts** — one mark per cell, shaded by how many rows fell in it | density, honestly, with a legend a reader can convert back to counts | individual rows disappear; the cell size is a choice that changes the shape of the result | when the finding is a rare point rather than the mass, since a lone row shades one cell almost imperceptibly | | **scattering marks slightly** — a small random offset per mark | the visible number of marks, so rows sharing a position stop hiding each other | position: every mark is now wrong by up to the offset | as soon as the reader needs to read a value off the axis, or the offset approaches the resolution that matters | | **a summary mark** — a contour or a fitted shape describing where the mass is | the structure of the bulk, legibly | every individual row, and it imposes a shape in sparse regions where little supports one | when the point of the chart was the rows that do not fit the summary | | **drawing a sample** — plot a random subset | the visual character of a single row, and readability | the count: density can no longer be read off the page at all | when rare structure is the finding, because the sample may not contain it | The habit to build is not "reach for the transparent version". It is to say, out loud and in the caption, **which rows the picture is about and what the remedy did to them.** ## Why transparency is the one people over-trust It is the smallest edit and it looks principled: every row is still drawn, nothing is thrown away, and the accumulation of shade genuinely does encode density — while the density is low. The trap is that its useful range is narrow. Once a few dozen marks share a pixel, that pixel is as dark as it can get, and so is its neighbour, and the accumulated-shade argument has quietly stopped applying in exactly the region the reader most wants to understand. A chart can be honest in its sparse corners and saturated in its centre at the same time, which is worse than being uniformly bad, because the sparse corners lend it credibility. ## The category case, where the silhouette also lies about composition When colour is bound to a category and the marks overlap heavily, the visible colour in a dense region is mostly **the order the marks were painted in**: the last category drawn covers the others. A reader will interpret that region as dominated by whichever category that happens to be, and re-running the drawing with the categories in a different order produces a different conclusion from identical data. Transparency softens this and does not fix it, because saturation reasserts itself. The honest moves are the same as above — count per cell, or summarise — plus one more that costs nothing: check whether the conclusion survives reversing the drawing order. ## A short procedure 1. **Estimate the coincidence**, at least roughly: how many rows land in one pixel in the densest area? That number, not taste, picks the remedy. 2. **Under a handful per pixel**, smaller marks and transparency are enough. 3. **Well past that**, move to counts per cell, and put the count legend on the chart so the density is readable rather than impressionistic. 4. **If the finding is a rare row**, do not use anything that dissolves rows into density; use a summary for the mass and draw the rare rows on top of it as themselves. 5. **If you sampled, say so on the chart**, with the sample size. A sampled cloud that does not admit it is the one remedy that misleads without leaving any trace at all.

  • How would you decide between counts per cell and drawing a sample?
    By what the chart has to support. If the reader needs to compare how much is where, keep density and use counts per cell. If they need to see what a single row looks like — its shape, its position, its colour — a labelled sample keeps that and gives up density. Needing both usually means two charts, not one cleverer one.
  • The dense region of a chart is mostly one colour and colour is bound to a category. What would you check first?
    Whether the region is showing composition or drawing order. Overlapping marks paint over one another, so the last category drawn wins the pixels. Redraw with the order reversed: if the dominant colour changes, the picture was reporting the draw order. Then switch to a per-cell count of each category, or summarise.
  • Is scattering the marks ever the right answer?
    Yes, where the axis is discrete and the exact position carries nothing — values on a small set of levels, say, where the question is how many rows sit at each level. The offset falsifies a coordinate no one reads. It is the wrong answer wherever a reader will take a number off the axis, because then the offset is an error you introduced.

Stamping the same ink stamp onto one sheet of paper. A dozen impressions and the darkness tells you roughly how many; ten thousand and the sheet is black, and no lighter ink distinguishes ten thousand from a hundred thousand. At that point the only honest report is to rule the sheet into squares and write the number of impressions in each — which is exactly what counting per cell does, and exactly why it hands back density by giving up the individual impression.

saying these in an interview costs you the question

  • Claims transparency solves overplotting at any density
  • Says scattering the marks is free because the offset is small
  • Reads the silhouette's outline as the shape of the distribution
  • Draws a sample without labelling the chart as sampled
  • Picks counts per cell purely because it looks cleanest
  • Assumes the visible colour shows which category is commonest