skip to content

Three Dispositions, Three Biases

Dropping shrinks the sample and tilts it, a constant fill stacks a spike wherever that constant sits, and carrying the last value forward walks across boundaries. All three cost something.

on this pageshow

questions

4

A numeric column has 1,500 of 10,000 cells absent, and every one is filled with zero — what does that cost?

level: juniorimportance: must knowfreq 72%

answer

  1. a fill is a claim, not a repair
  2. the constant does not vanish into the data
  3. a spike lands where the constant sits
  4. near the centre narrows, far from it widens
  5. filled and measured become the same value

basics

~20 s

Filling with a constant stacks 1,500 identical values at one point, and where that point sits decides the damage: a zero fill in a column centred above zero drags the average down and leaves filled cells indistinguishable from measured zeros.

solid answer

~50 s

A constant fill does not remove the holes; it converts 1,500 of them into confident readings of zero. Three things follow. The column now carries a spike exactly where the constant sits, so 15% of it is one repeated value. The average moves toward that constant in proportion to the share filled, and the apparent variability moves too — **narrower** if the constant sits near the centre of the measured values, **wider** if it sits far from them, as zero does in a column of readings centred near 100. And nothing in the column afterwards separates a filled cell from a measured one, so every later check reads the substitution as data. Zero is the worst constant to pick precisely when zero is also a plausible reading. If you fill, record which cells you filled in a companion column so the decision stays visible.

go deeper

for a junior

Know that a fill writes a real value into the cell and that later steps cannot tell it from a measurement. Be able to say that zero is not a neutral choice when zero is also a plausible reading.

for a middle

Explain the arithmetic: 15% of the rows land on one point, the average moves toward the constant in proportion to the share filled, and the spread narrows or widens according to whether that point sits inside or outside the bulk of the data.

for a senior

Show how you would make the fill survive handover — a companion column recording which cells were filled — and how you would defend the constant from outside the column rather than from its own distribution.

for a principal

The tradeoff to argue is who the fill serves. Filling removes a choice from every downstream consumer in exchange for making their code simpler; leaving the holes keeps the choice with them at the price of every consumer having to handle absence.

## The fill is a claim, not a repair A cell with nothing recorded is itself a statement about the world: **no measurement exists for this row**. Writing a constant into that cell replaces the statement with a different one — *the measurement was this value* — and every step after the fill reads the second statement as fact. Nothing in the column says a substitution happened. That is the whole of the problem; everything else is arithmetic that follows from it. With 1,500 of 10,000 cells filled, 15% of the column is now a single repeated value. Call that value the **fill constant**. ## Where the constant sits decides the damage The most common thing said about a constant fill — that it pulls the spread in — is true of only one case. What actually happens depends on where the fill constant sits relative to the values that were genuinely measured. - **Near the centre of the measured values.** The column's average barely moves, because you added mass where the mass already was. But 15% of rows now sit on exactly one point, so the column looks far more tightly concentrated than the real data ever was, and any statement about variability is understated. - **Far from the centre.** Zero, in a column of readings centred near 100, is this case. The average is dragged toward zero roughly in proportion to the share filled, and the column gains a second cluster with nothing in between. Variability is now **overstated**, not understated, and a chart of the column shows two humps where the world has one. - **Outside the plausible range on purpose**, chosen so it supposedly cannot be mistaken for data. This is the worst of the three: it moves the average violently, it dominates any ordering of the column, and the first reader who has not been told the convention treats it as a reading. A value chosen to be obviously fake is only obviously fake to the person who chose it. | fill constant | effect on the average | effect on the spread | what a later reader sees | |---|---|---|---| | the column's own centre value | roughly unchanged | understated | a suspiciously peaked column | | zero, in a column centred well above zero | dragged down | inflated | two clusters, one sitting at zero | | an out-of-range stand-in code | dominated by the code | inflated severely | one extreme value treated as real | | no fill, holes left alone | unchanged | unchanged | absence, visibly absent | ## Zero is the trap, because zero is also a reading In most numeric columns, zero is a value the world can genuinely produce: no purchases that day, no rainfall that hour, a balance that really is empty. Once the fill runs, a substituted zero and a measured zero are the same number in the same column. There is no predicate that separates them — the **absence test**, the predicate that asks cell by cell whether a value is absent, now answers false for both, because the substituted cell holds a real value. Anyone counting zeros, filtering on zero, or reasoning about how often the reading is zero is answering a different question than they think. ## What the fill honestly buys Filling is not always wrong. It buys three things: 1. **Every later step gets a value**, so no operation downstream has to decide what absence means to it, and none of them can decide differently from one another. 2. **The table keeps its rows.** Nothing is deleted, so a sample that was hard to collect stays whole. 3. **Sometimes the constant is known rather than guessed.** A row with no charges genuinely has a charge total of zero; the absence there is a reporting artefact and the fill is a correction, not an invention. The test is whether you can defend the constant without reference to the rest of the column. ## Making the decision survive the handover The cheapest honest version of a fill is a second column beside the first that records, per row, whether that cell was filled. It costs one column and it restores the thing the fill destroyed: a reader can exclude filled rows from a trend, keep them in a total, and see how much of the answer rests on substituted values. Without it, the decision lives only in the code that ran, and the code is not what gets handed on. ## What varies between tools Two things differ enough that you should say which is in play before predicting a result. First, some surfaces take a fill constant **per column** and some accept one value for a whole table at once, in which case a constant that is sensible for one column is silently written into every other. Second, tools differ on whether the fill constant must match the column's **declared representation** — the one physical form every value in that column is stored as. Some refuse a constant of the wrong form, some quietly re-store the column in a wider one to accommodate it, and a column of arbitrary references — one whose cells are references to anything at all — accepts the constant without complaint and leaves you with two kinds of thing in one column.

  • Why is filling with the column's own centre value not the safe default it looks like?
    It preserves the average by construction, which is exactly why it feels safe, but it puts 15% of the rows on a single point and so understates the column's real variability. It also weakens every relationship the column had with the others, because the filled rows now carry a value chosen without reference to anything else in the row.
  • What makes zero a particularly dangerous constant in a column that can genuinely read zero?
    The substituted zero and the measured zero become the same value in the same column, so no predicate can separate them afterwards. Anyone counting zeros, filtering on zero, or estimating how often the reading is zero is silently answering a question about your fill instead of about the world.
  • When is a constant fill defensible rather than an invention?
    When the constant is known from outside the column rather than inferred from it — a row with no charges genuinely totals zero, an unticked optional box genuinely means not selected. The test is whether you can justify the value without looking at the rest of the column's distribution.

Filling holes with one constant is like patching a photograph with a single flat colour. The picture is complete and nothing is obviously torn, but 15% of it is now one shade the camera never saw, and nobody looking at the print can tell which part you painted.

saying these in an interview costs you the question

  • Zero is a neutral fill because it adds nothing
  • Filling makes the missing-data problem go away
  • Filling with the column's average leaves the distribution intact
  • A filled cell can always be recognised again later
  • Any constant is fine as long as it is outside the real range
open as a page

A 40-column table has roughly 3% of each column's cells absent; why does dropping every incomplete row delete most of the table?

level: middleimportance: must knowfreq 66%

basics

~20 s

Holes are spread across columns, so a row survives only if all 40 of its cells are present. At 3% per column and independent holes that is 0.97 to the fortieth power, about 30%. Row-wise removal compounds every column's absence rate into one much larger loss.

open as a page

Readings from 300 devices sit in one table and holes are filled by carrying the last value forward; what silently goes wrong?

level: seniorimportance: should knowfreq 54%

basics

~20 s

The carry walks down the rows in whatever order they currently sit, not down each device's own sequence, so one device's last reading fills another device's hole. Within a single device it also invents readings for the whole time that device was off.

open as a page

A companion column recording every filled cell is proposed as a pipeline-wide rule — what do you weigh before committing?

level: principalimportance: should knowfreq 42%

basics

~20 s

Recording each fill turns an invisible decision into data a consumer can act on. The costs are one column per filled column, a rule every downstream step must honour, and a boundary decision about where the recording stops.

open as a page