A companion column recording every filled cell is proposed as a pipeline-wide rule — what do you weigh before committing?
answer
- a fill is a decision, not a measurement
- the record must travel with the data
- per column, per row, or not at all
- leaving the holes needs no bookkeeping
- retrofitting means reprocessing history
basics
~20 sRecording each fill turns an invisible decision into data a consumer can act on. The costs are one column per filled column, a rule every downstream step must honour, and a boundary decision about where the recording stops.
solid answer
~50 sThe case for is that a filled cell and a measured one are the same value, and the person who filled is not the person who reads. A companion column is the only record that travels with the data rather than living in code nobody reads, and it lets each consumer decide per use — exclude filled rows from a trend, keep them in a total. The case against is real cost: the column count roughly doubles for tables where most columns can be filled, every join, schema and write has to carry them, and consumers who ignore them impose the cost without taking the benefit. A rule that is not enforced is worse than no rule, because it makes unmarked data look trustworthy. Settle it before the pipeline is written — retrofitting means every historical row is unmarked, so consumers face two regimes or a reprocessing job.
go deeper
Know that once a cell is filled nothing in the data says so, and that a second column recording the fill is how that fact is kept. Recognise that leaving the holes alone is a legitimate option.
Explain the concrete costs on both sides: the column count, the discipline every step must keep, and what a consumer can do with the marker that they cannot do without it.
Show where the rule is enforced rather than merely agreed, how far the marker propagates through derived columns, and what happens to data written before the rule existed.
Argue the real trade: the pipeline making an unrecorded claim on behalf of consumers who cannot see it, against a convention nobody enforces making unmarked data look trustworthy. A team that cannot commit to enforcement should leave the holes and say so.
## Why this is a standing decision rather than a per-job one Every fill destroys the same piece of information: the fact that nothing was recorded. Once the value is written, the cell is an ordinary value and no predicate distinguishes it. The decision about whether that fact is worth preserving cannot sensibly be made job by job, because the consumer who needs it is downstream of a pipeline they did not write, reading a table whose provenance they cannot see. Either the convention holds across the pipeline or it is worthless — a marker present on some columns and absent on others tells a reader nothing, because they cannot distinguish *was not filled* from *nobody marks this one*. ## The case for - **It is the only record that travels with the data.** Code, notebooks, tickets and commit messages all stay behind when a table is handed on, exported or joined into something else. - **It moves the choice to the consumer.** One reader wants filled rows out of a trend line; another wants them in a total so the total still ties to the source. With a marker both are possible from the same table; without it, whichever the pipeline chose is imposed on everyone. - **It makes the size of the assumption visible.** A consumer can ask what share of a result rests on filled values. That number is often the most important thing about a report and is otherwise unknowable. - **It makes the fill auditable after the fact.** A wrong constant chosen a year ago can be located and undone if you know which cells it touched. ## The case against - **Width.** One marker per filled column roughly doubles the column count of any table where most columns can arrive with holes. That is paid in every read, every write, every schema, every transfer. - **Discipline.** The rule only works if every step that fills also marks, and every step that merges tables carries the markers through. One step that does not silently converts marked data into unmarked data, which is worse than never having marked, because unmarked now reads as *nothing was filled*. - **Consumers who ignore it.** If most readers never look at the markers, the cost is paid for a minority — a real argument, though it says more about who should pay than about whether the record should exist. - **Where it stops.** A value derived from a filled value is itself tainted. Propagating the marker through every transform is genuinely hard, and a marker that only covers the raw layer promises more than it delivers. ## Shapes that trade the costs differently | shape | cost | what it gives up | |---|---|---| | one boolean column per filled column | doubles the width | nothing; it is the complete form | | one column per row naming the filled fields | one extra column total | a consumer must parse it rather than filter directly | | a row-level count of filled cells | one narrow column | which fields were filled | | mark only judgment fills, not mechanical ones | small | a reader must know which fills count as mechanical | | no fill at all, holes left alone | none | requires every consumer to handle absence | That last row is the one most often forgotten, and it is frequently the right answer. **Leaving the holes alone needs no bookkeeping**, because nothing was destroyed. It is the disposition that keeps the choice with the reader at the cost of requiring every consumer to represent and handle absence — which, depending on what those consumers are, may be no cost at all or may be the entire reason the fill was proposed. Filling is a decision to take a choice away from downstream in exchange for making their code simpler. Whether that trade is right is exactly what is being argued here. ## How to settle it 1. **Ask who the consumers are, and whether they can represent absence.** If they can, the fill is buying convenience rather than correctness, and the default should be to leave the holes. 2. **Separate judgment fills from known-value fills.** A charge total of zero on a row with no charges is a correction and needs no marker. A chosen constant, a centre value, or a carried observation is a judgment and does. 3. **Decide the propagation rule explicitly.** Either markers cover only the layer where filling happens — and you say so — or every derived column carries the taint, and you accept the cost of computing it. 4. **Enforce it where it cannot be skipped.** A rule that relies on reviewers remembering degrades to noise within a year; one checked at the boundary where tables are published holds. 5. **Decide before the pipeline exists.** Adding the rule later leaves every historical row unmarked and therefore ambiguous, which means either reprocessing history or carrying two regimes and explaining the cutover to every consumer forever. ## The question underneath The strongest version of the argument is not about columns at all. It is about whether the pipeline is allowed to make an unrecorded claim on behalf of people who cannot see it being made. Framed that way, the width cost is usually the cheaper side — but the honest counter is that a convention nobody enforces makes the data look more trustworthy than it is, and a team that cannot commit to enforcement is better off leaving the holes and saying so.
- When is leaving the holes alone better than filling and recording it?When every consumer can represent and handle absence, and when the absence itself carries information — a gap that says the device was not reporting, or the question was not answered. Leaving it costs nothing, destroys nothing and keeps the choice with the reader; the fill is then buying convenience rather than correctness.
- Why is one row-level record of which fields were filled sometimes preferred to one column per filled column?It holds the table's width flat when many columns can arrive with holes, which matters for storage, transfer and every schema the table passes through. The price is that a consumer has to parse the record rather than filter on a plain condition, so casual readers are less likely to use it.
- Why must this rule be settled before the pipeline is written?Because adding it later leaves every historical row unmarked and therefore ambiguous: a reader cannot tell an unfilled cell from a cell filled before the rule existed. The choices are then to reprocess history or to carry two regimes and explain the cutover to every consumer indefinitely.
- What breaks the value of the convention fastest?A single step that fills without marking, or a merge that drops the markers while keeping the values. From then on unmarked data no longer means nothing was filled, and the convention has become a claim the data cannot support — which is worse than never having made it.
saying these in an interview costs you the question
- Record every fill everywhere; more metadata is always better
- Downstream consumers can work out which cells were filled
- Filling is a technical detail, not a decision worth recording
- A comment in the pipeline code is a sufficient record
- Leaving the holes alone is never acceptable in production
- Markers on the raw layer cover the derived columns too