A ratio column has holes: some rows were never measured, others divided zero by zero — can you still tell them apart?
answer
- two causes, possibly one pattern
- the format reserves it for undefined results
- borrowed as the missing marker too
- one predicate cannot split them
- evidence lives in the computation, not the column
basics
~20 sOnly under a design that carries two distinct markers. Where the single absence marker is the floating-point format's reserved pattern, an unrecorded observation and an undefined arithmetic result land on the same pattern and no predicate can separate them.
solid answer
~40 sIt depends on how many markers the design has. A tool whose absence marker is the pattern the floating-point format reserves for undefined results is reusing, for "nobody measured this", the very pattern that arithmetic produces when it has no answer. Both land in the same cell, the absence test answers true for both, and the column keeps no evidence of which is which. A tool that defines its own never-recorded marker *and* leaves the format's undefined-arithmetic value as itself has two distinct objects and two predicates, so the two causes stay separable. The distinction is worth caring about because the two demand different responses: a gap is a question for whoever produced the data, an undefined result is a bug or a missing guard in your own expression.
go deeper
Recall that a hole in a computed column has more than one possible origin: the source may have had nothing, or the arithmetic may have had no answer. They are not the same problem.
Explain the mechanism — where the absence marker is the pattern the floating-point format reserves for undefined results, both causes write the same pattern and one predicate answers true for both.
Show the diagnostic discipline: on a single-marker design the causal information has to be preserved at computation time, by guarding the expression or keeping the inputs, because the finished column cannot be interrogated.
The call to own is where in a pipeline undefined results are allowed to become absence at all — permitting it makes every downstream hole count ambiguous, and forbidding it costs a guard on every division a team writes.
## Two holes that look identical A hole in a computed column has at least two quite different origins: - **nothing was recorded** — the source never had a reading for that row, and the gap arrived with the data; - **the arithmetic had no answer** — the expression ran, and for that row it was undefined, most commonly a zero divided by a zero. These demand opposite responses. The first is a question for the producer: why is the feed incomplete, is it systematic, does it correlate with something? The second is a question for you: the expression needs a guard, or the denominator needs a rule, or the upstream step that produced the zero is wrong. Reporting them as one number sends the wrong team after the problem. ## What a single-marker design collapses The floating-point format reserves a not-a-number pattern for results it cannot define — that is what a zero divided by a zero produces. Some designs, having that pattern already sitting in the format, borrow it as their marker for "no value here". The consequence is not subtle: - both causes write **the same pattern** into the cell; - the absence test answers **true** for both, because it is asking about that pattern; - the column afterwards contains **no evidence at all** of which cause produced which hole. So a report reading "1,240 absent" over a computed column is answering a question nobody asked. It may be 1,240 missing source readings, or 1,240 division bugs, or any mixture, and the column itself will never say. ## What two markers keep apart A design that defines its own typed absence marker — a single absence value the library owns, usable in a column of any representation — can leave the format's undefined-arithmetic pattern alone to mean what the format says it means. The two then coexist as distinct objects: | property | one marker (borrowed pattern) | two markers (typed absence, plus the format's undefined value) | |---|---|---| | unrecorded observation | the reserved pattern | the tool's own absence marker | | undefined arithmetic result | the same reserved pattern | the format's not-a-number value | | predicates needed to separate them | none available | two, one per object | | what the column remembers | only that something is missing | which of the two causes applied | | what an audit can report | a single hole count | a gap count and a failed-computation count | This is the clearest case in the whole subject where the number of markers a design carries is not an implementation detail but a difference in what your data can tell you afterwards. ## Operating on a single-marker design If the tool you are on has one marker, the distinction has to be preserved **outside** the column, because nothing inside it survives. Three habits do it, in increasing order of effort: 1. **Keep the inputs.** As long as the operand columns are still there, you can reconstruct the cause after the fact: a row whose inputs were both present but whose result is absent had an undefined computation; a row with an absent input had a gap. 2. **Guard the expression rather than letting it fail.** Compute the ratio only where the denominator is present and non-zero, so the undefined case never reaches the output column and the holes there mean one thing. 3. **Count at the point of computation.** How many rows were excluded because an operand was missing, and how many because the denominator was zero, are two numbers you can record while you still know them — and cannot recover once the column is written. The general principle behind all three is that **the evidence lives in the computation's history, not in the column**. Once a single-marker column is handed to someone else, the causal information is gone, and no amount of inspection recovers it. ## What this looks like in an interview The question is usually put as a scenario — a dashboard shows a rising count of missing conversion rates, and someone wants to know whether the feed broke. The strong answer names the fork immediately: *it depends whether the tool has one marker or two*. Then it says what follows from each branch, and finishes with the operational point — on a single-marker design the answer cannot come from the finished column, so go back to the inputs or to the step that produced it. The weak answer is the confident one: "absent means the source had no data". It is true of some of the rows some of the time, it is exactly what a rising division bug looks like, and it is the reason this question gets asked.
- Under a design with one marker, what evidence is left in the data that separates the two causes?None inside the column. The only evidence is outside it: the operand columns, if they still exist, let you reconstruct which rows had a missing input and which had present inputs but an undefined result, and counts recorded at the moment of computation preserve the split directly. Once the inputs are discarded and only the computed column is handed on, the distinction is unrecoverable.
- Why does the distinction change what you do next, rather than being a curiosity?Because the owners differ. An unrecorded observation is a question for whoever produces the data — coverage, a broken feed, a population that genuinely has no value. An undefined arithmetic result is a defect in your own expression, usually a missing guard on a zero denominator. Merging them into one hole count sends the wrong team to investigate and hides a bug behind a data-quality story.
saying these in an interview costs you the question
- Assumes an absent cell always means the source had no reading
- Believes the absence test separates a data gap from a division bug
- Thinks every tool keeps the two causes as distinct objects
- Says the cause is recoverable by inspecting the finished column
- Reports one hole count without asking which kind they are