skip to content

A boolean column, a text column and a time column each need a way to record nothing, so what can serve each?

level: middleimportance: should knowfreq 44%

answer

  1. inside the value needs a spare pattern
  2. two boolean patterns, both taken
  3. a time column needs its own absence object
  4. references accept every kind of absence
  5. outside the value serves every column kind

basics

~20 s

Only encodings that live outside the value serve every column kind. A packed boolean has no spare pattern, a time column needs its own absence object, and text holds the language's empty reference only where cells are references.

solid answer

~50 s

Absence encoded *inside* the value needs the value format to have a pattern spare, and most formats do not. A boolean packed one bit per row has two patterns and both are taken, so absence has to come from a validity track beside it or a marker the tool owns. A time column is usually a count of ticks from an origin, so a design that reserves patterns inside values needs a separate absence object for time — a different object from the numeric one, which is why a check written for one can walk straight past the other. Text is the odd case: where cells are references to arbitrary objects, the host language's empty reference fits directly, but the column then holds references to anything and can accumulate several different absence-ish objects at once. A validity track or a typed marker is the only answer that serves all three unchanged.

go deeper

for a junior

Recall that some column kinds cannot hold a hole in every tool, and that the reason is storage rather than policy. Knowing that a boolean or a time column is a harder case than a decimal one is the takeaway.

for a middle

Explain the rule that decides it: a marker inside the value needs the format to have a pattern spare, while a validity track or a marker the tool owns works for any column at its own width.

for a senior

Show the audit consequence — an absence count is only as good as the check that produced it, and a column of arbitrary references can hold several different absence-ish objects that no single equality test will find.

for a principal

The tradeoff to own is uniformity: insisting that every column keep a typed representation with absence tracked beside it costs conversion work at every boundary, and buys hole counts that mean the same thing everywhere.

## The question is per column kind, not per table "Can this column hold a hole?" has no single answer for a table, because the answer depends on two things at once: **what the column's values are stored as**, and **which absence encoding the tool uses**. An encoding that lives inside the value can only work where the value format has a pattern to lend. An encoding that lives outside the value works everywhere. Walking three column kinds makes the rule concrete. ## A boolean column A boolean has exactly two states. Two storage choices are common: - **Packed, one bit per row.** Both patterns are taken. There is nothing spare, so an absence marker cannot live inside the value at all. - **One byte per row.** Most patterns are unused, so a tool *could* reserve one — but doing so means the column's values are no longer what a plain reader of a byte would assume, and the convention has to travel with the data. In practice, a boolean column that admits absence is a column with a validity track beside it or a typed absence marker the library defines. This is why some tools simply cannot store a hole in a boolean column: asked to, they turn it into a column of arbitrary references, and you have silently left the world of uniform typed storage. ## A text column Text is variable-length, so it is stored either as references to string objects, or as a values buffer with offsets into it. - Where the cells are **references to arbitrary objects**, the host language's own "nothing here" object — its empty reference — drops straight into the cell and the tool needs nothing else. The cost is that the column will accept anything: one loader leaves the language's empty reference, a numeric step leaves a reserved floating-point pattern, a third path leaves the tool's own marker, and the column now holds several different absence-ish objects. A check written for one of them under-reports the holes. - Where text is stored in its **own buffer with offsets**, absence is recorded beside the values in a validity track, exactly as for numbers, and the column stays uniform. ## A time column A time-typed column is almost always a whole number of ticks from an origin, dressed up with an interpretation. Under a design that reserves patterns inside values, that column needs an absence object of its own — **the absence marker a time-typed column uses, which is a different object from the numeric one** — because it must decode as a time, print as a time and order as a time. The practical consequence is a trap: - the general absence test recognises both, because it was written to; - code that tests for the *numeric* absence object by identity does not recognise the time one, and vice versa; - so a hand-rolled audit that counts holes with an equality against one marker reports a time column as complete when it is not. ## Summary of what serves what | column kind | pattern inside the value | validity bit beside it | the language's empty reference | the tool's typed marker | |---|---|---|---|---| | whole-number | no spare pattern | yes, width unchanged | only as a column of references | yes | | floating-point | yes, the format reserves one | yes | only as a column of references | yes | | boolean | not when packed | yes | only as a column of references | yes | | text | depends on the storage form | yes | yes, where cells are references | yes | | time | needs its own distinct object | yes | only as a column of references | yes | Read the table column-wise rather than row-wise and the rule falls out: **the two encodings that live outside the value serve every column kind, and the two that live inside the value are conditional on the format.** ## Why a column of arbitrary references is not the general answer It is tempting to treat "just make it a column of references" as the universal escape hatch, because it accepts every absence-ish object there is. Three things argue against it: 1. **The single-form guarantee is gone.** Every cell is a reference to anything, so nothing enforces that the column still contains one kind of thing, and nothing raises when it stops. 2. **Absence stops being one thing.** Several different absence-ish objects can coexist in the one column, so the number of holes depends on which check you ran. 3. **The declared representation stops carrying information.** Downstream code that reads the representation to decide how to handle the column learns nothing from it. The interview-grade summary is short: a tool that offers a marker outside the value — a validity track, or one typed absence marker — can admit absence in any column at that column's own width, while a tool that encodes absence inside the value must find a spare pattern per format and does not always have one.

  • Why does a design with one typed absence marker have an easier time across column kinds?
    Because the marker is the library's own value rather than a pattern borrowed from a particular format, it does not have to be negotiated per representation. One marker is admissible in a whole-number, boolean, text or time column, each of which keeps its own width and its own form, and one predicate finds every hole in the table regardless of the column it is looking at.
  • What actually goes wrong when a column of arbitrary references is used as the general answer?
    The column loses the guarantee that every cell holds one kind of thing, which is the guarantee whole-column work is built on. It will also accept more than one absence-ish object at the same time, so the hole count depends on which check you ran, and its declared representation no longer tells downstream code anything useful about the contents.

saying these in an interview costs you the question

  • Assumes the numeric absence object also works in a time column
  • Says a packed boolean column has a spare pattern for absence
  • Believes any text column can hold the language's empty reference
  • Treats a column of references as a free general solution
  • Thinks the encoding choice is invisible above the storage layer