skip to content

Codd argued that a single NULL marker is not sufficient and proposed distinguishing two kinds of missing information. What was that distinction, why did he consider one marker inadequate, and how do practitioners deal with it today?

level: seniorimportance: nice to knowfreq 25%

answer

  1. A-mark = missing but applicable
  2. I-mark = missing and inapplicable
  3. Codd RM/V2, four-valued logic, never adopted
  4. Date/Darwen: remove nulls, don't add marks
  5. today: decompose, status column, or subtype

basics

~20 s

He separated 'missing but applicable' (a value exists, unknown to us) from 'missing and inapplicable' (no value can exist), proposing two markers and four-valued logic. One marker conflates them, so queries cannot tell them apart. Today the usual answer is decomposition into separate tables or an explicit status column.

solid answer

~50 s

In his later work (RM/V2) Codd argued that SQL's single NULL blurs two different facts: a value that exists but has not been recorded - an employee's phone number - and a value that cannot exist at all - a termination date for a current employee. He proposed distinct A-marks (applicable, missing) and I-marks (inapplicable), with a four-valued logic to reason over them. The proposal was never adopted, and it drew heavy criticism - most prominently from Date and Darwen, who argued the fix compounds the problem: if one marker already gives counter-intuitive results, two markers and a four-valued logic give more of them, and users can reason about neither. Their alternative is to keep missing values out of the logic entirely. In practice we handle it structurally rather than logically: move the sometimes-inapplicable attribute into its own table so that row presence encodes applicability, or carry an explicit status column saying why the value is absent. Both replace a logical distinction with data you can query.

go deeper

for a junior

Know that one NULL covers both 'unknown' and 'does not apply', and that this ambiguity is a recognised weakness.

for a middle

Name Codd's two-marker proposal and give a concrete example where conflating the two absences makes a query unanswerable.

for a senior

Explain why it was rejected, present the Date/Darwen counter-position, and choose between decomposition, a status column and subtyping for a given case.

for a principal

Set modelling policy: applicability determined by entity kind becomes subtyping, reason-for-absence that must be reported becomes an explicit column, and decomposition is reserved for sparse or separately-governed attributes.

## The two absences Codd's observation is simple once seen. Consider an employees table with a termination_date. For a departed employee whose paperwork was never filed, the date exists in the world and is missing from the database - **missing but applicable**. For a currently employed person there is no termination date at all, and never was - **missing and inapplicable**. Both are stored as NULL, so the two facts are indistinguishable to any query. A count of employees with no recorded termination date mixes 'still here' with 'left, unrecorded', and no query can separate them, because the information needed to do so was destroyed at write time. The same shape recurs constantly: a maiden name for someone never married, a shipping address for a digital product, a second-approver id on a workflow that requires only one approver for small amounts. ## Codd's proposal In *The Relational Model for Database Management: Version 2* (1990) Codd introduced two markers - commonly called the A-mark (missing but applicable) and the I-mark (missing and inapplicable) - and extended the logic accordingly. With two kinds of unknown the truth tables grow: instead of three truth values you reason over four, and every operator, constraint rule and aggregate needs defined behaviour for both marks. ## Why it was rejected No mainstream SQL engine implemented it, and the critique was fierce. Date and Darwen's argument is that the difficulty with NULL is not that there is one marker instead of two, but that putting 'no information' inside the logic at all produces results users cannot predict: a predicate and its negation that both fail to return a row, constraints that accept what they appear to forbid, aggregates whose behaviour differs from the arithmetic that produced them, and optimizer rewrites that are valid in two-valued logic but not in three. Doubling the marks doubles that surface. Their prescription runs the opposite way: forbid missing values in the database and represent absence with structure. There is also a pragmatic objection. Two markers push a modelling decision into every insert: the application must now know *why* a value is absent, at a point where it usually does not. If the application does know, it can record that knowledge as ordinary data - which is exactly what the practical solutions do. ## How practitioners handle it **Vertical decomposition.** Move the optional attribute into its own table keyed by the entity id, storing a row only when the attribute applies and is known. Presence of a row then means 'applicable and known', and absence means 'not applicable or not known' - which sounds like the same ambiguity but is not, because applicability is usually derivable from the entity's own state (a current employee has no termination row *because* they are current). Taken to its limit this is sixth-normal-form-style modelling, and the cost is real: many narrow tables and an outer join per attribute to reassemble a row, which keeps it a targeted tool rather than a default. **Explicit status column.** Keep the nullable value column and add a NOT NULL column recording the reason for absence: KNOWN, UNKNOWN, NOT_APPLICABLE, REFUSED. This is common in survey, medical and regulatory data, where 'patient declined to answer' is itself a finding that must be reported separately from 'not asked'. It keeps absence queryable with ordinary two-valued predicates and lets you constrain the pairing - value present exactly when the status is KNOWN. **Subtyping.** When applicability is determined by the entity's kind rather than by chance, model the kinds as separate tables. A physical product has a shipping weight; a digital product does not, and its table simply has no such column. This is the cleanest resolution where it fits, because inapplicability disappears at the schema level instead of being represented at all. ## What to say in an interview Name the distinction and the terminology, state that the proposal was not adopted and why (added complexity without removing the underlying unpredictability), then move quickly to practical handling: decomposition, an explicit reason-for-absence column, or subtyping, chosen by whether applicability is determined by the entity's kind or by the state of the data. Mentioning the Date/Darwen counter-position shows you know this is a live disagreement rather than settled folklore.

  • What is the main objection to adding a second marker and moving to four-valued logic?
    That it multiplies the very unpredictability it aims to fix. Three-valued logic already produces results users misread - constraints that accept what they look like they forbid, complementary filters that do not cover every row - and four-valued logic doubles the rules every operator, constraint and aggregate must define. Critics argue absence should be modelled as data or structure, not encoded in the logic.
  • When is vertical decomposition the wrong answer to inapplicable values?
    When the attribute is merely optional rather than inapplicable and is read on the common path. Splitting it out means an outer join on every read, more objects to maintain, and transactional writes across tables, all for a distinction nobody queries. Reserve it for sparse attributes, for attributes with genuinely different lifecycles, or where 'does this apply' must itself be queryable.

saying these in an interview costs you the question

  • Claiming SQL implements two distinct NULL markers
  • Presenting four-valued logic as the accepted modern solution
  • Saying unknown and inapplicable are the same thing in practice
  • Proposing sentinel values in the same domain as real data to encode the difference

context