skip to content

questions

5

Edgar Codd published a numbered set of rules a database system must satisfy to be legitimately called relational. Why was that list written, and what do Rule 0 (foundation), Rule 1 (information) and Rule 2 (guaranteed access) require?

level: juniorimportance: should knowfreq 38%

answer

  1. 1985 checklist vs 1970 model
  2. Rule 0 = entirely relational, no side door
  3. Rule 1 = values in tables, one way only
  4. Rule 2 = table + key + column
  5. no meaning in row order or pointers

basics

~20 s

Codd published them in 1985 because vendors marketed non-relational products as relational. Rule 0: manage all data entirely through relational capabilities. Rule 1: all information appears as values in tables. Rule 2: every value is reachable by table name, key value and column name.

solid answer

~50 s

Codd defined the relational model in 1970; by 1985 vendors were calling hierarchical, network and file-based products relational, so he published Rule 0 plus twelve rules as an objective test. **Rule 0 (foundation):** a relational DBMS must manage all its stored data entirely through its relational capabilities — no non-relational side door. **Rule 1 (information):** every piece of information, including metadata, is represented in exactly one way — as values in columns of rows of tables. No hidden pointers, no row ordering carrying meaning, no significance in physical placement. **Rule 2 (guaranteed access):** every single value is addressable by the combination of table name + primary key value + column name. Nothing requires positional access like 'the fourth row'. Together they force a single, uniform, value-based way to store and reach data. They are a yardstick, not a conformance standard: no commercial engine passes all thirteen.

go deeper

for a junior

Know the origin story and be able to state Rules 0, 1 and 2 in one sentence each, with 'no meaning in row order' as the concrete takeaway.

for a middle

Connect Rule 2 to why every table gets a primary key, and give a real example of information hidden outside table values.

for a senior

Discuss where SQL engines depart — bags versus sets, keyless tables, external/LOB storage — and why those departures are usually pragmatic rather than sloppy.

for a principal

Use the rules as review vocabulary: name what a proposed design gives up (uniform representation, logical addressability) rather than citing rule numbers for their own sake.

## Why a rule list exists Edgar F. Codd, an IBM mathematician, published the relational model in 1970. Its idea: present data purely as **relations** (tables of rows and columns) and let users say *what* they want rather than *how* to navigate to it. Through the 1970s and early 1980s the word relational became a marketing label, attached to products that were really hierarchical (tree-shaped, parent-to-child pointers), network/CODASYL (record sets linked by pointers the program follows), or just indexed files with a query veneer. In 1985 Codd published a checklist — Rule 0 plus twelve numbered rules — so buyers could test the claim. ## Rule 0 — the foundation rule A system claiming to be relational **must manage its stored data entirely through its relational capabilities**. The load-bearing word is *entirely*. A product can expose perfectly good SQL and still fail: if some data can only be created, read, or maintained by a record-at-a-time navigational API, or lives in a structure the query language cannot address, then relational access is a facade over something else. Everything else on the list elaborates this one. ## Rule 1 — the information rule **All information in the database is represented explicitly and in exactly one way: as values in columns within rows of tables.** Three consequences: - *Metadata is data.* Table, column and constraint definitions are themselves held in tables, not in a separate format only the engine understands. - *No hidden state.* Physical pointers linking records, file placement, or the order rows happen to sit in must not carry meaning a user depends on. A relation is a set; if 'the row after this one' matters, that fact must be an explicit column value. - *One representation.* Two different mechanisms for holding the same kind of fact (say, some customers in a table and some in an attached document store the query language cannot see) violate it. ## Rule 2 — guaranteed access **Every atomic value is guaranteed to be logically reachable by naming three things: the table, the primary key value of the row, and the column.** No value should require positional or navigational access — no 'fifth row', no 'follow this pointer'. This is why the model insists every relation has a key: without one, some rows are indistinguishable and therefore not individually addressable. It also implies the access path is a *logical* address, not a physical one; how the engine actually finds the row is its own business. ## Where real engines land SQL engines respect these rules approximately, not exactly: - SQL tables are **bags**, not sets — duplicate rows are allowed, and a table with no key breaks guaranteed access outright. This is why practitioners insist every table gets a primary key: it is Rule 2 restated as a coding standard. - Large objects, file-system-backed columns, and external table wrappers stretch Rule 1's 'one way'. - Row order is not guaranteed, and any application that relies on insertion order without an ordering column is quietly depending on physical placement — exactly what Rule 1 forbids. ## How to use them today Nobody ships a Codd-conformance report. What the rules give you is a vocabulary. When someone proposes storing meaning in row order, or hiding a business fact in a filename, or reaching data through an interface the query language cannot see, these three rules name precisely what is being given up: uniform representation and guaranteed logical addressability, which are what make declarative querying and independent schema evolution possible in the first place.

  • Why does the guaranteed-access rule effectively require every table to have a primary key?
    The rule addresses a value by table name, primary key value and column name. Without a key, two rows can be identical or indistinguishable, so there is no value that names one of them uniquely and those rows cannot be individually addressed or updated. SQL permits keyless tables, which is one of the places practice departs from the model.
  • A team stores an ordering by relying on physical insert order rather than a sequence column. Which rule does that break, and why is it a practical problem?
    It breaks the information rule: meaning is carried by physical placement rather than by a value in a column. Practically, the engine gives no ordering guarantee without an explicit sort, so a plan change, a parallel scan, or a table rewrite can reorder rows and silently change application behaviour.

Rule 2 is a postal address system: every house is reachable by street + house number + apartment, never by 'the fourth door on the left after the oak tree'.

saying these in an interview costs you the question

  • Saying the twelve rules are a standard databases are certified against — they are Codd's own yardstick, and no engine passes all of them.
  • Claiming SQL tables are sets in the strict sense; they permit duplicate rows unless a key or DISTINCT forbids them.
  • Treating 'the order rows come back in' as stable without an explicit ordering.
  • Thinking guaranteed access is about performance or indexes — it is about logical addressability, not access speed.

context

open as a page

Codd's fourth rule requires a dynamic online catalog based on the relational model. What does that require, and how do relational engines satisfy it today?

level: middleimportance: should knowfreq 28%

basics

~20 s

The database's own description — tables, columns, constraints, privileges — must be stored as ordinary relational data and queried with the same language as user data, kept live and authoritative. Engines satisfy it with system catalogs and the standard INFORMATION_SCHEMA views.

open as a page

Codd's third rule requires systematic treatment of null values. What exactly does that rule require of an engine, and name a mainstream database behaviour that violates it.

level: middleimportance: should knowfreq 30%

basics

~20 s

It requires one single marker for missing or inapplicable information, supported uniformly for every data type and independent of any real value such as zero or empty string. Oracle treating the empty string as NULL for character columns violates it.

open as a page

Codd's view-updating rule says every view that is theoretically updatable must also be updatable by the system. Why do real engines fall short of that, and what do they offer instead?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Because for many views a change to the view has no unique translation back to base rows — joins, aggregates and set operations are ambiguous or lossy. Engines auto-update only simple single-table views and require an explicit handler, such as an INSTEAD OF trigger, for the rest.

open as a page

Codd's non-subversion rule forbids any low-level interface that bypasses the integrity rules enforced by the higher-level relational language. Which of Codd's rules do mainstream SQL engines actually fail, and how much should that influence a database choice?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Non-subversion means bulk loaders and record-level APIs must not skip constraints or authorisation. Mainstream engines commonly fall short on view updating, integrity independence, distribution independence and full non-subversion. It matters as a checklist of where your team must compensate, not as a product-selection score.

open as a page