Edgar Codd published a numbered set of rules a database system must satisfy to be legitimately called relational. Why was that list written, and what do Rule 0 (foundation), Rule 1 (information) and Rule 2 (guaranteed access) require?
answer
- 1985 checklist vs 1970 model
- Rule 0 = entirely relational, no side door
- Rule 1 = values in tables, one way only
- Rule 2 = table + key + column
- no meaning in row order or pointers
basics
~20 sCodd published them in 1985 because vendors marketed non-relational products as relational. Rule 0: manage all data entirely through relational capabilities. Rule 1: all information appears as values in tables. Rule 2: every value is reachable by table name, key value and column name.
solid answer
~50 sCodd defined the relational model in 1970; by 1985 vendors were calling hierarchical, network and file-based products relational, so he published Rule 0 plus twelve rules as an objective test. **Rule 0 (foundation):** a relational DBMS must manage all its stored data entirely through its relational capabilities — no non-relational side door. **Rule 1 (information):** every piece of information, including metadata, is represented in exactly one way — as values in columns of rows of tables. No hidden pointers, no row ordering carrying meaning, no significance in physical placement. **Rule 2 (guaranteed access):** every single value is addressable by the combination of table name + primary key value + column name. Nothing requires positional access like 'the fourth row'. Together they force a single, uniform, value-based way to store and reach data. They are a yardstick, not a conformance standard: no commercial engine passes all thirteen.
go deeper
Know the origin story and be able to state Rules 0, 1 and 2 in one sentence each, with 'no meaning in row order' as the concrete takeaway.
Connect Rule 2 to why every table gets a primary key, and give a real example of information hidden outside table values.
Discuss where SQL engines depart — bags versus sets, keyless tables, external/LOB storage — and why those departures are usually pragmatic rather than sloppy.
Use the rules as review vocabulary: name what a proposed design gives up (uniform representation, logical addressability) rather than citing rule numbers for their own sake.
## Why a rule list exists Edgar F. Codd, an IBM mathematician, published the relational model in 1970. Its idea: present data purely as **relations** (tables of rows and columns) and let users say *what* they want rather than *how* to navigate to it. Through the 1970s and early 1980s the word relational became a marketing label, attached to products that were really hierarchical (tree-shaped, parent-to-child pointers), network/CODASYL (record sets linked by pointers the program follows), or just indexed files with a query veneer. In 1985 Codd published a checklist — Rule 0 plus twelve numbered rules — so buyers could test the claim. ## Rule 0 — the foundation rule A system claiming to be relational **must manage its stored data entirely through its relational capabilities**. The load-bearing word is *entirely*. A product can expose perfectly good SQL and still fail: if some data can only be created, read, or maintained by a record-at-a-time navigational API, or lives in a structure the query language cannot address, then relational access is a facade over something else. Everything else on the list elaborates this one. ## Rule 1 — the information rule **All information in the database is represented explicitly and in exactly one way: as values in columns within rows of tables.** Three consequences: - *Metadata is data.* Table, column and constraint definitions are themselves held in tables, not in a separate format only the engine understands. - *No hidden state.* Physical pointers linking records, file placement, or the order rows happen to sit in must not carry meaning a user depends on. A relation is a set; if 'the row after this one' matters, that fact must be an explicit column value. - *One representation.* Two different mechanisms for holding the same kind of fact (say, some customers in a table and some in an attached document store the query language cannot see) violate it. ## Rule 2 — guaranteed access **Every atomic value is guaranteed to be logically reachable by naming three things: the table, the primary key value of the row, and the column.** No value should require positional or navigational access — no 'fifth row', no 'follow this pointer'. This is why the model insists every relation has a key: without one, some rows are indistinguishable and therefore not individually addressable. It also implies the access path is a *logical* address, not a physical one; how the engine actually finds the row is its own business. ## Where real engines land SQL engines respect these rules approximately, not exactly: - SQL tables are **bags**, not sets — duplicate rows are allowed, and a table with no key breaks guaranteed access outright. This is why practitioners insist every table gets a primary key: it is Rule 2 restated as a coding standard. - Large objects, file-system-backed columns, and external table wrappers stretch Rule 1's 'one way'. - Row order is not guaranteed, and any application that relies on insertion order without an ordering column is quietly depending on physical placement — exactly what Rule 1 forbids. ## How to use them today Nobody ships a Codd-conformance report. What the rules give you is a vocabulary. When someone proposes storing meaning in row order, or hiding a business fact in a filename, or reaching data through an interface the query language cannot see, these three rules name precisely what is being given up: uniform representation and guaranteed logical addressability, which are what make declarative querying and independent schema evolution possible in the first place.
- Why does the guaranteed-access rule effectively require every table to have a primary key?The rule addresses a value by table name, primary key value and column name. Without a key, two rows can be identical or indistinguishable, so there is no value that names one of them uniquely and those rows cannot be individually addressed or updated. SQL permits keyless tables, which is one of the places practice departs from the model.
- A team stores an ordering by relying on physical insert order rather than a sequence column. Which rule does that break, and why is it a practical problem?It breaks the information rule: meaning is carried by physical placement rather than by a value in a column. Practically, the engine gives no ordering guarantee without an explicit sort, so a plan change, a parallel scan, or a table rewrite can reorder rows and silently change application behaviour.
Rule 2 is a postal address system: every house is reachable by street + house number + apartment, never by 'the fourth door on the left after the oak tree'.
saying these in an interview costs you the question
- Saying the twelve rules are a standard databases are certified against — they are Codd's own yardstick, and no engine passes all of them.
- Claiming SQL tables are sets in the strict sense; they permit duplicate rows unless a key or DISTINCT forbids them.
- Treating 'the order rows come back in' as stable without an explicit ordering.
- Thinking guaranteed access is about performance or indexes — it is about logical addressability, not access speed.