skip to content

What is a composite key, and what makes an attribute 'prime' rather than 'non-prime' in a relation?

level: middleimportance: should knowfreq 42%

answer

  1. composite = unique only in combination, still minimal
  2. prime = in SOME candidate key, not just the primary one
  3. non-prime = descriptive payload
  4. extra attribute on a key = superkey, weaker constraint
  5. primeness is per relation, not per column name

basics

~20 s

A composite key is a candidate key made of two or more attributes that are only unique together. A prime attribute is one that belongs to at least one candidate key of the relation; every other attribute is non-prime. Prime status is per relation, not per column type.

solid answer

~60 s

A **composite (compound) key** is a candidate key with more than one attribute: no single member is unique, but the combination is, and no member can be dropped without losing uniqueness. A classic case is an enrolment relation keyed by student, course and term, or any relation resolving a many-to-many relationship where the key is the pair of participant identifiers. An attribute is **prime** if it appears in *at least one* candidate key of the relation, and **non-prime** otherwise. The test spans all candidate keys, not only the primary key: if email is an alternate key, email is a prime attribute even though it is not part of the primary key. The distinction matters because the classical normal forms are stated in terms of dependencies of non-prime attributes on keys, so you cannot even evaluate them until you know which attributes are prime. Practically, composite keys also propagate: a table referencing a three-attribute key must carry all three attributes, which is one of the usual arguments for introducing a single-attribute identifier instead.

code

text · 7 lines
text
Enrolment(student_id, course_id, term, grade, registered_on)
rule: a student takes a given course at most once per term

candidate key : {student_id, course_id, term}      (minimal, composite)
superkey only : {student_id, course_id, term, grade}
prime         : student_id, course_id, term
non-prime     : grade, registered_on

go deeper

for a junior

Give a concrete composite example such as enrolment or order line, and state the one-line prime versus non-prime test.

for a middle

Show minimality explicitly by removing each attribute in turn, and state that primeness ranges over all candidate keys.

for a senior

Discuss propagation cost of wide keys into referencing relations and interfaces, and insist the composite rule stays enforced even when a generated identifier is added.

for a principal

Position identity width as a system-wide interface decision affecting messages, caches and external contracts, not just table design.

## Composite keys Most discussions of keys implicitly assume a single attribute, but the definition of a candidate key never mentioned attribute count. A candidate key is a minimal set of attributes that is unique per tuple, and that set may hold two, three or more attributes. When it does, it is called a composite or compound key. A typical example is a relation recording that a student took a course in a given term: `Enrolment(student_id, course_id, term, grade, registered_on)` No single attribute is unique. A student appears many times, a course appears many times, a term appears many times. The triple `{student_id, course_id, term}` is unique if the rule is that a student may take a given course at most once per term. It is also minimal: drop `term` and repeat takers collide, drop `course_id` and a student's other courses collide, drop `student_id` and classmates collide. So it is a genuine composite candidate key. Composite keys arise naturally in three places: link relations resolving many-to-many relationships, weak entities identified only in the context of a parent (order line identified by order plus line number), and time-variant or versioned data where the key includes a period or version attribute. ## Minimality is what separates a composite key from a pile of columns A frequent error is to bolt extra attributes onto a key "to be safe". If `{student_id, course_id, term}` is already unique, then `{student_id, course_id, term, registered_on}` is still unique, but it is not a candidate key; it is a non-minimal superkey. Declaring it as the identifier is harmful: it weakens the constraint the database enforces, since two rows differing only in registration timestamp would now be allowed, and it forces every referencing relation to carry a fourth attribute. Order of attributes inside the key is irrelevant to the relational model. It matters only to physical access structures, which is a separate concern from key taxonomy. ## Prime and non-prime attributes Once candidate keys are known, every attribute of the relation falls into exactly one of two classes: - **Prime attribute**: it is a member of at least one candidate key. - **Non-prime attribute**: it is a member of no candidate key. In the enrolment relation, `student_id`, `course_id` and `term` are prime, while `grade` and `registered_on` are non-prime. Notice that the test quantifies over *all* candidate keys. Suppose a relation has candidate keys `{emp_id}` and `{email}`; then both `emp_id` and `email` are prime, even though only one of them is the primary key. Candidates who think only the primary key's attributes count get this wrong routinely. Primeness is a property of an attribute *within a relation*, not an intrinsic property of a column name. The same attribute may be prime in one relation and non-prime in another: `course_id` is prime in `Enrolment` but non-prime in a relation `CourseOffering(offering_id, course_id, room)` where the offering identifier is the key. ## Why the terminology exists The prime/non-prime split exists because dependency theory needs it. The classical normal forms are phrased as conditions on how non-prime attributes may depend on candidate keys, so the vocabulary is a prerequisite for talking about them at all. Even outside formal normalization, the distinction is a useful design lens: prime attributes carry identity and therefore propagate into other relations and into application URLs and messages, while non-prime attributes are descriptive payload that can change without changing what the row *is*. ## Design consequences Composite keys are correct but heavy at the edges. Every relation that references a three-attribute key must repeat all three attributes in its own reference, and a further reference to that relation may repeat four. Wide identity also travels into caches, message payloads and external interfaces. This propagation cost, not any theoretical defect, is the practical argument that leads teams to add a single generated identifier while still enforcing the composite natural key as an alternate key. Dropping the composite uniqueness rule when you add the generated identifier is the mistake that turns a clean model into one with duplicate business rows.

  • If a relation has candidate keys {emp_id} and {email}, is email a prime attribute?
    Yes. An attribute is prime when it belongs to at least one candidate key, and the test ranges over every candidate key rather than only the primary key. So both emp_id and email are prime here, and every remaining attribute is non-prime. The choice of which key is primary changes nothing about primeness.
  • Someone adds a timestamp column to an already-unique composite key. What is the effect?
    The result is still unique, so it remains a superkey, but it is no longer minimal and therefore no longer a candidate key. The practical damage is that the enforced rule becomes weaker: two rows that differ only by timestamp are now permitted even though the business says they are the same fact. It also forces every referencing relation to carry the extra attribute.

saying these in an interview costs you the question

  • Believing a key must be a single attribute, so composite keys are a workaround rather than a key
  • Calling any unique multi-attribute set a composite key without checking minimality
  • Testing primeness only against the primary key and ignoring alternate keys
  • Treating attribute order within the key as part of the logical definition
  • Dropping the composite uniqueness rule after adding a single-attribute identifier

context