What is a column family in a wide-column store, and why does grouping columns into families affect storage, retention and read cost?
answer
- declared up front, few of them
- qualifiers are open-ended
- stored and read together
- retention set per family
basics
~20 sA column family is a named group of columns declared in the schema; the columns inside it are created freely by writes. Stores keep a family's data together and apply settings such as version retention per family, so grouping decides what a read touches.
solid answer
~50 sIn stores that use them, a **column family** is a small, schema-declared group, and a column's full name is `family:qualifier`. The families are fixed and few; the **qualifiers** inside them are open-ended and created by writing. Families matter physically: a store typically keeps each family's cells **together on disk**, so a read of one family does not pay for another, and settings such as **how many versions to keep or how long to keep them** are set **per family**. That gives the design rule: put columns that are read together, and share a retention need, in the same family, and keep the number of families small, because each one adds its own files and flush and compaction work. In stores keyed by a partition key plus clustering columns, the older term "column family" simply names a table.
go deeper
Know that a column name has a family part declared in the schema and a qualifier part created by writes, and that families are few.
Explain that a family's cells are stored together and that retention and tuning are set per family, and what that means for which bytes a read touches.
Group real columns into families by access pattern and retention, and explain the flush and compaction cost of adding families.
Be ready to judge schema-level trade-offs such as one family versus several for a multi-tenant table, weighing read cost against operational overhead.
## Two levels of column naming Many wide-column stores name a column in two parts: - a **column family** — declared when the table is created or altered, and deliberately **few** in number; - a **column qualifier** — any name at all, created simply by writing a cell with it, so a family can hold millions of distinct qualifiers across the table. The full column name is written `family:qualifier`. The split lets the schema stay small and stable while each row is free to carry whatever qualifiers it needs. ## Why the family is a physical unit The family is not just a namespace. Stores in this lineage typically treat it as a unit of **storage and configuration**: - **Storage locality.** A family's cells are kept together in their own files, so a read that asks for one family does not read the bytes of another. Separating a large, rarely read blob family from small, hot metadata columns cuts the cost of the common read. - **Retention.** Garbage-collection rules — keep the newest *n* versions, or keep versions newer than an age — are set **per family**, not per column. Putting a column that needs one version in a family configured to keep a thousand wastes storage on 999 cells nobody reads. - **Tuning and access control.** Compression, caching hints and, in some stores, permissions are also set at the family level. ## The costs of many families Each family adds its own set of files for every key range. More families means: 1. more in-memory buffers and more files to flush, with families of one range competing for the same memory budget, so a busy family can trigger flushes that also write small files for quiet ones; 2. more compaction work, per family per range; 3. reads that need several families touching several sets of files. Guidance across the lineage is therefore to keep the count small: one family is common, a few is normal, and each extra one should be justified by a different access pattern or a different retention rule. ## A worked grouping | columns | read together? | retention | family | |---|---|---|---| | name, email, status | yes, on every profile read | latest only | `profile` | | login events | only in an audit screen | 90 days | `audit` | | avatar image bytes | rarely, and large | latest only | `media` | A profile read touches only `profile`; the audit screen scans only `audit`; the large images never inflate either. Retention is set once per family. ## The term in the other key shape In stores keyed by a hashed partition key plus clustering columns, "column family" is a **historical name for a table**, and there is no second level of grouping inside it. The storage lessons carry over at the table level: data read together belongs in the same table and partition, and retention (for example a default time-to-live) is set per table. When reading older material, check which meaning is in use. ## Interview angle A good answer names the two-level naming, says the family is the unit of storage and retention, gives the "few families, group by access pattern and retention" rule, and notes the terminology clash between the two shapes.
- Can two families in the same table hold the same row key?Yes. The row key is shared; each family stores its own cells for that key. A row can have cells in one family and none in another, and a read can ask for just the families it needs.
- When would you move a column into its own family?When it is read on a different path from the rest, is much larger, or needs a different retention rule. If none of those holds, an extra family only adds files, flushes and compactions.
saying these in an interview costs you the question
- Treating a column family as just a naming prefix with no physical effect
- Creating a new column family for every attribute of an entity
- Setting version retention per column rather than per family
- Assuming column family means the same thing in every wide-column store