skip to content

How do a wide-column store's two primary-key shapes, a hashed partition key with clustering columns versus one sorted row key, place and order rows?

level: middleimportance: must knowfreq 70%

answer

  1. one key, two jobs
  2. hash decides the node
  3. clustering orders inside it
  4. byte order across ranges
  5. prefix scans vs point partitions

basics

~20 s

With a hashed partition key, the hash picks the servers and clustering columns sort rows inside that partition. With one sorted row key, the whole keyspace is kept in byte order and split into contiguous ranges, each served by one server.

solid answer

~50 s

The family uses two key shapes, and in both the key does two jobs: **placement** and **on-disk order**. - **Hashed partition key + clustering columns.** The partition key is hashed to a position on a ring, which picks the replicas. Inside the partition, rows are kept sorted by the clustering columns. So a query must name the full partition key, and can then read a slice of rows in clustering order. Partitions themselves are not in any useful order. - **One sorted row key.** Rows are kept in lexicographic byte order across the whole table, and the keyspace is split into contiguous ranges (tablets or regions) that split as they grow, each served by one server. Any key prefix is a contiguous range, so scans across many rows are cheap. The trade: hashing spreads load but loses global order; a sorted keyspace keeps order but makes sequential keys land on one range.

go deeper

for a junior

Know that the key decides both where a row is stored and its order, and name the two shapes, partition key with clustering columns and a single sorted row key.

for a middle

Explain what a read must supply in each shape, which range reads are cheap, and why the hashed shape loses global order while the sorted shape keeps it.

for a senior

Predict each shape's failure modes, an unbounded or hot partition versus a sequential key hammering one range, and translate a design from one shape to the other.

for a principal

Be ready to argue which shape a workload needs, global ordered scans against default load spreading, and what the choice commits the team to for years.

## The key does two jobs In every wide-column store the primary key decides **where** a row is stored and **in what order** it sits relative to its neighbours. Getting it right is most of the design, because both jobs are fixed at write time. The family implements this in two recognisable shapes. ## Shape 1: hashed partition key plus clustering columns The primary key is split in two: - the **partition key** — one or more columns whose values are hashed to a token; the token decides which nodes hold the **partition** (and its replicas); - the **clustering columns** — the rest of the key, which order the rows **inside** the partition. A partition is therefore a sorted run of rows that all live together. Consequences: - A read must supply the **whole partition key**; without it the store does not know where to look. - Inside a partition you can read one row, or a **slice** of rows by a range on the clustering columns, in stored order (or reversed). - Across partitions there is no useful order: two partition keys that are alphabetically adjacent land wherever their hashes send them. - Hashing spreads keys evenly by default, so sequential partition keys do not pile onto one node. Some stores in this shape also offer **partition-level values** that are stored once per partition and shared by every row in it — useful for a header attribute of the entity the partition represents. ## Shape 2: one sorted row key over range-split tablets There is a single **row key**, compared as bytes. The whole table is kept in **lexicographic order**, and the keyspace is cut into **contiguous ranges** (called tablets or regions) that split automatically as they grow and are spread over servers, with each range served by one server at a time. - A **point read** needs the full row key. - A **prefix** or **start/end range** is always a contiguous run of rows, so scanning "everything for customer 42" is one sequential read if the key starts with the customer id. - Because order is global, **sequential keys** (timestamps, counters) all land at the end of the keyspace, on the one range currently holding it — the classic hotspot this shape has to design around. - Composite keys are built by concatenating fields with delimiters or fixed widths, and numbers must be encoded so that byte order matches numeric order. ## Side by side | | hashed partition key + clustering | one sorted row key | |---|---|---| | placement | hash of the partition key picks the nodes | the key's position in a sorted keyspace picks the range | | order across keys | none useful (hash order) | global byte order | | order within a key | clustering columns | columns within the row; related rows are adjacent | | cheap range read | a slice inside one partition | any key prefix or start/end range | | default load spread | even, from the hash | follows the key distribution; sequential keys concentrate | | typical failure | one partition growing without bound or taking all traffic | a monotonically increasing key hammering the last range | ## Mapping one design onto both A "readings per device" table illustrates the equivalence: 1. Hashed shape: partition key = device id, clustering column = reading time. One partition per device, readings sorted by time. 2. Sorted shape: row key = `deviceId#time`. One row per reading, all of a device's rows adjacent and in time order. Both answer "readings for device X between two times" with one sequential read. They differ in what happens across devices: the hashed shape scatters devices; the sorted shape keeps them in device-id order, which enables a scan over a range of devices but makes key distribution your problem. ## Why interviewers ask A candidate who can say which shape a store uses, and what that implies for range reads and hotspots, can reason about any product in the family without memorising its syntax.

  • In the hashed shape, can you read rows from many partitions in key order?
    Not efficiently. Partitions are placed by hash, so the store has no global key order to walk; a scan across partitions returns them in token order. If a read needs order across entities, that order must live inside one partition's clustering columns or come from a different table keyed for it.
  • Why must numbers in a sorted row key be encoded carefully?
    The key is compared as bytes, so the string "10" sorts before "9". Fixed-width, zero-padded or big-endian binary encodings make byte order match numeric order, which range reads over ids or times depend on.
  • What is the unit that moves between servers in each shape?
    In the hashed shape it is a token range, a slice of the hash ring holding many partitions. In the sorted shape it is a contiguous key range that splits when it grows. In neither shape is a single partition or row split across servers.

saying these in an interview costs you the question

  • Saying the partition key only identifies a row and has nothing to do with placement
  • Expecting a range query across partition keys to return rows in key order in the hashed shape
  • Believing sequential row keys spread evenly in a range-split keyspace
  • Treating clustering columns as secondary indexes rather than the order rows are stored in
  • Thinking one very large partition or row is split across several servers automatically