skip to content

Why must a wide-column design bound how large one partition or row can grow, and how does adding a time bucket to the key bound it?

level: middleimportance: must knowfreq 54%

answer

  1. a unit never split
  2. one server holds it all
  3. growth without end
  4. time in the partition part
  5. reads now span buckets

basics

~20 s

A partition or row is never split across servers, so one that grows forever becomes slow to read, compact and repair and can hit size limits. Adding a time bucket to the grouping part of the key starts a fresh partition each period.

solid answer

~50 s

The unit a wide-column store places on a server, a **partition** in hash-partitioned stores or a **row** in range-split ones, is never split across servers. If the key groups by something that only grows (all readings of a device, all messages of a channel), that unit grows forever: reads of it slow down, background merging and repair must handle it whole, memory pressure rises, and stores enforce or recommend size ceilings. The fix is to add a **time bucket** (day, hour, month) to the **grouping** part of the key: `device + day` instead of `device`. Each period starts a fresh partition, so size is bounded by data per period. Choose the bucket from the write rate so a bucket stays well under the limits, and remember that reads spanning several periods must now query several buckets.

code

pseudocode · 6 lines
pseudocode
function readRange(deviceId, fromTime, toTime):
    results = []
    for day in daysBetween(fromTime, toTime):          # one partition per device-day
        partitionKey = (deviceId, day)
        results += slice(partitionKey, max(fromTime, startOf(day)), min(toTime, endOf(day)))
    return results                                     # already in time order, day by day

go deeper

for a junior

Know that keying time-series data only by source id makes one partition or row grow forever, and that adding a time period to the key fixes it.

for a middle

Explain why the unit is never split, what degrades as it grows, and how to size a bucket from write rate and row size.

for a senior

Design the read path across buckets, handle entities with very uneven rates, and catch unbounded growth in a schema review.

for a principal

Be ready to set sizing rules for a team's schemas and to weigh bucket granularity against read fan-out for the main dashboards.

## Why size matters for one unit Every wide-column store has a unit it never divides between servers: - in hash-partitioned stores, the **partition** — all rows sharing a partition key; - in range-split stores, the **row** — a range may split between rows, never inside one. That unit lives on one server (and its replicas). If it grows without bound: - **reads** that touch it do more work and use more memory; - **background merging** has to process it as a whole; - **repair and streaming** between replicas move it as a whole; - **hotspots** form, because all traffic for it hits the same servers; - stores **recommend or enforce ceilings** on the size of a partition or row, and operations degrade well before a hard limit. ## Where unbounded growth comes from The key groups by an entity whose data only accumulates: - `device_id` for sensor readings, - `channel_id` for chat messages, - `user_id` for an activity log, - a single row per device with a new column per reading. ## The time bucket Add a period to the **grouping** part of the key: | before | after | |---|---| | partition key `device_id`, clustering `reading_time` | partition key `(device_id, day)`, clustering `reading_time` | | row key `deviceId#time` in one row per device | row per `deviceId#day`, or one row per reading | Now each device-day is its own unit. Its maximum size is readings per day times reading size — something you can calculate and keep far below the limits. ## Choosing the bucket size 1. **Estimate the write rate** per entity at its busiest, not on average. 2. **Multiply by bucket duration and row size** to get the worst-case unit size. 3. **Pick a bucket** that keeps that well under the store's guidance. 4. **Check the reads**: a dashboard showing "last 24 hours" touches one or two daily buckets but twenty-four hourly ones. Too small a bucket makes reads fan out; too large brings the problem back. For entities with very different rates (one busy device among thousands of quiet ones) a size-based bucket, starting a new one after N rows and recording the current bucket elsewhere, is sometimes used. ## Reading across buckets A range read that spans periods now: 1. computes the list of buckets the range touches; 2. queries each (in parallel, or newest first until the limit is reached); 3. merges results in order. This is application logic the design must plan for. ## Interview angle The interviewer wants the causal chain — unit never split, so unbounded growth hurts one server — then the bucket as the fix, sized from the write rate, with its read-side cost.

  • How would you read "the latest 50 readings" when the data is bucketed by day?
    Start with today's bucket, read newest first, and step back one day at a time until 50 readings are collected or a lookback limit is reached. Most reads finish in one bucket; quiet devices may need several.
  • Does a time bucket also help with hotspots?
    Only partly. It bounds size, but all current writes for an entity still go to its current bucket. Spreading a very hot entity's writes needs splitting the current bucket further, which is a load-spreading technique rather than a size fix.

saying these in an interview costs you the question

  • Assuming a very large partition or row is split across servers automatically
  • Keying time-series data only by the source id with no time component
  • Sizing the bucket from average rather than peak write rate
  • Forgetting that range reads must now query and merge several buckets