skip to content

For time-series data in a wide-column store, when would you store one row per event rather than one row per entity and period holding many cells?

level: seniorimportance: nice to knowfreq 30%

answer

  1. tall and narrow vs wide
  2. per-row overhead
  3. bounded row size
  4. how reads slice the data
  5. updates to one event

basics

~20 s

One row per event is simpler, copes with uneven or unbounded streams and slices any time range. One row per entity and period packs events compactly and reads a period in one fetch, but must stay bounded in size.

solid answer

~50 s

Both shapes store the same events. **Tall and narrow**: one row per event, the time in the key, a few columns per row. It is easy to implement, handles very uneven or unbounded streams, and range reads are just key ranges. **Wide**: one row per entity and time period, each event a column or a timestamped cell in that row. It stores fewer rows, amortises per-row overhead, compresses better and reads a whole period in one row fetch, but the row must stay bounded, so the period must be sized from the write rate, and hot entities write to one row. In the hashed-partition shape the same choice appears as many clustering rows inside a partition versus a partition holding packed collections or serialised blocks. Choose wide when periods are read whole and rates are predictable; tall when rates vary or reads slice finely.

go deeper

for a junior

Know that events can be stored one row per event or packed into one row per entity and period.

for a middle

Explain what each shape costs in per-row overhead, compression, size limits and read slicing.

for a senior

Choose a shape from read granularity, rate skew and correction patterns, and design a roll-up from tall to wide when both are needed.

for a principal

Be ready to set schema conventions for a team's time-series data and to justify them in storage cost and operational risk.

## The two shapes | | tall and narrow | wide | |---|---|---| | unit per event | one row | one column or one timestamped cell | | key | entity + event time (+ id) | entity + period | | rows | many, small | few, large | | reading a period | scan a key range | fetch one row | Both answer "events for entity E between T1 and T2"; they differ in how the data is packaged. ## Why wide rows can be faster - **Per-row overhead**: every row repeats its key and bookkeeping; packing many events into one row pays that once. - **Compression**: similar values stored together compress better. - **One fetch**: a dashboard that always shows a whole hour or day reads a single row instead of thousands. ## Why tall rows are often safer - **No size ceiling to manage**: each row is one event, so a burst from one entity cannot overflow a row. - **Uneven entities are fine**: a device sending 1,000 times more than its peers needs no special handling. - **Fine-grained reads**: "the last 5 minutes" reads only those rows, not a whole period's row. - **Simpler updates and deletes**: correcting one event touches only its own row. - **Simpler code**: no period arithmetic to find where an event lives. ## Hashed-partition stores In stores with a partition key and clustering columns the question looks slightly different: a partition already is a sorted run of rows. The equivalent choice is between 1. **many small clustering rows** in a partition bounded by a time bucket (the usual design), and 2. **packing** a period's events into one row as a collection or a serialised block, which reduces per-row overhead but makes each update rewrite or append to the packed value. Bounding the partition still matters either way. ## Choosing 1. **How are reads sliced?** Whole periods favour wide; arbitrary short ranges favour tall. 2. **How even are write rates?** Predictable rates make wide rows safe; spiky or unbounded rates favour tall. 3. **Are events corrected after writing?** Frequent edits favour tall. 4. **Is storage cost dominant?** Very high volumes with whole-period reads favour wide. 5. **What does the team find simplest?** Tall designs are easier to get right first. A common pattern combines them: write events tall for ingestion, then roll completed periods into wide, compacted rows for long-term reads. ## Interview angle Show that you know both shapes, can name what each optimises, and choose from the read and write patterns rather than habit.

  • How can you get the benefits of both shapes?
    Ingest events as tall rows, then run a job that rolls each completed period into one wide or serialised row and deletes or expires the tall rows. Recent data stays flexible; history becomes compact and cheap to read.
  • What limits the period size of a wide row?
    The store's guidance on row or partition size and the busiest entity's write rate. The period must be short enough that the busiest entity's row stays well under that guidance.

saying these in an interview costs you the question

  • Choosing wide rows without checking the busiest entity's write rate
  • Assuming tall and narrow rows always perform worse
  • Packing events into one row when individual events are corrected often
  • Believing the choice only matters for storage cost, not for reads