In a wide-column design with a layout per read path, what must happen to the duplicated copies when a field that forms part of one layout's key changes?
answer
- a key cannot be updated
- delete old, write new
- you need the old value
- source of truth drives copies
- verify and rebuild
basics
~20 sA key cannot be updated, so every layout keyed by the changed field must delete the row at its old key and write it at the new one. That needs the old value, read from the source of truth or carried in the change event.
solid answer
~50 sIn a wide-column store a key is an address, so a change to a field that is part of a layout's key is a **move**: delete the row at the old key, write it at the new key. For example, orders keyed by `status` in an "orders by status" layout must leave the `PAID` slice and enter `SHIPPED` when status changes. That creates two needs. First, the writer must know the **old value** to find the old key: read the **source-of-truth** layout before writing, or carry old and new values in the change event. Second, every layout must apply the move, and a failure between the delete and the write leaves a missing or duplicated copy. Treat one layout as authoritative, derive the others from it (inline or from a change stream), make each move idempotent, and run a periodic comparison that can rebuild a drifted layout from the source.
go deeper
Know that changing a field that is part of a key means deleting the old row and writing a new one.
Explain why the writer needs the old value and what happens to reads if only the delete or only the write succeeds.
Design how layouts derive from a source of truth, with idempotent moves, durable intent, comparison and rebuild.
Be ready to set rules on which fields may appear in keys and when state changes should be modelled as appended events instead.
## Duplication has a sharp edge Query-first design stores the same data in several layouts, each keyed for a read path. Updating a **non-key** field is simple: write the new value into every copy under the same keys. Updating a field that is **part of some layout's key** is not an update at all. ## Why a key change is a move The key decides where the row is stored and in what order, and stores do not rewrite keys. So in each layout that uses the field: 1. **delete** the row at its old key (writing a delete marker); 2. **write** the full row at its new key. If only the write happens, the old copy lingers and reads return the row twice, once under each value. If only the delete happens, the row vanishes from that read path. ## You need the old value To delete at the old key you must know it. Options: - **Read before write**: fetch the current row from the source-of-truth layout, then issue the delete and write. This adds a read to every such update and opens a race if two updates interleave. - **Change events with before and after images**: a change stream that carries the old and new values lets a consumer move the row without reading. - **Keep the previous value in the row itself**, updated atomically with the new value where the store allows single-row atomic updates. ## Keeping all layouts in step | practice | why | |---|---| | one authoritative layout | a clear place to rebuild the others from | | derived layouts updated from its changes | one path of truth instead of many writers | | idempotent moves | a retried delete or write is harmless | | durable record of intent | an interrupted move can be completed | | periodic comparison | finds drift from bugs or partial failures | | rebuild capability | a drifted layout can be regenerated from the source | Multi-row atomicity, batches and safe write ordering are general wide-column limits covered elsewhere; the point here is that **a key field change touches every layout that uses it, twice**. ## Designing to reduce moves - Prefer **immutable** fields in keys: ids, creation times, categories that never change. - Model a changing state as **events**: instead of moving an order between status slices, append a status-change row and read the latest. - For a layout that serves "items currently in state X", accept that it is a **work queue** with frequent deletes, and size it for the delete markers that result. ## Interview angle Strong candidates spot that a mutable key field turns updates into moves, name the need for the old value, and describe how the source of truth, idempotent moves and reconciliation keep layouts consistent.
- What goes wrong if two status updates for the same order run read-before-write concurrently?Both read the old status, both delete that key, and each writes its own new key, so the order appears in two status slices. Serialise changes per order, for example with a version check on the source row or one consumer per order key in the change stream.
- How does an event-based model avoid the move?Instead of relocating the row, you append a new row recording the change, keyed by entity and time. Reads take the latest state per entity. Nothing is deleted, so no copy can be left behind at an old key.
saying these in an interview costs you the question
- Updating a key field as if it were an ordinary column
- Writing the new copy without deleting the old key in every layout
- Keying a layout by a field that changes frequently without planning for deletes
- Having several services write each layout independently with no source of truth