skip to content

An entry the server understands as a collection has grown for eight months — do you split it across keys or cap it on write?

level: middleimportance: should knowfreq 54%

answer

  1. different remedies, different questions
  2. is the tail still wanted?
  3. bound at write time, not on a schedule
  4. splitting buys size, costs fan-out

basics

~20 s

Cap it when the tail is no longer wanted and the bound can be enforced on the write path. Split it across keys when every member still matters and only the single-entry packaging is the problem.

solid answer

~50 s

The two remedies answer different questions. Capping bounds the entry at write time: the writer adds the member and immediately drops whatever now falls outside the bound, so the entry can never grow past a size you chose. That is right when the tail is genuinely disposable. Splitting keeps everything but changes the packaging: derive part of the key from the data, such as a date segment or a numeric range, so one logical collection becomes many small entries and readers compute which ones they need. Splitting costs fan-out — a read that was one operation becomes several — and a migration, because the entry that already exists has to be spread out again while callers are reading it. If the server treats values as opaque bytes instead, the cap has to be applied by the application before it writes, and two writers can each trim a stale copy.

code

pseudocode · 9 lines
pseudocode
// capping: the bound is re-established on the same path as the append
store.append(key, member)
store.truncate(key, keep_newest = 5000)

// splitting: the bound moves into the key, so every entry is small by construction
bucket = day_of(event.occurred_at)
key    = "acme:audit:" + tenant + ":" + bucket
store.append(key, member)
// the reader now computes the buckets it needs and reads each one

go deeper

for a junior

Recall that a collection under one key grows until something stops it, because an unschematised keyspace has no declared size for anything. The two things that stop it are a bound applied when you write and a key that spreads the members out.

for a middle

Explain which remedy each situation calls for and what each one costs. Capping discards the tail and keeps readers unchanged; splitting keeps everything and hands the reader fan-out plus the knowledge of how a bucket is derived.

for a senior

Own the migration. The entry already exists and callers are reading it, so say how you spread out or trim it without a stall, how long both shapes coexist, and what the store's execution model does to that pass while it runs.

for a principal

The judgment call is whether the collection belongs under one key at all. An append-only sequence with a retention rule behind it is a durability requirement, and agreeing to bound it is often the moment to move it rather than to shard it.

## Two remedies, two different questions An entry is one key plus the value it addresses, and this one has been appended to for eight months by a write path that never asked how large it was getting. There are two honest remedies and they are not interchangeable, because they answer different questions about the data. **Capping** asks whether the tail is still wanted. If the collection is a recent-activity list, a rolling window of events, or a feed nobody pages past the first screen of, the oldest members are dead weight and the fix is to stop keeping them. **Splitting** asks, if everything really is still wanted, whether it has to arrive as one unit. If every member matters but no caller ever needs all of them at once, the entry's size is a packaging accident and the fix is to package it differently. | | capping on write | splitting across keys | |---|---|---| | what it assumes | the tail is disposable | every member is still wanted | | where the bound lives | in the write path | in the key | | what the reader changes | nothing | computes which entries it needs | | what it costs | the discarded members | fan-out, and a migration | | what it fixes | the entry's size, permanently | the size of any single operation | ## Capping is an invariant, not a cleanup A cap is only a cap if every write re-establishes it. The writer adds the member and, on the same path, drops whatever now falls outside the bound. Because the server understands the value as a collection here, dropping the excess is something the server can do without the value leaving the machine. The entry is over its bound for at most the gap between those two operations, and its size becomes a property of the design rather than a property of a schedule. The common wrong answer is a nightly job that trims collections which have grown. That fails twice. The entry is unbounded between runs, so the size you observe at three in the morning is not the size that was serving traffic at peak. And the job's own removal of the excess is proportional to however much accumulated, so the worse the drift, the more expensive the correction — the cleanup becomes its own oversized operation. ## Splitting moves the bound into the key Splitting means deriving part of the key from the data — a date segment, a numeric range, a bucket index — so that one logical collection becomes many entries, each small by construction. What you buy is that no single operation is ever large. What you pay: - **Fan-out.** A read that was one operation becomes one per bucket the caller needs, and every reader now has to know how the bucket is computed. - **More entries.** The same data spread over more keys costs more in total, because each entry carries the store's own per-entry overhead; what that overhead is belongs to the memory subject rather than this one. - **Topology.** Where the keyspace is spread across nodes, the buckets need not sit together and an operation spanning them may not be expressible at all. How keys are assigned to nodes is a separate subject; the point here is only that splitting can hand you that problem. - **A migration.** The entry that already exists has to be read out and spread out again while callers are reading it, which usually means writing both shapes and reading the new one with a fallback for a while. ## When the server does not understand the value Change that premise and the recommendation changes with it. On a store that hands values back as opaque bytes there is no partial write, so capping is done by the application: read the whole value, trim it, write it all back. The entry's size is paid twice on every single write, and two concurrent writers can each trim a stale copy, so the cap is not even reliable. Splitting becomes the stronger remedy in that world, because it is the only one that makes any individual operation small. A splitting-or-capping question that does not say which kind of store it is about simply has no answer, and this is the single variation most worth naming out loud in an interview. ## The third habit Both remedies are cheap before the shape ships and expensive after eight months of traffic, because both now require a migration of live data. The habit that prevents needing either is deciding up front what one key is allowed to hold — a byte bound and a member bound, written where the write path can see them. A store's own per-entry size ceiling is not that bound: where a store enforces one it sits far above the size at which the entry is already an operational problem, and those numbers differ by orders of magnitude between stores, so it is a backstop rather than a design input.

  • How does the answer change if the server treats the value as opaque bytes?
    Capping stops being server-side work. The application reads the whole value, trims it and writes it back, which pays the entry's size twice on every write and lets two concurrent writers each trim a stale copy. Splitting becomes the stronger remedy, because it is the only one that actually makes an individual operation small.
  • Why is a nightly cleanup job not a cap?
    A cap is an invariant enforced on every write; a job is a periodic correction. The entry is unbounded between runs, so peak traffic sees a size nobody bounded, and the job's own removal of the excess is proportional to whatever accumulated. Growth that outruns the interval is never bounded at all.

saying these in an interview costs you the question

  • Just delete the old members with a nightly job
  • Give the entry a deadline and the size takes care of itself
  • Splitting is always better, smaller entries are always faster
  • The store will refuse the write once the collection gets too large
  • Trimming in the reader is as good as bounding on write