skip to content

A durability standard fixes the copy count and acknowledgement rule on a cluster whose storage bill has doubled - which levers cut the total, and in what order?

level: seniorimportance: should knowfreq 44%

answer

  1. measure per stream before cutting
  2. attack the multiplicand first
  3. bytes at the source divide everything
  4. windows nobody chose are free wins
  5. posture changes are last and explicit

basics

~20 s

Measure per-stream first, then act on the multiplicand before the multipliers: fewer bytes per record, then retention windows nobody chose, then closed history moved to cheaper object storage, then dead streams retired. Changing the copy count is a posture decision, not a cost lever.

solid answer

~50 s

Start by finding where the bytes are - per-stream footprints on this kind of estate are heavily skewed, and a uniform cut wastes effort on the tail. Then work the product from the left: **(1) bytes per record**, because every byte removed is divided out of the whole multiplication and out of catch-up traffic too; **(2) retention windows nobody actually chose**, where a cluster default was inherited by streams that never needed replay, fixed with per-stream overrides; **(3) moving closed history to cheaper object storage**, which keeps the replay reach while taking those bytes off the expensive volumes; **(4) retiring streams nobody reads**, which still multiply happily. Only then does the copy count come up - and lowering it is a **change to the durability posture**, owned by whoever set the standard, made explicitly and never dressed up as an efficiency win.

go deeper

for a junior

The key idea to carry away is that the footprint is a multiplication, so removing bytes at the source is worth far more than it looks - each one avoided is avoided on every copy, every day.

for a middle

Be able to say which factor each lever attacks, and why the acknowledgement rule is not on the list at all: it governs what a write waits for, not how many copies are stored.

for a senior

Show the working order: measure per-stream first, take the levers that cost no capability, and treat a copy-count change as a posture decision that goes back to whoever owns the standard.

for a principal

The durable fix is a small set of tiers, each with its copy count, window and the footprint it implies, plus the review that stops a new stream inheriting the heaviest tier by accident.

## Find the streams that are the bill Before any lever is pulled, get per-stream stored bytes and per-stream ingest. Estates like this are almost always skewed: a handful of streams carry most of the footprint, usually because something large is being carried in each record, or because a generous retention window was inherited by a stream that never replays anything. Acting uniformly across every stream is the classic waste - it spends review time on the tail and leaves the head untouched. The frame for everything below is the multiplication itself: `stored bytes = ingest rate x retention window x copy count`. Every lever is an attack on one of those three factors, and the order below follows from two rules: **act on the multiplicand before the multipliers**, and **prefer levers that cost no capability**. ## The levers, in order 1. **Reduce the bytes in each record.** This acts on the leftmost factor, so it is divided out of the entire product, and it is the only lever that also reduces the catch-up traffic every copy carries. Typical wins are structural rather than clever: stop carrying a large document inside the record when a reference to where it lives would do, drop fields nobody consumes, and stop republishing whole entities where a change would do. Whatever byte you never write is a byte you do not store `copy count x window` times. 2. **Correct retention windows nobody chose.** A cluster-wide default quietly becomes the window for every stream created afterwards, including streams whose consumers are strictly online and would never replay an hour, let alone a week. Per-stream overrides that match the actual replay requirement are usually the largest single reduction available, and they cost nothing that anyone was using. 3. **Move closed history to cheaper object storage.** Where the platform supports it, the older part of a stream leaves the expensive volumes while remaining readable, so the replay reach is preserved and only the recent part carries the full multiplier. This changes where bytes live rather than how many exist, so it reduces the expensive line and not the total byte count. 4. **Retire streams nobody reads.** Dead streams multiply exactly as enthusiastically as live ones. Ownership is the hard part, not deletion. 5. **Keep the heavy posture only where it is earned.** A standard applied uniformly means the least valuable stream on the cluster carries the same copy count and window as the most valuable. Splitting the standard into a small number of tiers is not a weakening of it - it is what makes it survivable. ## Why that order | Lever | Which factor | What it costs you | |---|---|---| | Fewer bytes per record | the rate | producer work; nothing operationally | | Retention matched to real replay need | the window | replay reach you were not using | | Closed history to cheaper storage | where bytes sit | slower reads of old records | | Retiring dead streams | the whole product | nothing, once ownership is established | | Lowering the copy count | the copy count | durability - this is a posture change | One subtlety is worth stating because it is a common interview trap: **cutting record bytes and shortening the window by the same percentage remove the same number of stored bytes**. The reason the first is ranked ahead is not arithmetic. It is that reducing bytes also reduces the traffic each copy carries and takes away no capability at all, while shortening the window takes away replay reach that somebody may be relying on quietly. ## Savings that are not savings - **Changing what a write waits for.** The acknowledgement rule decides latency and the failure behaviour of writes. It does not change how many copies are stored, so it removes no bytes at all. - **Lowering the floor on caught-up copies required for a write.** Same error in a different place: that floor governs whether writes are accepted, not how many copies exist on disk. - **Compressing after the fact.** Whatever reduces bytes must do so before the record is written, or the multiplier has already been applied on the way in and the traffic already paid. - **Deleting a stream's records manually** instead of fixing the window that keeps recreating the problem next week. ## What is not on the table here Two things are deliberately out of this conversation. The first is the **rate** - how bytes stored and bytes moved are priced, and which boundary crossings are metered, belongs to whoever owns the platform's billing model; this exercise produces quantities. The second is the **posture** itself. If the arithmetic has been worked, the levers are exhausted and the total still does not fit, then the honest next step is to go back to whoever set the durability standard and say so. Lowering a copy count or a window because a bill is uncomfortable, without that conversation, is how a cluster ends up with a durability guarantee nobody has reviewed and everybody assumes.

  • How do you find which streams to act on?
    Pull per-stream stored bytes and per-stream ingest rate across the cluster and rank them. The distribution is normally heavily skewed, so a short list carries most of the footprint. Rank the list by bytes rather than by how noisy the owning team is, and take the effective window and copy count for each rather than the cluster defaults.
  • Why does moving closed history to cheaper object storage not reduce the byte count?
    Because it relocates bytes rather than removing them - the same records still exist and are still readable. What changes is that only the recent part of the stream occupies the expensive volumes and carries the full copy multiplier, while the older part sits somewhere cheaper. It buys cost, not capacity, and reads of old records get slower.
  • When is changing the copy count the right answer after all?
    When the arithmetic has been worked, the no-capability-cost levers are exhausted, and the data genuinely does not justify the posture it inherited. Then it is a durability decision: raised with whoever owns the standard, recorded with what it trades away, and applied as a per-stream override rather than quietly lowering the cluster default for everyone.

saying these in an interview costs you the question

  • Cuts the copy count first and calls it an efficiency win
  • Spreads the cut evenly instead of finding the streams that dominate
  • Shortens a retention window without asking who replays from it
  • Thinks the acknowledgement rule changes how many bytes are stored
  • Believes moving closed history to cheaper storage removes the replay reach
  • Thinks smaller records only help the writer, not stored bytes or catch-up traffic