skip to content

questions

3

Within one cluster, a stream's copy count goes from two to three - what does the third copy cost in stored bytes and traffic?

level: juniorimportance: must knowfreq 60%

answer

  1. two lines on the bill, not one
  2. disk and traffic both multiply
  3. two to three is fifty per cent
  4. one-off backfill, then forever
  5. the copy count is write amplification

basics

~20 s

The third copy adds a complete extra set of the stream's stored bytes - about fifty per cent more disk, since the base was two - plus one more transfer of every incoming byte, continuously, and a one-off backfill of everything already retained.

solid answer

~50 s

A copy is a full stored replica of the stream on one node, so raising the copy count from two to three means every retained record and every future record exists three times instead of twice. That is roughly a **fifty per cent** rise in that stream's disk, not a third, because the base is two. It also buys traffic twice over: a **one-off backfill** that moves the whole retained volume onto the new copy when you make the change, and then **catch-up traffic** carrying every incoming byte one extra time for as long as that copy exists. Where copies deliberately sit in different failure domains, that continuous carry is the part that can cross a metered boundary. What the third copy does not buy is ingest capacity, faster writes, or read capacity you can assume.

go deeper

for a junior

Recall that a copy is a whole second or third set of the same bytes, not a pointer to them, and that raising the count from two to three raises that stream's disk by about half.

for a middle

Explain both lines: stored bytes multiplied by the copy count, and one extra transfer of every incoming byte for as long as the copy exists. Separate the one-off backfill from the continuous carry.

for a senior

Show that you treat a copy-count change as a capacity event on a running cluster - the backfill lands now, on live hardware, and is proportional to what is already retained rather than to the ingest rate.

for a principal

Frame the copy count as a standard rather than a per-stream whim: which classes of data justify which multiplier, and who pays the resulting storage and traffic when the change is made across an estate.

## What a copy count actually multiplies A **copy** here means one full stored replica of a stream's data on one node, inside a single cluster - not a mirror held in a second cluster and not a restorable archive. The **copy count** is how many of those the cluster keeps for that stream. Raising it from two to three is not a bookkeeping entry: it adds another complete set of the stream's bytes. Every record already retained now exists three times, and every record that arrives from now on is written three times. Two consequences follow, and they are routinely confused with each other: - **Stored bytes.** The stream's footprint goes from twice its raw retained size to three times it. Relative to where you were, that is a rise of about **fifty per cent** - not a third. The base is two, and the most common arithmetic slip on this subject is dividing by the destination instead of the origin. - **Traffic.** Copies stay current by pulling **catch-up traffic** from the copy that leads the stream. A third copy means one extra transfer of every incoming byte, continuously. This is **write amplification**: one byte in becomes three bytes stored and two bytes transferred. ## The one-off and the forever The change has a cost the moment you make it and a cost that never stops. Teams reliably budget for neither, and almost never for the first. | What happens | When | Roughly how much | |---|---|---| | Backfilling the new copy | once, when the copy count is raised | the stream's entire retained volume, moved once | | Keeping the new copy current | continuously | one extra transfer of every incoming byte | | Holding the new copy's bytes | continuously | one extra full copy, for the whole retention window | | Rebuilding the copy after its node is replaced | on every replacement | the retained volume again | The backfill is the sharp edge: raising the copy count on a stream that already holds weeks of records schedules a large transfer immediately, on a cluster that is simultaneously serving live traffic. It is a capacity event, not just a settings change. ## "The disks are already provisioned" This is the usual objection and it is an accounting error rather than an argument. Capacity the third copy occupies is capacity nothing else can use, and it is occupied for the whole retention window - which is nearly always longer than the person approving the change is picturing. The same applies to the transfer: bytes moved between copies are charged wherever the transfer crosses a boundary that somebody meters, and the copies you deliberately place apart for survivability are exactly the ones whose catch-up traffic crosses such a boundary. Which boundaries are metered, and in what unit, is a platform-billing subject and not part of this arithmetic; what belongs here is that the multiplier is real and predictable before you commit to it. ## Where the shape of the claim changes This is a class of products, not one product, and the multiplier does not look the same everywhere: - Where a stream is split into parts, each with a leading copy and a **caught-up set** of followers, the copy count is a number you choose per stream and can change later - this is the case the arithmetic above describes directly. - Some designs keep exactly **one paired copy** rather than a configurable number. The multiplier there is fixed at two and is not a lever at all. - In **a design where shared durable storage replaces per-node copies**, the broker layer has no per-record multiplier. The store underneath keeps and charges for its own copies, but you neither set that count nor see it, so the lever you are reaching for does not exist at the broker. - Where a record is **deleted once it has been acknowledged**, there is no retention window to multiply by. The multiplicand is **backlog depth (stored undelivered bytes)**, which is small when readers keep up and enormous when they stop. ## What the third copy is and is not worth It buys tolerance of one more simultaneous loss among the nodes or volumes holding that stream, and it widens the margin before the cluster can no longer satisfy a write rule that demands several current copies. It does not buy ingest capacity - it consumes some, because each incoming byte is now written and transferred more times. It does not make writes faster; waiting for more copies can only make them slower. And it is not read capacity you can assume: many designs serve reads from the copy that leads the stream, while others can serve a reader from a nearby copy, so whether a third copy relieves anything on the read side depends entirely on the platform in front of you. The habit worth forming is small: before a copy count changes, write the multiplication down - rate, window, copies - and say the resulting number out loud. Nobody multiplies before choosing three, and that is the whole reason this question gets asked.

  • Does a third copy multiply the traffic that readers generate as well?
    No. Reader traffic scales with how many readers there are and how much each fetches, not with the copy count. The copy count multiplies the write side: stored bytes and the catch-up traffic that keeps copies current. On platforms that can serve a reader from a nearby copy, the copy count changes where reader traffic flows, not how much of it there is.
  • How does this arithmetic change where durability comes from shared storage beneath the brokers?
    The per-broker multiplier disappears - there is no copy count you set per stream, so there is nothing to raise from two to three. The underlying store still keeps several copies and charges for them, but that count is neither yours to choose nor visible to you, so the cost conversation moves entirely to how many bytes you store and for how long.
  • Why is raising the copy count on an old, large stream riskier than on a new one?
    Because the one-off backfill is proportional to what is already retained. A new stream backfills almost nothing; a stream holding weeks of records schedules a transfer of that entire volume, at once, on a cluster that is still serving live writes and reads. The steady-state cost is identical - the transition is not.

Keeping a third copy of an archive is not photocopying one page. It is renting a third filing cabinet for as long as the papers are kept, paying a courier once to fill it from the existing two, and then paying that courier again for every new page that arrives.

saying these in an interview costs you the question

  • Says a third copy costs a third more rather than half again more
  • Thinks copies only cost disk, forgetting the traffic each one carries
  • Assumes identical copies are stored once because the bytes match
  • Treats the extra copy as free because the volumes are already provisioned
  • Forgets the one-off backfill of everything already retained
  • Believes a third copy adds ingest or read capacity
open as a page

A stream ingests 400 GB a day, is kept for seven days, and the cluster keeps three copies - what stored-byte figure should planning use?

level: middleimportance: must knowfreq 56%

basics

~20 s

Multiply rate by window by copies: 400 GB a day times seven days times three copies is about 8.4 TB of stream data, then add headroom for per-record overhead, partial segments, peak days and rebuilding a copy - so provision well above the raw product.

open as a page

A durability standard fixes the copy count and acknowledgement rule on a cluster whose storage bill has doubled - which levers cut the total, and in what order?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Measure per-stream first, then act on the multiplicand before the multipliers: fewer bytes per record, then retention windows nobody chose, then closed history moved to cheaper object storage, then dead streams retired. Changing the copy count is a posture decision, not a cost lever.

open as a page