An in-memory store is overloaded, so replicas each holding the whole keyspace are added — which load does that relieve, and which not?
answer
- same data on every machine
- reads multiply, writes do not
- the whole keyspace still fits where
- copies cost the primary propagation
- only the reads you may answer stale
basics
~20 sFan-out across copies buys read request capacity and nothing else. Writes still land on one primary and are then propagated to every copy, and each copy holds the whole keyspace, so neither write throughput nor memory headroom improves.
solid answer
~50 sEach **replica** holds the whole keyspace and can answer reads, so adding copies multiplies the rate of reads the tier can serve — and only that. Writes still go to a single **primary** and are propagated onward, so every copy repeats the work the primary already did; write capacity does not grow. Memory headroom does not grow either, because every machine holds all of the data and the working set must still fit on one of them. Each added copy also gives the primary more propagation to do, although some designs let a copy feed further copies so that cost does not grow with every one added. What buys write throughput and headroom is splitting the keyspace across nodes, a different arrangement with different failures. And the read capacity is real only for the reads you are willing to answer from a copy that is behind.
go deeper
Know that each replica holds the whole keyspace, so it can answer reads but adds no room for more data, and that writes still go to one primary.
Explain where each kind of load lands: read requests spread across copies, while every write is performed once on the primary and again on each copy.
Do the arithmetic before buying machines — read-to-write ratio, the share of reads that may be answered by a lagging copy, and the propagation each copy adds to the feeding side.
Decide which problem you actually have. If the ceiling is memory or write rate, copies cannot move it, and adopting them buys a second failure mode instead of capacity.
## Two different answers to "one machine is not enough" There are two ways to add machines to a volatile tier, and they relieve opposite problems. - **Copies of the whole keyspace.** Every added node holds all of the data. One node — the **primary** — accepts writes; the others are **replicas**, fed by it, and can answer reads. - **Pieces of the keyspace.** Every added node holds a share of the data, and each key lives on exactly one of them. A candidate who has only read about this calls both of them "clustering" and blurs the two. They fail differently and they buy different things, so the first move in any question of this shape is to say which one is on the table. This question is about the first. ## What the fan-out buys The read path gets more machines. If the primary was saturated answering read requests, that work now spreads across the copies, and the ceiling rises roughly with the number of copies you are actually willing to route reads to. That is the whole purchase. It is worth being precise about latency, because this is where the claim usually gets overstated. A copy does not make an individual read intrinsically faster: the per-operation work is the same on either machine. What it does is relieve **queueing** on a primary that had more requests than it could answer promptly, which shows up as a large improvement when the primary was the bottleneck and as nothing at all when it was not. A copy placed nearer the caller also cuts network time — and a copy placed farther away adds it. ## What it does not buy | Pressure | Relieved by copies? | Why | |---|---|---| | Read request rate | Yes | Each copy can answer reads independently | | Write throughput | No | One primary accepts every write, then each copy performs it again | | Memory headroom | No | Every copy holds the whole keyspace, so the working set still fits one machine | | Per-read service time | Not directly | The same work per operation; only queueing and distance change | | Tolerating the primary's loss | Not by itself | A copy is a candidate for promotion, which is a separate mechanism with its own delay | The memory row is the one people get wrong most often, and the reasoning is one sentence: nothing was divided. Six copies of a keyspace that does not fit on one machine are six machines on which it does not fit. The write row has a second half that is easy to miss. Copies are not free to the primary that feeds them: every write must be sent onward to each copy, so a write-heavy tier gets **more** loaded by the addition, not less. Designs differ in how much — some let a copy feed further copies, so the primary's outbound cost stops growing after the first — and that variation is worth naming rather than asserting either extreme. ## The capacity you can actually use Read capacity on copies is usable only by reads that may be answered by a copy running behind the primary. - A read routed to the primary for correctness never touches a copy, so it gains nothing. - The number for a capacity plan is not total read volume; it is the stale-tolerant share of it. - A plan that quotes the first number and delivers the second is how this purchase disappoints. ## How to decide, before buying machines 1. **Measure the mix.** What fraction of operations are reads, and what fraction of those may tolerate an answer that is a little old? 2. **Name the pressure.** Is the tier short of request capacity, of memory, or of write rate? Only the first is what copies address. 3. **Price the propagation.** Every copy adds work on the feeding side; on a write-heavy tier that can exceed what the copies relieve. 4. **Count what you gain a second failure mode for.** Copies introduce a lag, a routing decision and a second thing to operate. If the reads that need capacity are the ones that cannot tolerate lag, you have bought the cost without the benefit. ## The neighbouring arrangement, in one line If the answer to step 2 is memory or write rate, the arrangement that addresses it is cutting the keyspace into pieces so each node holds a share and accepts writes for it — with its own constraints on operations that touch several keys at once. That is a different subject with a different failure list, and the useful interview move is simply to say which of the two the situation calls for, and why.
- Does adding a copy make an individual read faster?Not intrinsically — the per-operation work is identical on either machine. It relieves queueing on a primary that was saturated, which looks like a large latency win when the primary was the bottleneck and like nothing when it was not. A copy nearer the caller cuts network time; one farther away adds it, and answers with an older value.
- How many copies is too many?The point where feeding the copies competes with the work they were added to relieve. Every copy adds outbound propagation on the feeding side, and reads you are not permitted to answer from a copy never spread anyway. Some designs cascade — one copy feeds others — which moves that ceiling rather than removing it. Count the routable reads before counting machines.
saying these in an interview costs you the question
- Says adding copies of the keyspace raises write throughput.
- Claims more copies give the tier room for more data.
- Assumes each added copy is free to the primary feeding it.
- Treats copies and split keyspaces as one 'clustering' answer.
- Counts read capacity that correctness-critical reads can never use.