An in-memory store got replicas for read capacity, yet half its reads are pinned to the primary for correctness — what is your call?
answer
- capacity you cannot use is not capacity
- count the pinned share first
- pinned by need, or by nobody deciding
- change the state, not the copies
- which guarantees you refuse to need
basics
~20 sFan-out only buys capacity for reads you will answer from a lagging copy, so each pin returns part of it. Measure which classes are pinned and why, then change what the tier holds or how it is split rather than adding copies.
solid answer
~50 sStart by naming what was actually bought: copies raise the ceiling only for reads that may be answered from a copy running behind the **primary**, so a pinned read uses none of it. Measure the split — what share of read volume is pinned, and which class each pin protects — because in most tiers the majority of pins are defensive rather than necessary, applied because nobody ever classified the reads. The real levers are upstream of the fan-out: reduce what this tier holds as the only copy, so fewer reads have to be current; separate the correctness-critical read class onto its own path; or accept that the pressure was never read rate and get headroom by splitting the keyspace instead. Then set the standing posture, so classification is a reviewed decision rather than a per-caller improvisation.
go deeper
Know that a read sent to the primary does not use the copies at all, so the capacity those copies added goes unused for that read.
Be able to compute the usable share: the read volume that may be answered by a lagging copy, rather than total read volume.
Find out why each class is pinned. Most pins are defensive rather than necessary, and the cheapest capacity available is a class that never needed the primary.
Set the standing posture: what this tier may hold as the only copy, the default staleness budget, who approves a pinned class, and which guarantees you refuse to depend on because they vary by product.
## What the fan-out actually bought Copies of the whole keyspace add capacity for one thing: reads that may be answered from a snapshot that is behind the primary. Every read routed to the primary for correctness is outside that purchase. So the usable capacity is not `copies x read rate`; it is the share of read volume you are willing to have answered by a lagging copy. When half the reads are pinned, half the purchase is idle, and that is not a tuning problem — it is a statement that the workload and the arrangement do not match. ## Establish the facts before the argument 1. **Measure the pinned share by volume, not by count.** Ten pinned classes that carry two percent of traffic are a rounding error; one that carries forty percent is the whole story. 2. **Attribute each pin to a decision.** For each pinned class, what decision does the currency of the answer protect, and what happens if it is wrong? 3. **Separate the two kinds of pin.** Some pins protect a decision that genuinely turns on the presence or absence of a key right now. Others exist because someone had an incident, could not classify the reads, and pinned everything the service does. The second kind is usually the larger, and it is free to recover. 4. **Check what the tier holds.** Reads have to be current mainly when the tier holds the only copy of something. That fact is a design choice, not a constraint handed to you. ## The levers, in order of how much they change - **Unpin what was never at risk.** The cheapest capacity available is a read class that was pinned defensively. Unpin it against a named staleness budget and record the reasoning, so the next incident does not re-pin it by reflex. - **Change what the tier holds.** If fewer records live here as the sole copy, fewer reads have to be current. That may mean the record's truth lives in a durable system and this tier holds a derived view of it — a larger change, and the one that actually moves the ratio. - **Separate the paths.** Correctness-critical reads and capacity-hungry reads have opposite requirements. Giving them different tiers is sometimes cheaper than making one tier satisfy both, provided you are honest that a second tier of copies has exactly the same lag. - **Change the arrangement.** If the pressure turns out to be memory or write rate rather than read rate, copies were never going to help; splitting the keyspace across nodes addresses those, with its own constraints on operations spanning several keys. ## Why this drifts toward everything pinned Pinning is always the safe individual choice and its cost is collective. The person who pins gains certainty; the tier loses capacity that nobody attributes to them. Left alone, the pinned share only grows, because unpinning requires someone to assert that a read tolerates being behind and to own that assertion. The organisational fix is to make the classification the reviewed artefact rather than the pin: - a read class is **stale-tolerant by default**, with a named staleness budget; - a new pinned class names the decision it protects and gets the same scrutiny as any other design change; - the pinned share by volume is a reported number, so the cost is visible to the people who ask for more copies. ## The posture worth setting At this altitude the question is not how to route these reads; it is which guarantees the organisation is willing to build on. Three are worth deciding once: - **Whether this tier may hold sole-copy state at all**, and under what review. That single rule decides most of the pinning downstream. - **What the default staleness budget is**, so that teams argue about a number rather than about whether staleness is acceptable in principle. - **Which distribution guarantees you refuse to depend on**, because they vary across the products you might buy. Stores in this class differ on whether a write is acknowledged before any copy holds it or only after, on whether that choice can be made per call, on whether reads may be routed to copies at all — and some ship no replication of their own. A design that depends on a behaviour that varies is a migration hazard the day someone changes the product underneath it. ## The answer an interviewer is listening for Not a routing tweak. The recognisable senior-to-principal move is to say that the fan-out's value is bounded by the stale-tolerant share of reads, to measure that share, and then to work on the share rather than on the machine count — including being willing to conclude that the copies were the wrong purchase and that the actual ceiling is somewhere else.
- When is the honest answer that the copies were the wrong purchase?When the genuinely correctness-critical classes dominate read volume, or when the pressure was never read rate. Copies add read capacity, a lag, a routing decision and another thing to operate; if the ceiling is memory or write rate, splitting the keyspace addresses it and copies cannot. Saying so is cheaper than tuning a fan-out that was never going to help.
- What makes 'pin it to the primary' spread until it covers everything?It is the safe individual choice with a collective cost: the person pinning gains certainty and the tier loses capacity nobody attributes to them. The fix is to make classification the reviewed artefact — a new pinned class names the decision it protects, the pinned share by volume is reported, and someone may ask whether the tier should hold that state at all.
saying these in an interview costs you the question
- Adds more copies when the pinned share is the actual problem.
- Accepts pins that nobody has ever justified or reviewed.
- Assumes sole-copy state on this tier is always acceptable.
- Quotes a fan-out figure without subtracting the pinned reads.
- Assumes every store in this class can route reads to copies.