A restarted in-memory store can be filled from a point-in-time copy or by pre-loading a chosen key set. What does each buy?
answer
- one fills everything, one fills what you chose
- the copy is stale, the load is current
- restoring costs unavailability, not read load
- pre-loading is scheduled load you pace
- nothing to pre-load from without a system behind
basics
~20 sRestoring from a copy returns the whole keyspace as it was, with no read load behind the tier, but it costs time before the node serves and comes back stale. Pre-loading returns only what you chose, current.
solid answer
~50 sThey solve different halves of the same problem. Restoring from a point-in-time copy brings back everything the copy held, including entries you would never have thought to choose, and puts no read load on the system of record behind the tier - but it is only available if the posture kept a copy, in many designs the process does not serve until the load finishes, and what comes back is the keyspace as it was, not as it is. Pre-loading a chosen key set works on any store, including one that keeps nothing: you read the entries you care about from the system behind and write them in before opening to traffic. Everything loaded is current, and the read work is scheduled and paced by you instead of arriving all at once. They combine well: restore the bulk, then refresh the subset that must be current.
go deeper
Know that a restarted tier can sometimes be filled back from a copy that was kept, and otherwise has to be filled by writing entries into it again.
Explain the trade in both directions: the copy gives coverage but is stale and costs time before the node serves, while pre-loading gives current entries but only the ones you chose.
Show the operational half - holding the node out of rotation while it fills, pacing the pre-load against the headroom of the system behind, and combining the two mechanisms.
Frame it as which risk you are buying down: time to serve, freshness of what comes back, or load transferred behind the tier. Say which one this workload actually cannot afford.
## Two different fills After a restart the tier is empty. There are two ways to put entries into it before callers do it for you, and they are not variants of one another. - **Restoring from a point-in-time copy.** The tier reads back a copy of the keyspace taken earlier and comes up holding what that copy held. This is only available when the tier's **posture** kept one. Here you are a consumer of the copy: how the copy is produced, what it diverges by while it is being written, and how often it is taken are a separate subject. - **Pre-loading a chosen key set before taking traffic.** You decide which entries matter, read them from **the system of record behind the tier**, write them in, and only then let callers reach the node. ## What restoring from a copy buys, and costs It buys coverage and quiet: - You get **everything the copy held**, including the entries you would never have thought to choose. Nobody has to guess the working set. - The fill comes from the copy, so it puts **no read load on the system behind** at the moment of restart. - For entries with **no other home** - state that exists nowhere else - it is the only option there is, because there is nothing to pre-load from. It costs time and freshness: - **Recovery time.** In many designs the process does not accept callers until the load completes, so you have converted "empty and serving" into "unavailable and filling". That window grows with how much is stored, and on a large tier it can exceed the time an empty restart would have taken to settle. - **Staleness.** What the copy holds is the keyspace **as it was**, not as it is. For entries that are reconstructible copies of something durable, stale is harmless and will be corrected on the next write. For time-sensitive state, coming back stale can be worse than coming back empty, because an empty tier is obviously empty while a stale entry looks authoritative. - **Availability.** Not every store in this class offers a restore path at all. ## What pre-loading buys, and costs It buys currency and control: - It **works on any store**, including one that keeps nothing across a restart, because it uses the ordinary write path. - You **choose the set**: the busiest slice measured from live traffic, the entries belonging to active tenants, whatever the system behind can enumerate. - Everything loaded is **current as of the load**. - The read work it puts on the system behind is **scheduled work** - you pick the rate, the batch size and the hour - instead of arrival-driven work whose rate you do not control. It costs knowledge and delay: - You have to **know what to load**, and the honest source for that is measurement of the live keyspace, not intuition about which keys are hot. - It is still **load on the system behind**, just polite load; a pre-load that runs flat out is the same spike you were avoiding. - It **delays the node taking traffic**, and on a multi-node cycle that delay is multiplied by the number of nodes. - It can only reconstruct what is **reconstructible**. ## Side by side | | restoring from a copy | pre-loading a chosen key set | |---|---|---| | available when | the posture kept a copy | entries can be rebuilt from behind | | fills with | everything the copy held | only what you chose | | freshness | the keyspace as of the copy | current as of the load | | load on the system behind | none at restart | scheduled and paced by you | | delays serving by | the time to read the copy back | the time to load your chosen set | ## Choosing, and combining 1. Decide what the entries are. **Reconstructible copies** of durable data: either mechanism works, so choose on time-to-serve. **State with no other home**: there is nothing to pre-load from, so the copy is the only answer and "accept an empty restart" means accepting data loss rather than a latency blip. 2. Decide whether the node can be **held out of rotation** while it fills. If the routing in front of the tier lets you, do it - that is the difference between a controlled fill and a race against callers. 3. Consider doing both: restore to get the bulk back cheaply, then pre-load or refresh the subset that has to be current before opening to traffic. ## What varies between stores - Some stores keep nothing and have **no restore path**, so pre-loading is not one option of two - it is the only one. - Where restore exists, what dominates the time differs: reading back a whole copy and replaying an accumulated **write log** are not the same amount of work, and neither is proportional to the same thing. - Whether a node can be kept out of rotation while it fills is a property of the **routing in front of the tier**, not of the store.
- How would you pick the key set to pre-load?From measurement, not intuition. Sample what the live tier actually serves and take the slice that carries most of the reads, or enumerate an entity set the system behind can give you cheaply - active tenants, open sessions, the current catalogue. Then verify the choice by checking how much the load behind the tier drops after the pre-load finishes.
- When is restoring from a copy actively the wrong choice?When the restore takes longer than the empty-restart shock would have lasted, so you traded a degraded tier for an unavailable one; and when the entries are time-sensitive, because a stale entry that looks authoritative can be worse than no entry at all. Both are judgments about the data, not about the store.
- Does pre-loading avoid loading the system of record behind the tier?No - it reschedules that load. The same reads happen, but you choose the rate, the batching and the hour, and you can stop if the system behind shows strain. The gain is control over the shape, not removal of the work.
saying these in an interview costs you the question
- Assumes any in-memory store can be restored from a copy.
- Says pre-loading puts no load on the system behind.
- Forgets the node is often unavailable while it loads a copy.
- Treats entries restored from a copy as current.
- Picks the pre-load key set by intuition instead of measurement.
- Proposes pre-loading state that exists nowhere else to read from.