skip to content

Never the Only Copy

Volatility is the precondition the tier rests on: the data must be derivable from somewhere else. What you accept when you accept that, and what changes when it quietly stops being true.

on this pageshow

questions

4

A design holds a derived copy of catalogue data in a shared volatile tier. What must be true of every entry, and who guarantees it?

level: juniorimportance: must knowfreq 76%

answer

  1. an assumption, not a product feature
  2. who else still holds this?
  3. absence is a normal reply
  4. the second branch of every read
  5. stores differ in how entries go missing

basics

~20 s

Every entry must be reconstructible from a system of record that still holds what it was made from. No store enforces that — it is a promise the design makes, and the caller must stay correct when an entry is absent.

solid answer

~50 s

Two things, and only one of them is about the tier. Each entry has to be reconstructible from a system of record that still holds what the value was made from — and nothing in the store checks this, so the guarantee belongs to the design, not to the product. In practice that means every read path has a second branch that reaches the system of record, and that branch must be correct and fast enough at whatever rate misses arrive. Absence is a normal reply rather than an error: on stores that reclaim entries under memory pressure any entry can go early; on stores that refuse writes once they hit their memory ceiling the loss reaches you as a failed write instead; and what a store keeps across a restart varies by store and by configuration. The design has to be right under all of them.

go deeper

for a junior

Remember the one sentence: the tier answers from memory, and something else must still hold the data. If a read finds nothing, that is normal, and your code needs a path that gets the value another way.

for a middle

Be able to show the second branch in a real read path and say what it costs. Explain that the store never checks derivability for you, and that entries can go missing for reasons other than a lifetime expiring.

for a senior

Demonstrate that you check the precondition rather than assume it. Name which entry classes in a running tier have a system of record behind them, and say what the read path does in the seconds after the tier comes up empty.

for a principal

Treat derivability as a reviewable invariant, not a habit: a rule that every entry class in the tier names what regenerates it, checked when a class is added. Be ready to say what the organisation loses when that rule is not written down anywhere.

## What the word "derived" is doing in that sentence Calling the entries in a volatile tier a **derived copy** is not a statement about the tier at all. It is a claim about the rest of the system: that for every entry the tier holds, some **system of record** still holds the material that value was made from, and that something in your code knows how to make it again. If the claim is true, the tier is an optimisation, and losing it costs latency and load. If it is false for even one class of entry, the tier is holding the only copy of a fact, and the same loss costs you the fact. That is the whole precondition, and every other comfortable thing people say about this tier rests on it — that it is safe to restart, safe to resize, safe to replace, safe to lose. ## Nothing in the store enforces it The store has no idea which of your entries has a source behind it. It accepts an entry, holds it while it can, and hands it back when asked. Concretely: - The store cannot distinguish an entry it may drop from one nobody can rebuild, so it will treat both alike. - **Absence is a correct reply**, not a fault. A read that finds nothing is the store working as designed. - **Stores in this class differ in how an entry goes missing.** Some reclaim entries when memory runs short — any entry, not only the ones you would have chosen. Others refuse the write at the ceiling instead and keep what they already hold, so the shortage reaches you as a failed write rather than a missing entry. Some let you choose between the two. - What a store keeps across a restart also varies by store and by configuration, and that is a separate subject with its own trade-offs. The precondition here is stronger than any answer to it: the caller has to be correct when the entry is simply not there, whatever the reason. So "it will still be there, we only wrote it a minute ago" is not an argument. It is the assumption the precondition exists to replace. ## What you are accepting when you accept it 1. **Every read path has a second branch.** Somewhere in the code, a read that finds nothing reaches the system of record and produces the value. If that branch does not exist, the design is not using a derived copy — it is using the tier as the home of the fact. 2. **The system of record can serve that branch.** Not one miss: whatever rate misses actually arrive at, up to and including all of the traffic if the tier is empty. 3. **No write path records a fact only in the tier.** The moment a value is written to the tier and nowhere else, that entry class has left the derived-copy role, whatever the tier is called in the architecture diagram. 4. **Adding an entry class requires an answer.** Somebody has to be able to say, for the new class, what regenerates it — and that answer has to be reviewable rather than remembered. ## The same outage, priced both ways | Event | Precondition holds | Precondition false for a class | |---|---|---| | The tier comes up empty | Slower responses and extra load on the system of record | Those facts are gone | | The tier is unreachable | Reads go to the system of record; answers stay correct | The feature depending on that class is down | | Routine memory pressure on a store that reclaims | Entries are produced again on the next read | Silent loss, with no error raised anywhere | | Resizing, moving or replacing the tier | An operational task in a maintenance window | A data migration with a correctness risk | The right-hand column is not a worse version of the left. It is a different system, with a different availability story and a different set of people who need to approve a change to it. ## Where the claim quietly turns out to be false Even in a tier that was designed as a derived copy, some classes drift out of it: - **Counters and running totals accumulated in the tier**, which were never computed from anything and cannot be recomputed from anything. - **Values produced from inputs that have since changed or been removed** — the source row still exists, but re-deriving now yields a different value than the one the entry held. - **Values fetched from a third party** that rate-limits or charges per call: derivable in principle, not at the rate an empty tier would demand. - **Entries written by a feature nobody reviewed**, often the smallest one, and often the one whose loss is most visible to a user. ## How to answer this in an interview Lead with the precondition, not with a latency number. A strong answer is roughly: this tier holds a derived copy, so every entry must be reconstructible from the system of record; that is an assumption we are making, not something the store provides; the read path therefore has a second branch, the system of record has to be able to carry it, and we should be able to name, class by class, what regenerates each entry. Then — and only then — is it worth talking about how fast the tier is.

  • If the precondition holds, what does an outage of the tier actually cost?
    Latency and load, not correctness. Reads reach the system of record instead, so answers stay right while responses get slower and the source takes traffic it was not sized for. Whether it can carry that traffic is the real question, and it is a capacity question rather than a data question.
  • Is it enough that the value came from the system of record when it was written?
    No. The precondition is about now, not about the past. What matters is whether the system of record can produce the value again today — the inputs may have changed or been removed, and a value accumulated inside the tier never had a source at all, however it first arrived.
  • Does giving every entry a lifetime make the precondition safe to ignore?
    It does the opposite of what people expect. A lifetime bounds how long an entry lives, which only guarantees more disappearances, not fewer. It says nothing about whether the value can be produced again, and an entry can also go missing well before its deadline on a store that reclaims under memory pressure.

saying these in an interview costs you the question

  • Says the entry will be there because it was written recently.
  • Treats an expiring lifetime as the only way an entry disappears.
  • Assumes the store guarantees the data is derivable.
  • Plans only for the tier being unreachable, never for it being empty.
  • Calls the tier the source of truth for entries nothing else holds.
open as a page

A shared tier answers 90% of 8,000 reads per second; the system of record is sized for 900. What does derivability promise here?

level: middleimportance: must knowfreq 66%

basics

~20 s

Derivability promises each value can be produced again, not that the source can serve every miss at the tier's rate. Here an empty tier sends 8,000 reads per second at a source sized for 900: an outage, not a slowdown.

open as a page

How would you find out which entries in a running volatile tier could not be regenerated if it emptied?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Go class by class, not entry by entry: for each writer, name what produces the value again. Three outcomes — a system of record does, nothing does, or something does but not the value that was there.

open as a page

A tier introduced as a derived copy now holds facts nothing else can produce. What has changed, and how do you get out?

level: principalimportance: should knowfreq 44%

basics

~20 s

No code changed; what changed is where those facts live. The tier's availability is now the service's durability, routine memory pressure can remove a fact nothing rebuilds, and resizing it became a data migration. Exit per entry class, never wholesale.

open as a page