For a table in a shared identifier-keyed cache, how does its write rate decide the entry strategy?
answer
- ratio first, mechanism second
- never written needs no invalidation
- guarding turns readers into missers
- only enrolment commits with the row
- parity ratio means do not cache
basics
~20 sWrite rate picks the mechanism: data never written is stored once; rarely written rows need entries dropped or replaced at commit, and guarded during the write when mid-write reads matter. At parity, do not cache.
solid answer
~50 sThere is a short ladder, and the write rate says which rung you need. Data that is never written after load can be stored once with no invalidation path at all. Rows written rarely but read constantly want the entry dropped or replaced when the write commits — cheap, with a brief window in which a reader can still be handed the previous state. If a reader must not observe that window, the entry is guarded for the duration of the write: it is marked in progress, readers miss and go to the database instead, and the mark clears at commit or rollback. The strictest rung enrols the cache update in the same transaction, so the entry and the row become visible together at the cost of coordination on every write. At roughly equal read and write rates, the honest answer is not to cache the type.
go deeper
Remember the extremes: data nobody ever writes is easy to cache, and data written as often as it is read should not be cached at all.
Explain the rungs as mechanisms — store once, drop or replace at commit, guard during the write, enrol in the transaction — and what each promises when the write commits.
Demonstrate that you choose by measured ratio and by what a reader must never observe mid-write, and that you can defend refusing to cache a hot, frequently written table.
Weigh the coordination cost of the strictest rung against the value of the guarantee, and set a house default so teams do not reach for transactional coupling out of nervousness.
"Which strategy?" is really two questions: how often is this row written, and what must a concurrent reader be prevented from seeing? A shared identifier-keyed cache offers a short ladder of mechanisms, each costing more than the one below it and promising more at commit. ## The ladder 1. **Store once and never invalidate.** The type is treated as never written after it is loaded. The entry is written on first read and served until the process ends or capacity evicts it. There is no invalidation path, because the mechanism assumes there is nothing to invalidate. If the data does change, only a purge or a restart puts it right. 2. **Drop or replace the entry at commit.** The write goes to the database; when it commits, the layer removes the entry or writes the new state into it. Cheap, and adequate for most read-mostly data. Between the row changing and the entry going there is a window in which a concurrent reader can still be handed the previous state. 3. **Guard the entry for the duration of the write.** Before the write, the layer marks the entry as *in progress*. While that mark stands, readers do not use it — they miss and read the database, rather than being handed either the pre-write value or a half-applied one. The mark is cleared at commit, with the entry dropped or replaced, and at rollback, leaving the entry as it was. The cost is extra misses during the window and a store that supports the marking. 4. **Enrol the entry in the transaction.** The cache update takes part in the same commit as the database change, through a coordinating protocol, so the entry becomes visible at the moment the row does and a rollback undoes both. This is the only rung on which the entry has the same commit semantics as the row, and it needs a store that can participate plus coordination cost on every write. | mechanism | assumes | what a concurrent reader can be handed | cost per write | |---|---|---|---| | store once | the data is never written | the stored state, indefinitely | none | | drop or replace at commit | writes are rare | the previous state, briefly | one store operation | | guard during the write | writes are rare and the window matters | nothing from the cache; the read misses | mark, clear, plus extra misses | | enrol in the transaction | the entry must commit with the row | only committed state | coordination on every write | ## Reading the write rate - **Never written after load** — fixed reference data, code tables, records that are only ever inserted. Rung 1. - **Written rarely, read constantly**, hundreds or thousands of reads per write — rung 2 or 3. Choose 3 when a reader being handed the pre-write state for the duration of a long write would be a real defect; choose 2 when it would not. - **Written at roughly the rate it is read** — no rung helps. Every write pays invalidation and nearly every read pays a miss, so you have bought a memory footprint, an extra moving part and a new way to be wrong in exchange for very little. The right answer is not to cache the type. - **Write-heavy** — the same, only worse. An entry voided more often than it is read is pure overhead. Note what the ladder does *not* decide: how long a stale copy may survive, or how invalidation is ordered around the commit. Those belong to invalidation. The strategy only says which guarantee the mechanism attempts. ## Drop or replace? Dropping is the conservative default: the next reader misses and repopulates from the committed row. Replacing saves that miss and is tempting for a type read immediately after every write, but it pushes state into a shared store from a transaction that has not necessarily finished — which is precisely what rungs 3 and 4 exist to make safe. If you replace, do it under a mechanism that guards the window. ## What an interviewer is listening for - that you ask for the read-to-write ratio before naming a mechanism - that you describe mechanisms rather than reciting one product's list of strategy names - that you know "guarded" makes readers **miss**, not block, and not read the old value - that you are willing to answer "this table should not be cached" when the ratio says so - that only the transactional rung makes the entry and the row become visible together
- What does guarding an entry during a write actually do to a concurrent reader?It makes the reader miss. While the in-progress mark stands, the entry is unusable, so the read falls through to the database and sees whatever the engine's isolation gives it. Readers are not blocked and are not handed the pre-write copy; they simply pay the cost of a normal load.
- Why is treating a type as never written risky even when nobody writes it today?Because that rung has no invalidation path at all. If a row is ever changed — by a release, a data fix or another application — nothing in the layer notices, and the entry serves the old state until the process restarts or someone purges it by hand. Use it only where immutability is enforced, not merely observed.
- Two types have the same read-to-write ratio but you pick different rungs. What justifies that?The consequence of a reader seeing the pre-write state during the write. For a display counter, a brief window costs nothing and the cheap rung is right. For a value that gates a decision, that window is a defect, so the entry is guarded for the duration of the write or enrolled in the transaction.
saying these in an interview costs you the question
- Picks a strategy without asking the read-to-write ratio
- Thinks guarding an entry blocks readers instead of making them miss
- Believes dropping the entry at commit removes the window entirely
- Caches a table written as often as it is read and expects a gain
- Assumes replacing the entry on write is always safer than dropping it
- Treats a rarely written table as never written, so no invalidation path exists