Rows your cache holds are also written by another service and by operators; how do you stop it serving values no invalidation will clear?
answer
- who else writes this data
- no signal, no invalidation
- a marker column is a cheaper read
- the age becomes the contract
- detect by comparison, not by logs
basics
~20 sYou cannot invalidate a write you never see, so either consume a change signal from the other writer, check a cheap freshness marker before serving, or accept and publish a bounded maximum entry age. Otherwise do not cache that data.
solid answer
~40 sInvalidation only works for writes that pass through your code. A second service, a scheduled job, an import, an operator at a console, a restore or a replication stream all change rows without running anything of yours, so those entries have no invalidation path at all. The honest options are: consume a change event or change stream from the writer and treat it as the trigger; keep a version or last-modified column and read only that before serving a cached object; give entries a bounded maximum age and state that number as part of the contract; or leave the data uncached. Detection has to be active, since a stale hit emits nothing — sample served values against fresh reads and alert on the divergence rate and the age of the worst mismatch found.
go deeper
Take away the rule of thumb: you can only invalidate writes that go through your own code. If someone else can change the row, the copy can be wrong and nothing will tell you.
List the mechanisms and their costs — a change signal, a freshness-marker check, a bounded maximum age, or no cache — and say which of them need cooperation from the other writer.
Show the operational side: active divergence sampling because nothing errors, counters on invalidations issued versus applied, and a purge command plus a per-request bypass so an incident is not a redeployment.
Treat it as a boundary decision. A cache whose correctness depends on another team's write path is a coupling that must be documented and given a signal, or given up; otherwise the real budget is 'until someone redeploys'.
## The ownership question comes first Before choosing a cache for a dataset, answer one question: **do all writes to it pass through code you control?** If they do, invalidation is a programming problem — remove the entry before the write and again after the commit. If they do not, invalidation is not available at all, and no amount of care in your write path changes that. A second service, a scheduled job, an import, a data-fix script, an operator at a console, a restore, or a replication stream applying rows written elsewhere will all change the row without calling anything of yours. This is why "should we cache this?" is frequently a question about organisational boundaries rather than about performance. A dataset with several independent writers is a dataset whose cached copies have no reliable invalidation path. ## The signals you can actually get There are only a few honest ways to learn about a write you did not perform: 1. **Consume a change signal.** The writer publishes an event, or you read the engine's change stream, and treat that as the invalidation trigger. This is the strongest option, and it is only as reliable as the delivery: a signal you can silently miss is not a guarantee, so it still needs a backstop. 2. **Check a cheap freshness marker.** Keep a version or last-modified column on the row and read only that column before serving the entry. This converts a full load into a small one rather than eliminating it — worthwhile when the object is expensive to build, pointless when the row is small. 3. **Accept a bounded maximum age.** Give the entry a lifetime and state plainly that the data may be that old. This is the only mechanism that needs no cooperation from anyone, and it is the floor under the other two. 4. **Do not cache it.** Perfectly respectable, and the right answer whenever acting on an old value has consequences you would not accept. | Signal | Needs cooperation | Bounds staleness | Cost | |---|---|---|---| | Change event or change stream | yes, from the writer | to propagation delay | integration plus a delivery risk | | Freshness-marker check on read | no | to one round trip | a small query per read | | Bounded maximum entry age | no | to the configured age | staleness up to that age, by design | | No cache | no | perfectly | full cost on every read | ## Detecting what you cannot prevent A stale entry emits nothing, so detection has to be active. The technique that works is comparison: sample served values against a fresh read of the same rows, on a schedule, and record the divergence rate and the age of the oldest divergence found. That gives you two numbers a dashboard can carry and an alert can watch. Everything else people reach for is a proxy that does not measure this: hit rate does not fall when entries are wrong, eviction counts describe memory pressure, and error rates stay flat because nothing errors. It is also worth logging the invalidation path itself — how many removals were issued, how many were applied on each instance — because a silent drop to zero after a deployment is the most common way a working invalidation stops working. ## Reducing the exposure Design choices that shrink the problem, roughly in order of effectiveness: - **Cache the data whose writers you own**, and read the rest live. Divide the model along the write boundary, not along what happens to be slow. - **Prefer removal to replacement.** Removing an entry is safe under every outcome, including outcomes you did not model; replacing it asserts a value you may not be entitled to assert. - **Make the maximum age explicit and small enough to be honest**, then say the number out loud in the API's documentation. "Up to sixty seconds old" is a contract; "cached for performance" is not. - **Keep an escape hatch.** An operational way to purge a key or a whole region, plus a per-request bypass for support staff, turns an incident from a redeploy into a command. The uncomfortable conclusion is that for jointly-written data with no change signal there is no such thing as an invalidated cache — only a cache whose maximum staleness you have chosen, measured, and can defend. Where a signal does exist it shortens that number; it does not replace the need to state one.
- Why is a freshness-marker check a compromise rather than a fix?It still costs a query per read, so it removes the round-trip saving that a cache exists for; it pays off only when building the object is far more expensive than the small check. It also assumes every writer maintains the marker, which is the same cooperation problem in a smaller form.
- How do you decide the maximum entry age for data with writers you do not control?From the consequence of acting on an old value, not from load. Data feeding an irreversible action gets no cache; data on a screen a user has just edited needs a very small number; reference data tolerates a large one. Then state the number in the API contract instead of leaving it implicit.
- The other team offers to publish change events. What still needs a backstop?Delivery. An event you can silently miss during a restart or a partition leaves an entry nothing will clear, and no counter reports it. Keep a bounded maximum age underneath the events so a missed one self-heals within a known time, and monitor applied-versus-issued invalidation counts.
saying these in an interview costs you the question
- Assumes the database will notify the application of external writes
- Caches jointly-written data with no maximum age at all
- Expects hit rate to fall when entries diverge from rows
- Treats change events as guaranteed delivery with no backstop
- Thinks two services writing the same rows share a transaction
- Has no way to purge a key without a redeployment