skip to content

Deciding Against One

The case for not adding a shared tier: a second failure domain and consistency story to run, a hop on the hot path, and a latency problem that was never where the time went.

on this pageshow

questions

4

A team adds a shared in-memory store purely because memory is fast; what does the architecture now have that it did not before?

level: juniorimportance: must knowfreq 64%

answer

  1. speed is one line of the bill
  2. a component, not just memory
  3. a hop on every consulting request
  4. availability now composes across two things
  5. entries can be gone next request

basics

~20 s

A second running component on the hot path: one more thing to size, secure, upgrade and be paged about, one more way a request can fail, and a store whose entries can be gone by the next request.

solid answer

~40 s

Speed is one line of the bill, not the whole invoice. A shared in-memory store is a separate process with its own address, so every request that consults it crosses a hop and depends on something that can be slow, full or unreachable. Somebody now sizes its memory, decides who may connect, upgrades it and answers the page when it misbehaves. Its entries are volatile by design: an entry present for the last request may be absent for this one, and after a restart the tier may be empty. Stores in this class also differ in what they keep across a restart and in whether they reclaim entries or start refusing writes at their memory ceiling. None of that argues against adding one; it is the list the decision has to be worth.

go deeper

for a junior

Recall that an in-memory store is a separate component with its own address, not a library inside the service. It answers from memory, so it is fast and bounded, and what it holds can disappear.

for a middle

Explain the mechanics of the price: a network hop on every consulting request including misses, a new component in the availability chain, and volatility that the calling code has to tolerate rather than assume away.

for a senior

Demonstrate that you price the operating model too - who runs the tier, in what shape, and what the team is signing up to watch, upgrade and be paged for. Say what the callers do when it is empty.

for a principal

Frame it as a dependency the organisation takes on permanently. Once traffic relies on the tier, removing it is a project, so the argument for adding it has to survive the years in which it is simply there.

## What the claim leaves out Memory answers a read far faster than a cold read from disk, and that is true. It is also only one line of the bill. Speed is a property of the **medium**; everything else on the invoice comes from the fact that the design has gained a **component**. An in-memory store deployed as a shared tier is a separate process with its own address, its own memory limit, its own version and its own ways of behaving badly. A weak answer stops at the speed. A good one continues to what the architecture now has to run. ## Four things the architecture did not have yesterday - **A hop on the hot path.** Every request that consults the tier crosses the network to it and back, on hits and on misses alike. Where the read it replaces is genuinely expensive, that trade is easy. Where the **system of record** - the authoritative holder of the data - was already answering from its own memory one hop away, the trade can be close to nothing. - **A second failure domain.** A request that must consult the tier can now fail in a new place. Availability composes: a path depending on two components is no more available than the two together, unless the design has somewhere else to go when one of them does not answer. - **A second thing to operate.** Someone sizes its memory, decides who may connect, upgrades it, watches it and is paged when it misbehaves. Who that someone is depends on the shape: a single instance you run yourself, a replicated pair somebody must fail over, or a **managed in-memory service**. The operational burden differs enormously across those three, and an honest cost argument says which one is on the table. - **A second answer to one question.** Two places now hold something about the same fact, and the design has to state which one a reader may believe and for how long. Writing that down is work; skipping it is the source of most arguments about the tier six months later. ## Misses pay the hop too The hop is charged per consulting request, not per hit. A request that finds nothing in the tier pays the hop *and* the original read at the system of record. That is fine when hits are the common case and the underlying read is expensive; it is a straight loss when reads rarely repeat. So the honest framing of the cost is not "one fast lookup added" but "one network round trip added to every request on this path, whatever it finds". ## Volatility is a standing condition, not an incident The defining property of this class of store is that entries can be gone. They can be gone because a deadline the store enforces passed, because memory ran short, or because the process restarted. The caller cannot tell those apart and should not need to: the code has to work when an entry that was there last time is not there now. That is the precondition behind the word **derived copy** - a faster copy of data a system of record still holds - and it is why the tier is not a place to keep the only copy of something without deciding to, deliberately, with the loss understood. ## What varies from store to store Nothing in this class is uniform enough to assume, and asserting one store's behaviour as the model is the standard way to get this wrong: | Question the design must answer | Why there is no single answer | | --- | --- | | What survives a restart? | Some stores in this class keep nothing across one by design; others can write what they hold to disk and reload it, at a cost while they serve. | | What happens at the memory ceiling? | Some reclaim entries to make room for a write; others refuse the write and return an error. Calling code has to survive both. | | Who runs it, and in what shape? | A single self-run instance, a replicated pair someone must fail over, and a managed service are three different availability and operations problems. | ## So when is the bill worth paying? When the reads it would absorb repeat, are expensive at the system of record, and arrive often enough that the hop is cheaper than the work it replaces - and when the design can still answer, in some reduced form, with the tier empty or absent. The negative case is not that a shared tier is a bad idea. It is that **fast** is an argument for the medium, while the thing actually being added is a component. ## Saying it in an interview Lead with the component, not the speed. Name what the tier would hold, name what still holds it authoritatively, then name the hop, the failure domain and the operating burden as the price. Interviewers listen for a candidate who treats added infrastructure as a decision with a cost side rather than a free improvement.

  • Does the argument change if the copy lives inside each application process instead of in a shared tier?
    Partly. A per-instance copy adds no hop and no component to operate, so most of the operational bill disappears. In exchange there are as many copies as there are instances, they can disagree, and nothing reconciles them. Which arrangement a design wants is its own comparison.
  • What is the smallest honest statement of when a shared tier is worth the bill?
    When the reads it absorbs repeat, cost the system of record real work, and arrive often enough that the added hop is cheaper than the work it replaces - and when the design still produces an answer, however reduced, with the tier empty or unreachable.

It is the second fridge in the garage. The drinks are closer, but you now own a second appliance that can be left open, be unplugged, or be empty when guests arrive - and the kitchen fridge still has to hold everything that matters.

saying these in an interview costs you the question

  • It is only memory, so there is nothing to operate.
  • Putting a tier in front cannot make the system less available.
  • Whatever was written to the store will still be there later.
  • The extra hop is free because the store answers in microseconds.
  • Memory is far faster than disk, so any slow read gets faster.
open as a page

A service reads a derived copy from a shared tier that stops answering; what decides whether its requests degrade or fail?

level: middleimportance: must knowfreq 58%

basics

~20 s

Three facts decide it: whether the entries are a derived copy with a system of record behind them, whether the read path has a route to that source, and whether the source can absorb the tier's full read rate.

open as a page

A team proposes a shared in-memory store to fix a 900 ms p99, though its system of record answers reads in 4 ms warm; what would you measure before agreeing?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Attribute the 900 ms: where the time goes, how many source reads a request makes, what baseline the speed claim uses, and what share of reads repeat. A tier removes only time spent in the reads it absorbs.

open as a page

You inherit a shared tier that eight services read on their hot path; how would you decide whether it still earns its failure domain?

level: principalimportance: should knowfreq 41%

basics

~20 s

Inventory the entries by role, not by reader. Derived copies go if their sources can take the load; entries that are the only home of a fact cannot go at all. Then price eight services sharing one availability number.

open as a page