In-Memory Store Concepts
The volatile tier as a component you run: what it holds, what it drops under pressure, what it loses on a restart, and what one round trip costs. Not one product's command set.
on this pageshowhide
explore
- The Volatile Tier20 questions
- Memory as the Medium4 questions
- Never the Only Copy4 questions
- Four Roles, One Component4 questions
- Deciding Against One4 questions
- In-Process vs Shared4 questions
- Keyspace Design21 questions
- Naming & Namespacing4 questions
- Cardinality & Key Size4 questions
- Oversized Entries4 questions
- Traversing a Live Keyspace4 questions
- Secondary Access Paths5 questions
- Server-Side Value Shapes24 questions
- Opaque Bytes vs Structure4 questions
- Atomic Counters5 questions
- Maps, Sets & Lists5 questions
- Score-Ordered Sets5 questions
- Probabilistic Sketches5 questions
- Expiry Semantics21 questions
- Attaching a Lifetime4 questions
- Absolute vs Sliding4 questions
- When Memory Comes Back5 questions
- Time vs Pressure4 questions
- The Immortal Entry4 questions
- The Memory Ceiling26 questions
- Per-Entry Overhead4 questions
- Compact Representations4 questions
- Fragmentation & Resident Size5 questions
- Sizing the Working Set4 questions
- Refuse, Evict or Die4 questions
- Eviction Policy Families5 questions
- Atomicity & Concurrency21 questions
- Serialized Execution4 questions
- The Overwritten Change4 questions
- Grouping Several Steps4 questions
- Optimistic Check-and-Set5 questions
- Server-Side Logic4 questions
- Durability Options20 questions
- Three Honest Postures4 questions
- Point-in-Time Copies4 questions
- The Write Log4 questions
- Naming the Loss Window4 questions
- Restart & Warm-Up4 questions
- Spreading the Tier30 questions
- Replicating the Tier5 questions
- Reading a Replica4 questions
- Automatic Failover5 questions
- Split Brain4 questions
- Key-to-Node Assignment4 questions
- Operations Across Partitions4 questions
- More Than One Region4 questions
- The Access Cost Model21 questions
- The Round Trip4 questions
- Batching & Pipelining4 questions
- Pooling and Limits4 questions
- Blocking Operations5 questions
- One Slow Operation4 questions
- Ephemeral State Workloads21 questions
- Sessions Between Instances3 questions
- Locks & Leases5 questions
- Counters & Quotas4 questions
- Deduplication Records4 questions
- Transient Fan-Out5 questions
- Operating the Tier26 questions
- The Signals That Matter5 questions
- Diagnosing a Slowdown4 questions
- Finding the Memory5 questions
- One Tier, Many Users4 questions
- Scaling & Upgrading4 questions
- Who Can Reach It4 questions
questions
251 · 11 sectionsA team adds a shared in-memory store purely because memory is fast; what does the architecture now have that it did not before?
basics
~20 sA second running component on the hot path: one more thing to size, secure, upgrade and be paged about, one more way a request can fail, and a store whose entries can be gone by the next request.
How does a store that answers every read from RAM differ from a disk engine holding most pages in RAM?
basics
~10 sAn in-memory store guarantees every read it serves is a memory read; a disk engine usually gets one. The gap shows in the tail - a miss costs the engine a device access.
A design holds a derived copy of catalogue data in a shared volatile tier. What must be true of every entry, and who guarantees it?
basics
~20 sEvery entry must be reconstructible from a system of record that still holds what it was made from. No store enforces that — it is a promise the design makes, and the caller must stay correct when an entry is absent.
A team runs one in-memory store for several jobs - what four roles can it fill, and which one is a cache?
basics
~20 sAn in-memory store is rented for four jobs: a derived copy of data a system of record still holds, the sole home for ephemeral state, a coordination point, and a transient transport. Only the first is a cache.
A service reads a derived copy from a shared tier that stops answering; what decides whether its requests degrade or fail?
basics
~20 sThree facts decide it: whether the entries are a derived copy with a system of record behind them, whether the read path has a route to that source, and whether the source can absorb the tier's full read rate.
Two teams share one flat keyspace and one writes entries under the key user:1042 — what should that key carry instead?
basics
~20 sOn a shared flat keyspace the key string is the entire addressing model, so it must say whose it is: an owning prefix, then the entity, then the identifier, and a format version where the value's shape may change.
A store answers only by key, but a service must find an account by email address. What has to exist, and who maintains it?
basics
~20 sA second entry: an application-maintained index keyed by the email address, holding the account's key. The application writes, updates and repairs it, and nothing in the store's base data model builds such a lookup or notices when it is wrong.
Why is asking a serving in-memory store for every key matching a prefix in one operation not a harmless query?
basics
~20 sA whole-keyspace listing costs the distinct-key count, not the number of matches, and the reply is assembled whole before anything is sent. Where the store executes one operation at a time, every waiting caller pays it.
A key scheme in an in-memory store mints one entry per user per day per field. How do you size it before shipping?
basics
~20 sMultiply the factors out to a distinct-key count, then price each entry as key bytes plus value bytes and multiply again. Compare the result with the tier's budget, and identify which factor is unbounded before arguing about the bytes.
On a store that hands values back as opaque bytes, what does one entry grown to hundreds of megabytes cost to read, write and remove?
basics
~20 sEvery operation on that entry is proportional to its whole size: a read ships all of it, any change rewrites all of it, and removing it is real work that someone has to pay for.
On a store that can address named fields inside a value, what does writing one field buy over a whole-value read-modify-write round trip?
basics
~10 sWriting one field sends only that field and skips the read: one round trip instead of two, and two callers changing different fields of the same entry stop overwriting each other.
Why do in-memory stores offer an operation that adds to a stored number instead of leaving the arithmetic to the caller?
basics
~10 sServer-side arithmetic means the caller never holds the old number: one round trip replaces a read-modify-write round trip, and two concurrent callers each contribute a change instead of one silently overwriting the other.
A store hands back exactly the bytes it was given, so what does changing one field of a stored value take?
basics
~20 sChanging one field costs a read-modify-write round trip: read the whole value, deserialize it, change the field, serialize it again, write the whole value back. Two operations where a store that can address parts of a value needs one.
In a score-ordered set, where does the ordering come from, and what does it turn into a read instead of a sort?
basics
~20 sThe caller supplies a number - the ordering score - with every member, and the store keeps the members arranged by it as they are written. Position then becomes a read: the rank of one member, or a slice by position or by score band.
Should a user's 200 live room memberships, each checked individually, be one member collection under one key or 200 keys?
basics
~20 sOne member collection: the checks are membership tests the server answers without moving the collection, and the group is one unit of work. Choose separate keys only when a membership must expire on its own.
How does an entry with a fixed deadline behave differently from one with an idle deadline pushed forward by every access?
basics
~20 sA fixed deadline is set when the entry is written and never moves, so total life is bounded no matter how busy the entry is. An idle deadline is pushed forward by access, so it removes only entries that have gone quiet.
An entry's deadline passed ten minutes ago and nothing has read it since — what does a reader get, and is the memory back?
basics
~20 sA reader is served nothing: past its deadline the entry is never handed back. The memory is usually still held — stores free an expired entry when something touches it, a sweep reaches it, or the room is needed.
An entry written with a one-hour lifetime is missing ten minutes later - what else removes entries besides a deadline?
basics
~20 sTwo different events remove entries. Expiry is removal the application asked for by attaching a lifetime; eviction is removal the store forced on it to free room. Only the second can take an entry early.
A service needs an entry gone ten minutes after it is written. What does attaching a lifetime buy over deleting it later from application code?
basics
~20 sA lifetime attached at write makes cleanup the store's obligation: it stops serving the entry at the deadline even if the writing process crashes, is redeployed, or never reaches its delete path. Application-side deletion survives none of that.
Your design wants an entry kept alive while it is being read — which side actually pushes the deadline forward, and what does that cost?
basics
~20 sUsually the caller, not the store. Many stores do not move a deadline on a read, so an idle deadline is an extra write on every access, turning a read-heavy path into a write-heavy one.
A running in-memory store exhausts the memory available to it: what three outcomes can follow, and how does each appear to the caller?
basics
~20 sMemory exhaustion has three configured endings: the store refuses writes while reads still work, it removes entries to make room, or the operating system kills the process. Which one you get was a choice, usually a default nobody made deliberately.
What does one entry in an in-memory store cost beyond the bytes of its key and value?
basics
~20 sAn entry also costs a slot in the store's lookup structure, a small header of lengths and pointers, a lifetime field or record if it carries one, and whatever the allocator rounds each of its several allocations up to.
Only 5 million of a store's 200 million entries, all copies of database rows, are read hourly — which number sizes it?
basics
~20 sSize for the working set — the entries actually touched in a window — not total data. A database holds the rows, so an absent entry costs one extra read, and memory bought for untouched entries is never read.
An in-memory store's memory steps sharply when its small maps each gain their hundredth field — which two representations explain that step?
basics
~20 sMany in-memory stores keep two representations of one value: a compact form that packs elements together, and a general form with a lookup structure per element. Crossing the promotion threshold swaps them, so memory steps rather than climbs.
Why does a store report 2.4 GB of data size for 10 million entries whose payloads total 800 MB?
basics
~20 sBecause the fixed per-entry cost is charged 10 million times. Data size over entry count is 240 bytes, of which 80 is payload and 160 is lookup slot, header, key bytes and allocator rounding. Overhead scales with count, not bytes.
An in-memory store applies four submitted operations as one uninterleaved group; what does that promise, and what does it not?
basics
~10 sIt promises only that no other caller's operation is applied between the four. It is not a database transaction: there is no rollback, so when the third step fails, the first two stay applied.
Two callers read one entry holding a quota count, each adds ten, and both write back — what is stored, and what is reported?
basics
~20 sThe second write lands whole and the first caller's addition is gone. An entry that held 100 holds 110, not 120 — and nothing is reported: both callers were told their write succeeded. That default outcome is last-writer-wins.
A write presented with the version token read alongside the value is refused - what happened, and what must the caller do next?
basics
~20 sAnother caller changed that entry between the read and the write, so the store refused it instead of overwriting. The caller must re-read the entry, recompute the change from the new value, and write with the new token.
If a store makes every operation atomic, why can two callers that read one entry, change it and write it back still lose a change?
basics
~20 sPer-operation atomicity covers one operation, not a caller's sequence. A read and a write are two operations, and other callers' operations run in the gap between them, so the later write lands whole and the earlier change is gone.
Two stores both guarantee that one operation on one entry is atomic - how does run-to-completion execution deliver that, and how does per-entry locking?
basics
~10 sRun-to-completion executes one operation fully before starting the next, so nothing interleaves. Per-entry locking lets worker threads run concurrently, each holding the entry it operates on for that operation's duration. Same promise, different consequences.
A team says their in-memory tier "has persistence turned on" — what must you still ask before you can say what a crash costs?
basics
~20 s"Persistence is on" is not a loss window. Ask which posture is in force and the flush policy or copy interval that governs it — those numbers, not the word, say how many seconds of acknowledged writes a crash costs.
An in-memory store restarts empty and takes traffic again. What does the system of record behind it experience, and how large is that effect?
basics
~20 sUntil entries accumulate again, every read lands on the system of record behind the tier. The multiplier is the inverse of the miss rate: at a 95% hit ratio that system briefly sees twenty times its normal read load.
A store writes a whole copy of its keyspace to disk every 15 minutes while still accepting writes; which writes does that copy hold?
basics
~20 sA point-in-time copy holds the keyspace as of the instant it was cut, not when its file finished writing. Writes accepted after the cut are absent, so up to one whole copy interval of acknowledged writes can be missing.
What does an in-memory store hold after a restart if it keeps nothing, a periodic copy, a write log, or both?
basics
~20 sKeep nothing gives an empty store. A periodic whole copy gives the keyspace as of the last copy. A write log gives it replayed to the last flush to disk. Keeping both replays the log on top of a copy.
When an in-memory store is configured to keep a write log, what is appended while it runs and what happens to that log at start?
basics
~20 sA write log records every state-changing operation in the order it was accepted. At start the process replays that log from the beginning, applying each entry again, to rebuild the keyspace it held before it stopped.
A tier whose nodes each hold a share of the keyspace, not a copy: what limits an operation naming three keys?
basics
~20 sOne operation runs on one node, so it can only name keys that are all in the same partition. Split across partitions, the call is refused, composed by an intervening layer, or has to be split by the caller.
A volatile tier acknowledges a write before any replica holds a copy of it: what has the caller been promised?
basics
~10 sAcknowledgment means one node has the write in memory, nothing more. Until it propagates, the write exists in exactly one place, and if that node is lost the write is simply gone.
In a volatile tier with a primary and two replicas of the whole keyspace, why must a replica be promoted before writes resume?
basics
~20 sA replica holds a copy of the keyspace but is not the address that accepts writes, so writes resume only after some deciding party promotes one of them to primary. That promotion is a deliberate step, and it takes time.
When a keyspace is split so each key lives on one node, why must every caller resolve keys the same way?
basics
~20 sAssignment must be a deterministic function of the key, evaluated identically by every caller. If two callers disagree about where a key belongs, each reads and writes its own copy on a different node and neither sees the other's.
A caller updates a key on the primary of an in-memory store, then reads it from a replica and sees the previous value — why?
basics
~20 sA replica answers from its own copy, which trails the primary by the propagation lag, so a read issued inside that window returns the value the key held before the write. The write is not lost, only not yet visible there.
An in-memory store answers a lookup in microseconds, but the caller opens a fresh connection for every call - what does each call pay?
basics
~20 sConnection setup, paid before any work begins: at least one network round trip for the handshake, plus further trips where an encrypted transport or a credential exchange is required. Microseconds of server work end up buried under milliseconds of setup.
A request makes 500 single-key reads, each answered in microseconds, yet takes 300 ms overall — where did the time go?
basics
~20 sAlmost all of it went into the network. Each read pays one full crossing, so 500 sequential reads pay 500 crossings. The store's microseconds are far too small to explain the delay; the number of crossings is the cost.
One write in a run of two hundred sent to a volatile tier without waiting for replies fails — what happened to the others?
basics
~20 sThe other one hundred and ninety-nine were executed and their effects stand. Sending a run of operations without waiting for replies buys latency and nothing else: no atomicity, no isolation, no rollback — each operation succeeds or fails on its own.
A request needs forty values from a volatile tier one zone away — how does a pipelined run differ from one operation taking all forty keys?
basics
~20 sBoth collapse forty round trips into one, and that is all they share. A pipelined run stays forty independent operations with forty replies matched in order; one operation taking many keys is a single operation with a single reply and a single outcome.
While a caller is parked on a wait-with-deadline call, what does it consume on the store and what does it not?
basics
~20 sA parked caller consumes a connection slot for the entire wait and essentially no server processing time. The store registers the waiter and goes on serving everyone else, so every store-side signal stays flat while the caller's connection is unusable.
A quota counter on a shared volatile tier uses one key per subject per period and never gets a lifetime - what does that cost?
basics
~20 sDead counters accumulate. The keyspace grows with the whole history of subjects and periods rather than with the active ones, and on a tier that removes entries under memory pressure those dead keys crowd out live ones.
Why does a lease key that lets only one worker run a job carry a lifetime instead of living until its holder releases it?
basics
~20 sA lifetime bounds a crash. If the holder dies between claiming and releasing, an entry that only a release deletes blocks the job forever; the lifetime lets the store drop the claim so another worker can take it.
Why must per-user request state leave the application instance once a second instance exists, and what does the move cost?
basics
~20 sState held in one instance's memory is readable only by that instance, so the next request may land on a stranger. A shared tier makes all instances equivalent, at the price of a round trip per authenticated request.
A price feed is broadcast to recipients attached to a shared volatile tier; one drops for two seconds — what does it receive on reconnect?
basics
~20 sNothing from those two seconds. A broadcast reaches only the connections attached at the instant it is sent, and is then forgotten, so the reconnecting recipient starts from the next message — and neither side can tell a gap happened.
Why must a deduplication record be written before the guarded work rather than after it, and what must the store offer?
basics
~20 sA retry arrives while the first attempt is still running, so a record written after the work is absent exactly when it is needed. Claim it first, with a conditional create: a write that succeeds only if the key is absent.
A dashboard says an in-memory store answered in 0.3 ms while the caller reports 40 ms for the same calls; what does each number measure?
basics
~20 sThe store's number is server-side execution only, from the server starting the operation to the reply being produced. The caller's number adds waiting for a connection, both network hops and time queued before execution. That gap is where the slowdown lives.
A shared in-memory store reports stored-data size at 80% of its memory ceiling; why can that number not name whose entries to remove?
basics
~20 sStored-data size is one total for the whole keyspace, carrying no owner and no breakdown. Naming a holder is a separate measurement: walk the entries, estimate each one's size, and roll those estimates up per key prefix.
A dashboard for an in-memory store shows a single 'memory used' figure - which four quantities could it be, and what does each mean?
basics
~20 sFour quantities hide behind one 'memory used' figure: stored-data size (the server's count for its entries), resident footprint (what the operating system sees), the memory ceiling the server enforces, and the host or container limit above it.
Callers of an in-memory store are timing out, yet the server's record of operations past its configured time threshold is empty; how is that possible?
basics
~20 sThat record times execution only: its clock starts once the server begins an operation, so waiting never enters it. Callers can be timing out on pool wait, on queueing before execution, or on the network while every operation the server actually ran was genuinely quick.
An in-memory store is brought up with stock settings on a routable host — what reaches it, and what proves a caller's identity?
basics
~20 sStock settings in this class of store assume a private network. Many servers listen on every interface and several offer no identity step at all, so whatever can route to the address can read, overwrite and empty the keyspace.