Ten instances each hold an in-process copy of one 2 GB working set — what does that cost in memory, and what does a read cost?
answer
- count processes, not datasets
- one copy per instance
- instance count multiplies the bytes
- local read pays no network
- per-instance limit, not fleet total
basics
~20 sTen copies of a 2 GB set cost roughly 20 GB across the fleet, one copy inside each process, charged against each instance's own memory limit. The read itself costs no network hop: it is a local lookup.
solid answer
~50 sThe copy is held per process, so the memory bill scales with the instance count rather than with the size of the data: ten instances holding the same 2 GB set spend about 20 GB, and each 2 GB has to fit inside one instance's limit alongside everything else that process needs. Filling it lazily does not divide that by ten — behind a load balancer spreading requests evenly, every instance is eventually asked for most of the hot set, so each copy converges on most of it. What you buy is the cheapest possible read: no network, no serialisation, no second component that has to be up. What you do not buy is one answer, because ten copies were filled at ten different moments. A shared tier inverts both terms: one copy for the whole fleet, read alike, at the price of a hop and a dependency.
go deeper
Remember where the copy physically lives: inside each application process. Ten processes mean ten copies and ten times the bytes, and a read never leaves the process.
Do the arithmetic out loud and name what each placement buys: bytes times instance count with no hop, against one copy for the fleet behind a hop and a dependency.
Size it against the per-instance memory limit at peak instance count rather than a fleet total, and say what those bytes displace inside the process.
Once the fleet autoscales, the instance count is an input nobody fixes. The placement rule has to hold at the top of the scaling range, not at today's ten.
## Two placements for one copy A fast copy of data can sit in one of two places. An **in-process copy** is allocated inside the application process: the structure holding it is on the same heap as the code that reads it, and a read is a lookup in local memory. A **shared tier** is a separate component with its own address and its own memory, which every application instance reaches over the network. Both answer the same read. Almost everything about what they cost is different, and the memory arithmetic is the part an interviewer will ask you to do out loud. ## The arithmetic The unit that holds an in-process copy is the **process**, not the service. Ten instances are ten processes with ten heaps, so: - A 2 GB working set held in-process by ten instances costs roughly **2 GB x 10 = 20 GB** across the fleet. - Each of those 2 GB is charged against **one instance's own memory limit**, next to request buffers, connection state and everything else the process needs. The fleet total is the easy number; the per-instance fit is the one that takes the service down. - The growth term is the **instance count**, not the data. Doubling the fleet doubles this cost while the data stays exactly the same size. - Filling the copy lazily does not divide the bill by ten. Behind a load balancer that spreads requests evenly, every instance is eventually asked for most of the hot set, so each copy converges on most of it. Lazy filling delays the cost and shapes it to real demand; it does not shard it. - Per-entry bookkeeping is paid per copy too, so whatever overhead one entry carries, the fleet carries ten of it. A shared tier inverts the growth term: one 2 GB copy serves the whole fleet, and adding instances adds request load to that component rather than bytes to it. | | In-process copy, ten instances | One shared tier | |---|---|---| | Bytes for a 2 GB set | about 20 GB, one copy per process | about 2 GB, once | | Who pays for them | each instance's own limit | the tier's own memory budget | | Read path | lookup in local memory | a hop across the network | | Do instances agree | no, ten copies filled independently | yes, all read the same entry | | Extra dependency | none | the tier has to be reachable | ## What a read actually costs An in-process read crosses no network and needs no serialisation: the caller reads a value it already holds. That is the whole case for the placement, and it is a real one. Two precisions are worth stating before you claim it: 1. **In-process means the same address space.** A copy held by a sidecar process on the same host is *not* in-process. That read is still a round trip — a cheap one over the local interface, but a round trip with its own serialisation, its own failure mode, and one copy per host rather than per process. 2. **Cheap is not free.** Those bytes displace heap the process could have used to serve requests, and a much larger heap changes how the runtime behaves when memory gets tight. ## What the arrangement does not buy The ten copies were filled at ten different moments from whatever the system of record held then. Nothing in the arrangement reconciles them, so the fleet holds ten versions of the answer at once and which one a caller sees depends on which instance took the request. For a read where that matters, the in-process placement is the wrong one no matter how cheap it is — and no amount of memory arithmetic tells you whether it matters. ## Where the class varies Two things here belong to a specific store rather than to the arrangement, and a good answer scopes them instead of asserting them: - **What happens at the memory ceiling.** Some in-memory stores reclaim entries when they run out of room; others refuse the write and return an error to the caller. How the shared tier behaves when it fills up is a property of the store that was chosen, not of the fact that it is shared. - **Whether the software can be placed either way at all.** Some products in this class run only as a server you reach over the network; others ship as a library you can embed in the process. Whether "in-process" and "shared" are even the same software is a deployment question, and an answer should not assume they are. ## Saying it in an interview 1. State the multiplier first: bytes times instance count, charged per instance, not per service. 2. State the read cost: no hop for the in-process copy, one hop for the shared tier. 3. State what the arithmetic does not cover: ten copies are ten answers, one shared copy is one answer.
- Does filling the in-process copy lazily reduce the fleet's memory bill?Only as far as instances are asked for genuinely different keys. Behind a load balancer that spreads requests evenly, every instance is eventually asked for most of the hot set, so each copy converges on most of it. Lazy filling delays the cost and shapes it to real demand; it does not divide it by the instance count.
- Is a copy held by a sidecar process on the same host the same thing as an in-process copy?No. In-process means the same address space as the caller, so the read is a local lookup. A sidecar is a separate process: the read is still a round trip, cheap over the local interface but with its own serialisation and its own failure mode, and the copy is held once per host rather than once per process.
- The 2 GB does not fit inside the instance limit. What does that rule out?It rules out the in-process placement for that working set, not the tier. Either the copy is reduced to a subset small enough to fit every instance, or the data is read from a shared tier where one copy is held once and sized on its own budget. Raising every instance's limit to hold a full copy pays for the data ten times.
saying these in an interview costs you the question
- Sizes the copy once for the service instead of once per instance.
- Thinks instances running the same build share one copy in memory.
- Assumes each instance only holds its own tenth of the hot set.
- Thinks lazy filling divides the memory bill by the instance count.
- Calls an in-process copy free because there is no round trip.
- Treats a value read from one instance's copy as the fleet's answer.