A worker retains 40 GB of state but touches only 2 GB per cycle and has 16 GB of memory - which placement fits, and what could overturn it?
answer
- total decides possible, hot set decides fast
- no partial fit for live objects
- cache hit ratio at the real cache size
- restore time can outrank throughput
- same 2 GB or a different 2 GB
basics
~20 sForty gigabytes cannot live as objects in a sixteen-gigabyte process, so a disk-backed store with a memory cache fits, and the two gigabytes touched per cycle are mostly cache hits. Restore time or an unstable hot set can overturn it.
solid answer
~50 sRead the three numbers in order. The **total retained bytes** decide what is possible: 40 GB of entries cannot be held as live objects inside a 16 GB process, so the in-process memory placement is out before any performance argument starts. The **hot working set** - the entries actually touched per cycle - decides how the surviving option will perform: 2 GB against a cache of several gigabytes means most accesses are served from memory and only pay an encode and a decode, not a disk read. Then check the two things that override a throughput argument: how long the durable snapshot takes and how long a replacement worker needs before it serves a record, both of which scale with the 40 GB, not the 2 GB; and whether the 2 GB is the *same* 2 GB each cycle, because a scattered access pattern collapses the hit rate to roughly the cache-to-total ratio.
go deeper
Recall that live objects need the whole retained set to fit in the worker process, so a set several times larger than the process memory needs the disk-backed placement instead.
Separate the two sizes out loud: total retained bytes decide what is possible and what a snapshot carries, while the entries touched per cycle decide how the disk-backed placement performs.
Show the override: name restore time, snapshot duration and cache-hit stability as the things that can beat a throughput argument, and say which measurement you would take before committing.
Set the rule others follow: the retained size above which a job must change shape rather than change placement, and the recovery-time budget that makes that threshold a number instead of an opinion.
## Three numbers, three different jobs | quantity | example here | what it decides | |---|---|---| | total retained bytes on this worker | 40 GB | which placements are possible at all; snapshot size; restore time | | the hot working set touched per cycle | 2 GB | how the disk-backed placement will actually perform | | the worker process's memory | 16 GB | the hard ceiling for live objects, and the room available for a cache | The single most common mistake in this decision is using the first number where the second belongs. Total bytes tell you what a snapshot must carry and how long a replacement worker takes to become useful. They do **not** tell you how fast steady-state processing will be, because a disk-backed store with a cache is fast exactly to the degree that the accesses land in the cache. ## Why the total rules out the in-memory placement Holding the retained set as live objects means **all** of it, all the time. There is no overflow path from that placement: 40 GB does not partially live in a 16 GB process. Worse, the process also needs memory for the steps' own work, so the usable share is below the nominal limit, and a live set approaching it makes automatic memory reclamation pauses grow. So the choice here is not really between two placements; it is between the disk-backed store and making the retained set smaller - more workers to spread the keys, or a removal rule so old entries stop accumulating, both of which are other subjects. ## Why the hot set predicts performance Give the worker a cache of, say, 4 GB out of its 16. If each cycle touches the same 2 GB, nearly every access is a cache hit: it pays an encode or a decode, but no disk read, and the placement behaves close to memory. That is the happy case this example is describing. It stops being the happy case in two ways: - **The hot set is not stable.** If the 2 GB touched is a *different* 2 GB each cycle, there is no hot set. The hit rate falls toward the cache-to-total ratio - here roughly one in ten - and almost every access becomes a read of the store's local files. Your steady-state throughput measurement, taken on a warm cache, will not have shown this. - **The hot set grows with distinct keys.** The cache is sized in bytes; the working set grows as more distinct keys are active per cycle. The margin that looks generous today is the first thing consumed by growth. ## What can overturn the answer 1. **Restore time.** A replacement worker must materialise enough of its 40 GB to serve reads. Some engines load the store's files eagerly before processing resumes, some warm lazily or in the background - but the bytes have to arrive from durable storage either way, and the time counts against the job's recovery budget. If the budget is minutes and the restore is tens of minutes, the placement is wrong even though steady-state throughput was fine. 2. **Snapshot duration.** The periodic durable copy scales with the retained set. Where the store's local files are immutable once written, a snapshot can ship only what is new, which makes 40 GB affordable; where it must be written whole each time, it may not be. The mechanism belongs to recovery; the *cost* is a legitimate input to this choice. 3. **A latency target at the tail.** Averages hide the background file merge and the cache misses. If the contract is on the tail rather than the mean, a cache-hit-dominated average does not settle the argument. 4. **The processing model.** A record-at-a-time runtime pays the per-access cost once per record; a runtime that runs a continuous input as repeated short finite jobs touches state far less often per record, and may not hold it in a worker-local store at all. The same three numbers lead to different answers under the two models, so say which you are on. 5. **Growth in distinct keys.** Size the decision against the retained set you expect, not the one you have, because the placement is difficult to change later without discarding state. ## What to measure before committing - Distinct entries touched per cycle multiplied by bytes per entry, for the hot set - not records per second. - The store's cache hit ratio at the cache size you would actually give it, under a realistic key spread. - Tail latency, not mean, with the background file merge running. - Wall-clock time from a worker being replaced to it serving its first record, measured, not estimated. An answer that names the placement without naming the measurement that would change it is the answer an interviewer is probing for.
- How would you actually measure the hot working set?Count distinct entries touched per cycle and multiply by bytes per entry, rather than inferring it from records per second. Better still, run the candidate cache size against replayed traffic and read the hit ratio directly, watching tail latency rather than the mean, because a warm-cache average hides exactly the scattered-access case you are testing for.
- Why does restore time belong in a store-choice argument at all?Because the placement decides what must be present before the first record can be processed after a worker is replaced. That cost scales with total retained bytes, not with the hot set, so it is invisible to a steady-state throughput test. A recovery-time budget can rule out a placement that wins on throughput.
- The 2 GB is a different 2 GB every cycle. What changes?There is no hot set to exploit, so the cache stops being a performance argument: the hit rate tends toward the cache-to-total ratio and most accesses read the store's local files. You then size for disk-read latency on nearly every access, or you attack the retained set itself rather than the placement.
saying these in an interview costs you the question
- Chooses by total retained bytes and never measures what a cycle touches
- Assumes a cache smaller than the working set still gives cache-like latency
- Justifies the placement on steady-state throughput and ignores restore time
- Treats restart as free because a durable copy already exists elsewhere
- Sizes for today's distinct keys with no headroom for growth
- Believes 40 GB can partially fit in a 16 GB process