skip to content

Your platform team's standard answer to a single unit running out of room is to raise every worker's memory — what does that policy cost?

level: principalimportance: should knowfreq 40%

answer

  1. one multiple, charged everywhere
  2. the per-machine total is fixed
  3. fewer, larger processes mean less parallelism
  4. measure the median job, not the incident

basics

~20 s

It buys one fixed multiple of headroom and charges it on every worker process, in every job, for every run. It also trades parallelism when the machine total is fixed, and it hides the shape defect that will return at the next data size.

solid answer

~50 s

Raising the fleet-wide worker size is a purchase, not a fix: the resident thing that failed is sized by the data, so doubling the budget survives a doubling of the data and no more. Meanwhile the cost is paid everywhere — every worker process of every job holds the larger footprint for its whole run, whether it needed it or not. Where a machine's memory total is fixed, larger processes means fewer of them per machine, so parallelism falls and jobs that were healthy get slower. Larger budgets also lengthen reclamation pauses on runtimes that hold records as ordinary language objects, though engines that keep records as packed bytes they manage themselves feel that less. And the job with one indivisible group is never repaired; it fails again, later, louder. Prefer a scoped, expiring override for the one job, and change the default only on measurement of the median job.

go deeper

for a junior

Recall that raising memory is paid by every worker process in the cluster, not only by the job that failed, and that it buys a fixed amount of extra room rather than a permanent fix.

for a middle

Explain the mechanics of the trade: a fixed machine total divides into fewer, larger processes, so parallelism falls, and the resident thing that failed is sized by the data rather than by the setting.

for a senior

Show the scoped alternative — an expiring per-job override with a recorded reason — and the diagnosis you would insist on before any number changes.

for a principal

Own the distribution argument: change a default on the median job's measured footprint, keep overrides visible and expiring, and treat a recurring whole-group formulation as a platform gap rather than a team's mistake.

## What the policy actually buys A single unit of work fails because one thing did not fit in one process. Raising every worker's memory does make that thing fit, if the increase is larger than the shortfall. The question a lead has to answer is what was bought and at what price. What was bought is **one multiple**. If the failure was a group that must be held whole, the group's size is a property of the data — how many records carry that key, and how large each is. Doubling the budget survives a doubling of the group. It does not survive the following quarter, and it does not survive the next job that hits the same shape with a different key. Purchasing headroom against data growth is a losing race that has to be re-run every time it is lost. ## What it costs, and where The cost is not paid by the job that failed. It is paid by the fleet: - **Every worker process of every job** holds the larger footprint for the whole of its run, whether its working set needed the room or not. Multiply by process count and by hours to get the real number. - **Parallelism falls where the machine total is fixed.** A machine with a fixed amount of memory divides into fewer, larger worker processes. Fewer processes means fewer units of work in flight, and jobs that were perfectly healthy take longer — a slowdown attributed to the cluster rather than to the change that caused it. - **Reclamation pauses lengthen** where the runtime holds records as ordinary language objects, each carrying the runtime's own per-object bookkeeping and each reclaimed automatically: a larger region takes longer to sweep, and the worker's own work stops while it happens. Engines that keep records as a packed byte layout they allocate and interpret themselves are much less exposed to this, so the size of the effect depends on which representation your platform is standardised on. - **Scheduling gets harder.** In a shared pool, a job asking for larger processes waits for machines that can host them; bigger indivisible requests pack worse. Who grants that share is the resources subject, but the queueing is a real consequence of this decision. - **The defect is hidden.** The job whose step must hold a whole group is now passing, so nobody reformulates it, and the reformulation is the thing that would have made it durable. | Lever | Blast radius | Buys you | |---|---|---| | Reformulate the failing step | One job, one team | A fix that survives data growth | | Scoped override for that job | One job, on request | Time, with the defect still recorded | | Fewer units per worker | The jobs that opt in | Room per unit, at a cost in parallelism | | Fleet-wide default raise | Every job, every run | One multiple, charged to everyone | ## When raising the default is the right call It is not always wrong, and a lead who refuses it on principle is as unhelpful as one who reaches for it first. It is defensible when: 1. **Measurement says the shape is wrong for the median job**, not for the loudest one. If peak footprint across most jobs sits near the ceiling and spilling is routine and costly, the fleet is genuinely undersized and every job is paying in time for memory it should have had. 2. **The machines have memory that nothing can use.** If the per-machine total is not the binding constraint — cores are — then larger processes cost nothing in parallelism and the objection above does not apply. 3. **The engineering cost of the fix genuinely exceeds the memory cost**, which does happen for a job with weeks of life left in it. Then make it explicit and temporary rather than permanent and silent. ## The better default, and the thing to institutionalise The useful distinction is blast radius. A scoped override that applies to one job, carries a written reason and has an expiry date, keeps the cost proportional to the problem and keeps the defect visible. A fleet-wide raise is a decision about every job the organisation runs, and should be justified by a distribution rather than by an incident. Three things worth owning at this level: - **A record of every override and why.** Without it, the fleet default creeps upward by accumulation, one incident at a time, and nobody can say which increases are still needed. - **A supported way to express the operations that keep failing.** If teams keep writing whole-group possession because the alternative is awkward, the platform is the reason the failure recurs. - **A standard diagnosis path**, so that the failing unit's own numbers are read before anything is bought. The tell of a weak organisation here is not that it raises memory; it is that it raises memory without ever having looked at what the failing unit contained.

  • How would you decide the number instead of guessing it?
    From the distribution, not the incident: collect peak footprint and spilled bytes per unit of work across a representative week, look at where the mass sits and how much time spilling is actually costing, then size so the median job is comfortable and outliers get scoped overrides rather than a new floor for everyone.
  • Would running fewer units of work inside each worker be a better lever?
    Often yes, and it is cheaper to reverse. It gives each unit a larger share of the same budget without changing any machine's shape, and it is scoped to the jobs that ask for it. The price is parallelism inside each worker, so the job runs longer for the same total resource.

saying these in an interview costs you the question

  • Treats a fleet default as a cheap, local change
  • Sizes the default from one incident's peak
  • Ignores that a fixed machine total means fewer processes
  • Leaves overrides permanent and unrecorded
  • Claims more memory removes the need to reformulate the step