skip to content

A long job pins several intermediates and its later steps spill heavily on unchanged input — what should you check first?

level: seniorimportance: should knowfreq 44%

answer

  1. a pin has a lifetime
  2. nothing expires it automatically
  3. count live pins per phase
  4. spill rises with phase number
  5. release after the last reader

basics

~20 s

Check whether earlier pins were ever released. A pinned result occupies retained-result memory until the job releases it or the engine drops it, so copies from finished phases can still be squeezing the operators of later steps on every worker.

solid answer

~50 s

A pin has a lifetime, and nothing about a reader finishing ends it. Unless the job releases the pinned result, the copy stays in each worker's retained-result share for the rest of the run — so by the fifth phase the workers may be holding four intermediates nobody will read again, and the operators of the current step get whatever is left. The symptom is exactly this: unchanged input, rising spill figures in later steps, and held bytes that only ever grow. Two things make it worse. A loop that pins a fresh result each round accumulates one per iteration. And on engines where a pin is a request rather than a guarantee, the engine starts dropping pinned results to make room for operator work, so the branches are recomputed while their successors are still displaced — you pay both. The fix is to release each pin after its last reader, and to keep the number of live pins small enough to name.

go deeper

for a junior

Remember that asking the engine to keep a result is not temporary: it stays held until the job says otherwise, so keeping many results at once leaves less room for the work still to come.

for a middle

Explain the accumulation: pins from finished phases still occupy retained-result memory, so later operators get less working memory and write part of their working set to local disk.

for a senior

Diagnose it from the run's own numbers — spill rising with phase number, held bytes never falling, early branches reappearing — and pair every pin with its last reader before touching any sizing.

for a principal

Own the policy: who may pin on a shared pool, how pins are released, and what the worker shape assumes about how much of a budget results are allowed to hold, since an unreleased pin taxes every other job on that machine.

## A pin outlives its last reader **Pinning a result** tells the engine to keep a computed intermediate so later steps read it instead of recomputing the branch above them. What it does *not* do is expire. The copy is held in each **worker**'s **retained-result memory** — the part of that process's fixed budget holding results the job was told to keep — and it stays there until one of three things happens: 1. the job explicitly releases it; 2. the engine drops it to make room for operator work, on engines where a pin is a request rather than a guarantee; 3. the job ends. Nothing about the last reader finishing is one of those. The engine cannot know that no later step will read the intermediate again — the program may still describe steps that have not run. So in a long job with phases, the natural outcome is accumulation: phase one pins a prepared set, phase two pins a joined set, phase three pins an aggregated set, and by phase five every worker is holding three copies that nobody will read again, while phase five's operators draw on what is left of the same budget. ## The loop that never releases The sharpest version is iterative work. Each round derives a new result from the previous one and pins it, because each round's result is read more than once inside the round. If nothing releases the previous round's copy, the held bytes grow linearly with the iteration count while the working set of the current round stays the same size. The job runs well for the first several rounds and then degrades — with no change in the input, which is exactly what makes it confusing to diagnose. ## What the symptoms look like Read them from **the run's reported numbers** — whatever the engine reports per unit of work after a run: - **spill figures rising with phase number.** Operators in later steps cannot hold their working set and write part of it out, in sorted runs or per-key chunks, to **local scratch disk** — disk attached to the worker machine whose contents nobody may read once the job ends. - **held bytes that only ever climb**, never falling between phases. - **branches of early phases reappearing** in later steps. That is the double-payment signature: the engine dropped a pinned result under pressure and is recomputing it, while other pins keep displacing operator work. - **lengthening reclamation pauses** — the runtime stopping the worker's own work while it finds and frees objects nothing refers to — on engines that keep results as ordinary objects of the worker's language, because long-lived objects are exactly what makes that work harder. - **a worker killed at the ceiling the platform enforces** while the engine's own accounting says it had room. Held copies kept outside the accounted regions add to the process's real footprint, and the platform enforces its limit by killing the whole process rather than by failing one allocation. ## What engines do when retained-result memory fills This genuinely differs, and the difference decides which symptom you see: | engine behaviour when the held share is full | what you observe | |---|---| | drops the least recently used held result and recomputes on next read | early branches reappear; work repeats intermittently | | overflows the surplus to local scratch disk instead of dropping it | disk reads on every use of the copy; spill figures blurred with real spilling | | holds a packed byte layout it manages itself, and refuses beyond it | a hard failure or a fallback the engine reports explicitly | | has almost no notion of held results at all | the question does not arise; there is nothing to leak | Do not reason from the behaviour of the one you know. Establish which of these your run is showing before choosing a repair. ## Releasing, and where to put the release - **Release immediately after the last reader**, not at the end of the program. The whole point is to shorten the window during which the copy displaces operator working memory. - **In a loop, release the previous round's copy after the current round has been computed from it** — not before, or the round's own readers walk the branch again. - **Keep the live pins countable.** If you cannot say out loud which intermediates are currently held and who still reads them, there are too many. - **Prefer a narrower copy.** Project away unused columns before pinning; the copy shrinks in proportion and so does the window's cost. - **Re-examine whether each pin earned its place at all.** A branch that is a local, selective scan of a durable source is usually cheaper to walk again than to hold. ## The check to run first Before tuning anything, list every pin the job takes and pair it with the step that is its last reader. Most jobs in this state have at least one pin with no reader left at all, and releasing those is free — it costs nothing and returns memory to the operators that are spilling. Only when the list is clean is it worth arguing about worker shape or region boundaries.

  • Why does the engine not release a pinned result once its last reader has finished?
    Because it cannot know there is no later reader. The program is a description, and steps that have not run yet may still name the intermediate. Releasing on the engine's own judgment would mean sometimes recomputing an expensive branch the author explicitly asked it to keep. The lifetime is therefore the author's to manage, which is exactly why unreleased pins are such a common cause of late-phase pressure.
  • The held bytes are climbing and early branches are being recomputed at the same time. What is happening?
    You are paying on both sides. On engines that treat a pin as a request, the engine drops held results to make room for operator work and recomputes them on the next read. So some copies still occupy memory and displace operators, while others have been discarded and their branches walked again. Releasing the pins that no longer have readers usually resolves both symptoms at once.

saying these in an interview costs you the question

  • The engine releases a held result once nothing reads it
  • A pin only matters during the step that created it
  • Held copies cannot cause spilling, because they are separately accounted
  • Unchanged input means the slowdown must be the cluster's fault
  • Holding more results is safe as long as nothing has failed yet
  • Releasing a pin risks losing data the job still needs