skip to content

Debugging Remote Work

Diagnosing logic that ran on machines you no longer have: logs scattered across workers and gone with them, counters carried out of the run, and one failing record reproduced locally.

on this pageshow

questions

5

Why can the diagnostic lines a worker process printed be unrecoverable once a job run has ended and its machines are released?

level: juniorimportance: must knowfreq 70%

answer

  1. the machines were only borrowed
  2. local disk leaves with the host
  3. what was carried out survives
  4. decided before the failure, not after

basics

~10 s

A worker's diagnostic lines are written on the machine that ran it, and the cluster releases that machine when the work ends. Only what the run carried off its machines stays readable.

solid answer

~50 s

A **worker process** is one process on one machine that runs pieces of the job and owns the memory they use. When it prints a diagnostic line, that line lands on its own host: the process output, or a file on that machine's local disk. Neither is part of the job. Once the machine is released at the end of the run, or reclaimed mid-run, its local disk goes with it, and a process that crashed cannot flush anything afterwards. Some deployments forward lines off the host while it is still alive, but that forwarding is a separate system and this job does not depend on it. What is reliably there afterwards is what the run was built to carry out of itself: totals summed by the coordinating process, records written to shared storage, and whatever run figures the engine persists.

go deeper

for a junior

Recall that a worker's printed output lives on that worker's machine, and the cluster gives the machine back. Anything you want to read tomorrow has to leave the machine today.

for a middle

Explain the mechanics: local disk against shared storage, what the coordinating process retains, and why a crashed process is the worst case because its last lines may never be flushed.

for a senior

Show that you plan evidence in advance — a small set of totals reported as units finish, record-shaped material written to shared storage under the run's identity, and no diagnosis that depends on a live host.

for a principal

Frame it as a standing cost: every run pays for the evidence you always carry, and the alternative is an unexplained wrong number that consumers, not the owning team, absorb.

## The machines were borrowed A **job run** is one submission of one program to the cluster, from start to its last written output. It is executed by **worker processes** — each one process on one machine, running pieces of the job and owning the memory those pieces use — and coordinated by a single **coordinating process**, which turns the program into pieces, hands them out and tracks which finished, and whose death ends the run. When code inside a piece prints something, the text goes where the code is running: the worker's own process output, or a file on that machine's local disk. It is not routed into the job, it is not part of the result, and the engine makes no promise about it. The host was borrowed for the duration of the work, and the writing sits on the borrowed thing. ## Four ways the lines go away 1. **The run ends.** The cluster releases the machines it held. On a cluster raised for one job this is immediate; on a standing pool the process exits and its files become rubbish someone will clean. 2. **The worker process dies.** The failure you want to read is frequently the reason the process stopped, so the last and most interesting lines are the ones least likely to have been flushed anywhere. 3. **The machine is reclaimed mid-run.** On an **interruptible worker** — a cheaper machine the provider may take back at any moment — the host can vanish while the job is healthy, and the run simply redoes that work elsewhere. 4. **The disk is reused.** A worker's local disk also holds *spill*, the bytes it writes out when a working set will not fit in memory, and the intermediate files of a redistribution. That space is recycled aggressively; diagnostic files are not privileged over it. ## What survives, and what does not | Evidence | Where it lives during the run | Readable after the machines go | |---|---|---| | A worker's printed diagnostic lines | that worker's host | No, unless something moved them while the host was alive | | The worker's in-memory working set | the worker process | No | | Intermediate files on the worker's local disk | that worker's host | No | | A **carried counter** — a number every worker adds to and the coordinating process sums | the coordinating process | Usually, alongside the run's figures | | Records the job wrote to **shared storage** (the storage every machine can read, where input and output live) | shared storage | Yes | | Per-unit status and duration | the coordinating process | Depends on the engine | ## Where engines genuinely differ Do not assume one model here; this family disagrees. - **Whether a finished run's figures outlive it.** Some engines write a record of the run's numbers that a later reader can open; others expose them only through the live coordinating process, so the screen goes blank with the run. - **How much per-unit detail is kept.** A long run with hundreds of thousands of units is often summarised rather than preserved item by item. - **Whether intermediate data is durable at all.** In the two-phase disk-to-disk model, every phase writes its whole output to shared storage before the next phase reads it, so the intermediate data is itself evidence you can go back to. A record-at-a-time runtime materialises nothing between steps, and a continuous job built from repeated small finite runs keeps only what each slice committed. The same question has three different answers. ## The consequence: it is a decision, not a lookup Because the host is temporary, *what you will be able to see* is fixed before anything goes wrong. Practically: - Decide up front the few numbers worth carrying — records read, records emitted, records rejected by reason — so a total survives the worker that produced it. - Write anything record-shaped to shared storage under the run's identity, never to a worker's local disk, if you expect to read it tomorrow. - Expect the coordinating process's summary to tell you *that* a unit failed and roughly where, not *which record* did it. - Treat a diagnosis plan that begins with "we will look at the worker's output" as a plan that works only while the run is still alive. The common junior mistake is not misunderstanding the mechanism; it is assuming somebody has already solved it. On a cluster you do not own, nobody has, unless your team arranged it.

  • The run succeeded but the numbers look wrong, and the machines are gone. What can you still ask of it?
    The totals the coordinating process summed, the output records themselves in shared storage, and whatever per-unit status and timing the engine persisted. Those three are enough to say how much went in, how much came out and which units were unusual — but not to point at a record. If your program carried no totals, the output is the only witness left.
  • Does running on interruptible workers change the picture?
    It makes it sharper. A machine the provider can reclaim at any moment may disappear while the job is perfectly healthy, taking anything held only on it. Evidence must therefore be carried continuously — reported as each unit of work completes — rather than gathered at the end, because there may be no end for that host.
  • Why is a crashed worker the hardest case of all?
    Because the lines that would explain the crash are the last ones written, and a process that dies abruptly may never flush them. Even where something forwards output off the host, the forwarding is asynchronous, so the tail is exactly the part most likely to be missing.

A crew works out of a temporary site office. Notes pinned to its wall are the fastest way to see what happened today, but when the job ends the office is taken away and the wall goes with it. Only what was posted back to head office is there next year.

saying these in an interview costs you the question

  • Assuming every worker's output is collected somewhere by default
  • Planning to inspect a released machine's local disk after the run
  • Treating the coordinating process's summary as holding each worker's detail
  • Believing a reclaimed machine's files stay readable once it is taken back
  • Confusing the job's own output records with evidence about the run
open as a page

A run fails on one step of the job graph and the error names no record. How do you narrow it to a single input piece?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Start from what the failure names: the step and the failed unit of work. If the same unit fails each attempt, map it back to its input range, then run the step's function over that range in one process.

open as a page

Why is a counter that every worker adds to and a coordinating process sums better evidence than a line printed on a worker?

level: middleimportance: should knowfreq 58%

basics

~20 s

A carried counter is summed by the one process that outlives the workers and is normally written out with the run's figures, so the evidence survives the machine that produced it. Its limit is that it gives a total, never a record.

open as a page

A step rejects a small fraction of records in every run. What do you arrange in advance so the rejected ones can be examined?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Count every rejection, but keep only a bounded sample of the records themselves: a fixed cap per worker process per run, written to shared storage under the run's identity, with the reason attached and sensitive fields left out.

open as a page

Which evidence should a team carry on every job run, and which should it capture only on a deliberate re-run?

level: principalimportance: should knowfreq 40%

basics

~20 s

Carry always what is small and impossible to reconstruct afterwards: a few totals per step, the run's identity and shape, per-unit status. Raise the expensive evidence on demand, and only where a re-run can genuinely recreate the failure.

open as a page